Video processing method, electronic device, and storage medium
By performing keyframe overview analysis and frame-by-frame analysis of target areas in video footage, the problem of low processing efficiency of electronic devices for long footage is solved, and the user experience of the one-click video creation function is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-01-10
- Publication Date
- 2026-04-28
AI Technical Summary
When the user selects a long video clip, the electronic device takes a long time to analyze multiple image clips, resulting in low video processing efficiency and affecting the user experience of the one-click video creation function.
By performing partial keyframe analysis on each video in the image material, an overview analysis is first performed to determine the target area, and then frame-by-frame analysis is performed in the target area to extract highlight segments, reducing the workload of overall video analysis.
It improves the efficiency of electronic devices in processing videos, reduces the time required for one-click video creation, and enhances the user experience.
Smart Images

Figure CN120343182B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video data technology, and in particular to a video processing method, electronic device, and storage medium. Background Technology
[0002] With the development of image and video processing technologies, users can trigger electronic devices to further process photos and videos in their albums. For example, electronic devices can stitch together multiple image materials (including multiple pictures and / or multiple videos) to obtain a complete stitched video.
[0003] For example, electronic devices or third-party video processing software within them can have a one-click video creation function (or a one-click video creation service). This function automatically analyzes and extracts highlight segments from multiple image clips selected by the user using algorithms; then, it automatically generates an edited video based on these extracted highlight segments. These highlight segments, also known as "highlight clips," refer to single-frame images or video clips composed of multiple consecutive frames extracted from the aforementioned materials to record exciting moments. These highlights could be moments of remarkable action, such as a person smiling, a moment of victory, or an airplane landing.
[0004] However, when the overall duration of the footage selected by the user is long, it takes a lot of time for electronic devices to analyze the multiple footage selected by the user, resulting in low efficiency in video processing. Summary of the Invention
[0005] This application provides a video processing method, electronic device, and storage medium. By performing partial keyframe analysis on each video in the image material, it avoids the problems of high time consumption and low video processing performance caused by analyzing the complete video.
[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions.
[0007] Firstly, a video processing method is provided, the method comprising:
[0008] The electronic device receives a user's selection of multiple image resources from its gallery. These multiple image resources may include videos; alternatively, they may include both videos and pictures.
[0009] In response to a selection operation, the electronic device performs a first-number of I-frame overview analysis on each video from multiple image materials to obtain a first score. The first score is an aesthetic score for the corresponding image frame, and the first number corresponds to the video duration; an I-frame is an image frame that includes complete image information.
[0010] The electronic device determines the target region of each video based on the first score of each I-frame. The target region includes the first I-frame, as well as I-frames before and after the first I-frame within a first preset duration. The first I-frame is the I-frame with the highest first score. Frame-by-frame analysis is performed on the target region of each video in multiple image materials to obtain the second score of each image frame in the target region. Frame-by-frame analysis includes I-frame analysis or individual image frame analysis, and the second score is an aesthetic score for the corresponding image frame.
[0011] The electronic device determines the highlight segment of the corresponding video based on the second score corresponding to each video; wherein the highlight segment includes the second image frame, as well as frames within a second preset time period before and after the second image frame.
[0012] Among them, the second image frame is the second highest-scoring image frame in the target region; highlight segments from each video in multiple image materials are used to stitch together the target video set.
[0013] In this application, the electronic device first performs an overview analysis on the video to locate the target area that needs to be processed frame by frame. When extracting highlight segments, the electronic device only performs frame-by-frame analysis on the target area, rather than analyzing the entire video frame by frame. This reduces the workload of image frame analysis, thereby reducing the time required for the electronic device to achieve one-click video creation, improving the efficiency of video processing, and ultimately enhancing the user experience of the one-click video creation function.
[0014] In another possible implementation of the first aspect, a first number of I-frame overview analyses are performed on each video in the multiple image materials to obtain a first score, including:
[0015] Obtain the first position of the analyzed I-frame in the video;
[0016] Based on the first position and first score of the analyzed I-frames, the second position of the I-frame to be analyzed in the video is determined, and the first score of the I-frame at the second position is obtained, until the number of analyzed I-frames reaches a first number.
[0017] In this application, the electronic device can determine the first number of image frames that can be analyzed for each video within the suggested upper limit of the total analysis time. Under the limited performance of the electronic device, by analyzing the first number of image frames, the analysis efficiency of highlight segments can be improved, and a relatively reliable analysis effect can also be obtained.
[0018] In another possible implementation of the first aspect, the first position includes a first preset position and a second preset position of the video.
[0019] Based on the first position and first score of the analyzed I-frame, the second position of the I-frame to be analyzed in the video is determined, including:
[0020] The area between the midpoint between the first preset position and the second preset position and the starting position of the video is defined as the first region; the first region includes the I-frame of the first position, and the score of the I-frame of the first region is defined as the first score of the I-frame of the first preset position.
[0021] The area between the midpoint and the end of the video is designated as the second region; the second region includes the I-frame at the second preset position, and the score of the second region is used as the first score of the I-frame at the second preset position.
[0022] Calculate the product of the score for each region and the duration of the region, and use the midpoint of the region with the highest product as the second position.
[0023] In this application, based on the scores of the analyzed I-frames and the first position of the I-frames, regions with higher scores can be identified in the video. From these regions, the second position of the I-frame to be analyzed can be determined, thereby improving the quality of the I-frames sent for analysis and indirectly enhancing the effectiveness of the I-frame analysis results.
[0024] In another possible implementation of the first aspect, the first preset position is the 1 / 3 position of the video duration, and the second preset position is the 2 / 3 position of the video duration.
[0025] In this application, the first preset position is set at 1 / 3 of the video duration, and the second preset position is set at 2 / 3 of the video duration. This directly and effectively determines the quality of the region from the start position to the midpoint and the region from the midpoint to the end position of the video, thereby further determining the second position of the I-frame to be analyzed. This avoids the problem of low analysis effectiveness caused by randomly selecting the first position.
[0026] In another possible implementation of the first aspect, the first position is a non-preset position. Therefore, the first position is the position of the first I-frame sent for analysis (not the video frame), and the first position is the position of a specific I-frame during the I-frame analysis of the video.
[0027] Based on the first position and first score of the analyzed I-frame, the second position of the I-frame to be analyzed in the video is determined, including:
[0028] A third region containing the I-frame at the first location is determined based on the adjacent boundaries before and after the I-frame at the first location.
[0029] The adjacent boundary includes one of the following: the position of the I-frame adjacent to the first position and already analyzed, the start position of the video, the end position of the video, or the boundary of a connected region. A connected region is a region with the same score and is temporally continuous, with no overlap between connected regions. The third region is one of the connected regions in the video.
[0030] Obtain multiple connected regions from the video and calculate the score for each region. The score of a connected region is the average of the first score of the included I-frames and the scores of the connected regions. Obtain the product of each region's score and its duration, and use the midpoint of the connected region with the highest product as the second position.
[0031] In this application, during the I-frame overview analysis of a video, connected regions in the video can be dynamically determined based on the scores of the analyzed I-frames and their first positions. The second position of the I-frame to be analyzed is then determined from the connected region with the largest product of region score and region duration. The I-frame at this second position will not have a low score, thus indirectly improving the quality of the I-frames sent for analysis and further enhancing the effectiveness of the I-frame analysis results.
[0032] In another possible implementation of the first aspect, determining the third region corresponding to the I-frame at the first position based on the adjacent boundaries before and after the I-frame at the first position includes:
[0033] The first boundary of the third region is determined based on the first score and corresponding first weight of the I-frame at the first position, and the scores and corresponding second weights of the adjacent boundaries preceding the I-frame at the first position. The second boundary of the third region is determined based on the first score and first weight of the I-frame at the first position, and the scores and corresponding third weights of the adjacent boundaries following the I-frame at the first position. The scores and weights are correlated.
[0034] In this application, the boundary of the third region of an I-frame is determined based on the I-frame score and weight, as well as the scores and weights of adjacent boundaries of the I-frame. This makes the obtained third region more representative of the I-frame score, resulting in more accurate scoring of each region during the subsequent determination of the second position of the I-frame to be analyzed based on connected regions. The determined I-frame for analysis is thus more effective. The more accurate the target region determined based on the I-frame, the better the effect of determining the highlight fragments through frame-by-frame analysis based on the target region.
[0035] In another possible implementation of the first aspect, the target region of the corresponding video is determined based on the first score of the I-frame in each video, including:
[0036] For each video, obtain the first I-frame based on the first score of the I-frame in the video. Determine the candidate highlight segments of the video by including the first I-frame and having a duration equal to the preset suggested duration of the highlight segment. Define the candidate highlight segments and the I-frames within a third preset duration before and after the candidate highlight segments as the target region of the video; the third preset duration is less than the first preset duration.
[0037] In this application, the electronic device, based on the first image frame with the highest score and the suggested duration of the highlight segment, can determine a candidate highlight segment that matches the suggested duration, thereby identifying a target region containing the candidate highlight segment. Since this target region includes the first image frame, it warrants further analysis to determine the highlight segments in the video. Therefore, the determined highlight segments are more accurate.
[0038] In another possible implementation of the first aspect, the method further includes:
[0039] In response to the selection operation, the analysis parameters corresponding to multiple image materials are obtained. The analysis parameters include the actual duration of the corresponding video and the suggested upper limit value of the total analysis duration. The suggested upper limit value of the total analysis duration represents the suggested maximum duration required to analyze multiple image materials.
[0040] The first number of I-frames for each video is determined based on the actual duration of each video in the multiple image materials, the number of videos in the multiple image materials, and the recommended upper limit of the total analysis time.
[0041] In this application, the electronic device can determine the first number of image frames that can be analyzed for each video within the suggested upper limit of the total analysis time. Under the limited performance of the electronic device, by analyzing the first number of image frames, the analysis efficiency of highlight segments can be improved, and a relatively reliable analysis effect can also be obtained.
[0042] In another possible implementation of the first aspect, the first number of I-frames for each video is determined according to the actual duration of each video in the multiple image materials, the number of videos in the multiple image materials, and the suggested upper limit of the total analysis duration, including:
[0043] Based on the actual duration of each video and the preset correspondence, the basic number and maximum number of corresponding videos are determined. The basic number is the minimum number of I-frames required to ensure the analysis effect of the video, and the maximum number is the maximum number of I-frames that the duration allows for the analysis of the video. The preset correspondence represents the maximum number and basic number of I-frames corresponding to different threshold ranges of video duration.
[0044] The total number of analyses is determined based on the base and maximum number of analyses for each video, as well as the recommended upper limit for the total analysis time. The total number of analyses is the total number of I-frames that can be analyzed from all videos across multiple image materials.
[0045] Based on the actual duration of each video in the multiple image materials and the number of videos in the multiple image materials, the total number of analyses is allocated to each video, and the first number of I-frames of each video is obtained.
[0046] In this application, a first quantity is determined by the maximum number and the basic number of image frames in each video. The first quantity of image frames is then analyzed, which can ensure the analysis effect of the video and perform efficient analysis and processing of highlight segments under the limited performance and time of electronic devices.
[0047] In another possible implementation of the first aspect, the analysis parameters also include single-frame analysis duration; the single-frame analysis duration is the duration required to analyze one image frame.
[0048] Based on the base and maximum number of videos, and the suggested upper limit for the total analysis time, the total number of analyses is determined, including:
[0049] If the recommended upper limit for the total analysis time is less than the sum of the base number of I-frames in all videos across multiple image materials, the total number of analyses is the ratio of the recommended upper limit for the total analysis time to the analysis time of a single frame. If the recommended upper limit for the total analysis time is greater than the sum of the base number of I-frames in all videos across multiple image materials, and the duration of a second multiple of the recommended upper limit for the total analysis time is less than the sum of the base number of I-frames in all videos across multiple image materials, the total number of analyses is the sum of the base number of I-frames in all videos across multiple image materials.
[0050] If the duration of the second-multiplied recommended value of the total analysis duration is greater than the sum of the base number of I-frames in all videos across multiple image clips, and the duration of the second-multiplied recommended value of the total analysis duration is less than the sum of the maximum number of I-frames in all videos across multiple image clips, then the total number of analyses is the ratio of the duration of the second-multiplied recommended value of the total analysis duration to the analysis duration of a single frame; if the duration of the second-multiplied recommended value of the total analysis duration is greater than the sum of the maximum number of I-frames in all videos across multiple image clips, then the total number of analyses is the sum of the maximum number of I-frames in all videos across multiple image clips.
[0051] Among them, the second multiplier is greater than 0 and less than 1.
[0052] In this application, a first quantity is determined by the maximum number and the basic number of image frames in each video. The first quantity of image frames is then analyzed, which can ensure the analysis effect of the video and perform efficient analysis and processing of highlight segments under the limited performance and time of electronic devices.
[0053] In another possible implementation of the first aspect, the total number of analyses is allocated to each video according to the actual duration of each video in the multiple image materials and the number of videos in the multiple image materials, and the first number of I-frames of each video is obtained, including:
[0054] Iterate through each video in multiple image materials, updating the first value and the second value of each video until the first value is 0; where the initial value of the first value is equal to the total number of analyses, and the initial value of the second value is 0.
[0055] In this process, for each video, the second value of the video is incremented by 1, and the first value is decremented by 1.
[0056] Use the second value of each video as the first number of I-frames in the video.
[0057] In this application, each video is traversed to determine the first number of each video, which can reasonably allocate the total number of analyses to each video, so that no highlight segments of any video are missed.
[0058] In another possible implementation of the first aspect, each video in the multiple image materials is iterated over, and the first value and the second value of each video are updated until the first value is 0, including:
[0059] Before iterating through the first video among multiple image materials, if the second value of the first video is equal to the maximum number of I-frames in the first video, then skip the first video and iterate through the next video. Skipping the first video means that the second value of the first video is not incremented by 1.
[0060] In this application, if the number of image frames in a video reaches the maximum, no further image frames will be allocated to that video. Instead, the image frames will be allocated to other videos that have not reached the maximum number, which can improve the effectiveness of video analysis.
[0061] In another possible implementation of the first aspect, a frame-by-frame analysis is performed on the target region of each video in multiple image materials to obtain a second score for each image frame in the target region, including:
[0062] Based on a preset single-frame analysis duration and the number of all I-frames in the target region of all videos in the image material, the sum of the first analysis durations for each I-frame is obtained. Here, the single-frame analysis duration is the time required to analyze one image frame. Based on the single-frame analysis duration and the number of all image frames in the target region of all videos in the image material, the sum of the second analysis durations for each image frame is obtained.
[0063] If the remaining analysis time is less than the sum of the first time, perform frame-by-frame analysis on the I-frames of the target region of each video in the multiple image materials to obtain the second score of the I-frames in the target region; the remaining analysis time is equal to the time used for video analysis minus the total time spent on performing overview analysis.
[0064] If the remaining analysis time is greater than or equal to the sum of the first time, or if the remaining analysis time is greater than the sum of the second time, perform frame-by-frame analysis on the image frames of the target region in each video from multiple image materials to obtain the second score of the image frames in the target region.
[0065] In this application, a frame-by-frame analysis strategy for the target region is determined based on the analysis time of all I-frames in the target region, the analysis time of all image frames in the target region, and the remaining analysis time. This allows for the analysis of a limited number of image frames within the target region within the remaining analysis time, thereby improving the analysis efficiency of highlight segments while obtaining relatively reliable analysis results under the limited performance of electronic devices.
[0066] In another possible implementation of the first aspect, the method further includes:
[0067] Based on the target regions' scores from highest to lowest, and according to the time taken to analyze all image frames within each target region, the remaining analysis time is allocated to each target region until all remaining analysis time is allocated. If any target region is not allocated analysis time, then that target region will not undergo frame-by-frame analysis.
[0068] In this application, if there are target regions that have not been allocated analysis time, meaning that the remaining analysis time is insufficient to analyze all target regions, then these target regions will not be analyzed frame-by-frame. The target regions not allocated analysis time are determined based on a scoring ranking; target regions with low scores have lower value, and their corresponding highlight fragments are of poor quality, so their analysis is abandoned, further reducing the time consumed by these target regions. Alternatively, more time can be allocated to target regions with high scores to obtain higher-value highlight fragments, ensuring the effectiveness of the highlight fragments.
[0069] In another possible implementation of the first aspect, the highlight segments of the corresponding video are determined based on the second score corresponding to each video, including:
[0070] For each video, obtain the second image frame based on the second score of the image frame of the target region;
[0071] The image frames within a second preset duration before and after the second image frame are determined as highlight segments of the video; the duration of the highlight segment is less than or equal to the recommended duration of the highlight segment.
[0072] In this application, the electronic device can determine a highlight segment that meets the recommended duration of the highlight segment based on the second image frame with the highest second score and the recommended duration of the highlight segment. The highlight segment determined based on the second image frame and the recommended duration of the highlight segment is relatively accurate.
[0073] In another possible implementation of the first aspect, before performing a first number of I-frame overview analyses on each video from the multiple image materials to obtain a first score, the method further includes:
[0074] If the total analysis time of multiple image materials exceeds the preset upper limit of the total analysis time, M videos are randomly selected from the multiple image materials. Here, the total analysis time represents the total time required to analyze multiple image materials, and the preset upper limit of the total analysis time represents the maximum recommended time required to analyze multiple image materials; the time required to analyze M videos is less than or equal to the total analysis time, and M is less than the number of videos in the image materials.
[0075] In this application, when the total analysis time of multiple image materials exceeds the upper limit of the recommended value, a portion of the video can be randomly selected from multiple image materials for analysis. Under the limited performance and time consumption of electronic devices, efficient analysis and processing of highlight segments can be achieved.
[0076] In another possible implementation of the first aspect, a first number of I-frame overview analyses are performed on each video in the multiple image materials to obtain a first score, including:
[0077] If the sum of the actual durations of all videos in multiple image materials is greater than a preset duration threshold, a first number of I-frame overview analyses are performed on each video in the multiple image materials to obtain a first score.
[0078] In this application, when the sum of the actual durations of videos containing multiple image materials exceeds a preset duration threshold, this solution can reduce the workload of electronic devices in performing image frame analysis, thereby reducing the time required for electronic devices to achieve one-click video creation, improving the efficiency of electronic devices in processing videos, and ultimately enhancing the user experience of the one-click video creation function.
[0079] In a second aspect, an electronic device is provided, comprising a memory, a display screen, and one or more processors; the memory, the display screen, and the processors are coupled together; the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the method as described in any one of the first aspects above.
[0080] Thirdly, a computer-readable storage medium is provided that stores instructions which, when executed on an electronic device, cause the electronic device to perform any of the methods described in the first aspect.
[0081] Fourthly, a computer program product containing instructions is provided, which, when run on an electronic device, enables the electronic device to perform the method described in any one of the first aspects above.
[0082] Fifthly, embodiments of this application provide a chip, the chip including a processor, the processor being configured to invoke a computer program in memory to execute the method as described in the first aspect.
[0083] It is understood that the beneficial effects of the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect can be referred to the beneficial effects of the first aspect and any of its possible design embodiments, which will not be repeated here. Attached Figure Description
[0084] Figure 1 This is a schematic diagram illustrating an application scenario of a video processing method provided in an embodiment of this application;
[0085] Figure 2 This is a schematic diagram illustrating an application scenario of another video processing method provided in an embodiment of this application.
[0086] Figure 3 This is a schematic diagram illustrating an application scenario of another video processing method provided in an embodiment of this application.
[0087] Figure 4 This is a schematic diagram illustrating an application scenario of another video processing method provided in an embodiment of this application.
[0088] Figure 5 This application provides a schematic diagram of the hardware structure of an electronic device.
[0089] Figure 6 This application provides a schematic diagram of the software structure of an electronic device.
[0090] Figure 7This application provides a flowchart illustrating a method for various modules in an electronic device to perform operations such as algorithm initialization and obtaining analysis parameters.
[0091] Figure 8 This application provides a flowchart illustrating image analysis in a video processing method.
[0092] Figure 9 This application provides a schematic diagram illustrating the technical concept of a video processing method.
[0093] Figure 10 This application provides a flowchart illustrating video analysis in a video processing method.
[0094] Figure 11 This application provides a schematic diagram of I-frame extraction as an embodiment;
[0095] Figure 12 Another schematic diagram of I-frame extraction is provided for the embodiments of this application;
[0096] Figure 13 This application provides a schematic diagram of allocating I-frames to multiple videos;
[0097] Figure 14 This application provides a schematic diagram of a first preset position and a second preset position in a 15-second video.
[0098] Figure 15 A schematic diagram of a first position provided for an embodiment of this application;
[0099] Figure 16 A schematic diagram of region 3 provided in an embodiment of this application;
[0100] Figure 17 This application provides a schematic diagram of a connected region of a video.
[0101] Figure 18 This application provides a schematic diagram of a target area as an embodiment;
[0102] Figure 19 This application provides a schematic diagram of another target area for its embodiments.
[0103] Figure 20 A schematic diagram of a highlight segment provided in an embodiment of this application;
[0104] Figure 21 This provides a schematic diagram of another highlight segment for an embodiment of this application;
[0105] Figure 22 This application provides a schematic flowchart of the post-processing of highlight segments in a video processing method according to an embodiment of the present application. Detailed Implementation
[0106] In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can indicate: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0107] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.
[0108] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0109] First, some of the terms or terms used in this application will be explained.
[0110] A keyframe, also known as an intra-coded picture (I-frame), is a frame that contains complete image information. During decoding, only this frame needs to be decoded to extract the complete picture. An I-frame is typically the first frame of each group of pictures (GOP), moderately compressed, and used as a reference point for random access. An I-frame (keyframe) is an image frame that includes complete image information.
[0111] Highlight clips, also known as montage clips, are single-frame images or video clips composed of multiple consecutive frames extracted from videos or images to record exciting moments. Highlight moments can include moments of great action such as a person smiling, winning a competition, jumping, an airplane landing, or a goal scored in a ball game.
[0112] One-click video editing refers to an electronic device's ability to automatically analyze highlight segments from one or more image clips selected by the user, and then combine these highlight segments into a pre-edited video. In other words, the electronic device can extract multiple highlight segments from one or more image clips and combine them into a single video. The image clips selected by the user can be pictures or videos; alternatively, the image clips can include both pictures and videos.
[0113] The process by which an electronic device selects highlight segments from a video or image may include: acquiring aesthetic scoring parameters for each frame, such as image color, image texture features, image quality, frame interpolation with preceding and following frames, and edge change rate values; and then scoring each frame aesthetically based on these parameters. The electronic device can then use the highest-scoring single frame or multiple consecutive frames as a single highlight segment.
[0114] Currently, the one-click video creation function supports editing of images and videos. In implementing this function, users may select a large number of images or a long video. In this case, the electronic device analyzes the selected images to extract highlight segments, which is time-consuming and results in low video processing efficiency. Furthermore, this extended processing time leads to excessively long waiting times for the final video output, negatively impacting the user experience.
[0115] In view of the above problems, this application provides a video processing method. Using this solution, in the process of achieving one-click video creation, the electronic device can first perform an overview analysis of a first number of keyframes in each video from multiple image materials to obtain a first score for each keyframe. Then, based on the first keyframe with the highest first score, and keyframes within a first preset time period before and after the first keyframe, the electronic device can determine the target region of the video. Afterwards, the electronic device can perform frame-by-frame analysis of the target region in the video to obtain a second score for each image frame. Based on the second image frame with the highest second score, and image frames within a second preset time period before and after the second image frame, the highlight segment of the video is determined.
[0116] Using this solution, the electronic device first performs an overview analysis of the video to locate the target area that needs to be processed frame by frame. When extracting highlight segments, the electronic device only performs frame-by-frame analysis on the target area, rather than analyzing the entire video frame by frame. This reduces the workload of image frame analysis, thereby reducing the time required for the electronic device to achieve one-click video creation, improving the efficiency of video processing, and ultimately enhancing the user experience of the one-click video creation function.
[0117] The video processing method provided in this application can be applied to electronic devices with image processing capabilities. It should be noted that the image materials used for one-click video creation in this application embodiment can include both images and videos. Users can use the one-click video creation function to generate a video set from highlight clips of multiple images, or from highlight clips of multiple videos, or from highlight clips of multiple images and videos.
[0118] The aforementioned electronic devices can also be referred to as terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Electronic devices can be mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technologies or device forms used in the electronic devices.
[0119] The following example uses a mobile phone as an electronic device, combined with... Figures 1-4 Taking mobile phones as an example, this paper introduces the application scenarios and interface implementation of one-click video creation on electronic devices.
[0120] In one application scenario, a user can pre-capture multiple images using their mobile phone, which are then stored in the phone's gallery. The phone can also pre-acquire images transmitted from other devices. For example... Figure 1 As shown in (a), a gallery icon is displayed on the phone's home screen. When a user wants to create an edited video clip based on multiple image assets using the phone, the user can tap the gallery icon on the home screen. In response to the user's tap on the icon, the phone displays the following... Figure 1 The gallery interface 101 is shown in (b) above. This gallery interface 101 includes a "One-Click Movies" option. The phone then responds to the user's request... Figure 1 Clicking the "One-Click Blockbuster" option shown in (b) will display the following: Figure 1 The gallery interface 102 is shown in (c). This gallery interface 102 can include multiple recently captured images. Users can select... Figure 1 One or more image assets from the plurality of image assets shown in (c) are used as candidate image assets for one-click image generation. For example, the mobile phone responds to the user's request. Figure 1 The selection operation of some image materials in (c) can display as follows: Figure 1The gallery interface 103 is shown in (d). This gallery interface 103 includes all the image assets in the gallery. The gallery interface 103 may also include video generation options, such as a checkmark option.
[0121] The phone responds to the user's Figure 1 Clicking the checkmark option (d) indicates that the process analyzes the five image materials selected by the user in the gallery interface 103, extracts highlight segments from each image material, and generates a video set based on the selected highlight segments. During this process, the phone can display... Figure 2 The image library interface 201 is shown. This image library interface 201 includes the analysis material progress so that users can intuitively view the analysis progress.
[0122] In one example, after the phone generates a video set, it can display... Figure 3 The gallery interface 301 is shown in (a) above. The generated video set can be displayed in the gallery interface 301. The mobile phone can automatically play this video set in the gallery interface 301. Furthermore, as... Figure 3 As shown in (a), the gallery interface 301 may also include a video export option 302 for supporting the export of generated video sets. In response to a user's click on the video export option 302, the phone can save the video set in the gallery, allowing the user to view it. In response to a user's click on the video export option 302, the phone may also display... Figure 3 The video export interface 303 is shown in (b) of the diagram.
[0123] In one example, as the user selects image materials, the phone can provide prompts to help the user determine the appropriate number of image materials to choose. For example... Figure 1 As shown in (d), the image library interface 103 displays the prompt message "6 or more image materials will produce better results", so that users can know at least how many image materials to select to generate a better video set.
[0124] In one example, as the user selects image assets, the phone can provide prompts to let the user know the maximum number of image assets they can select. For example... Figure 4 As shown, the gallery interface 401 displays the message "A maximum of 30 image materials can be selected", so that users can know how many image materials they can select.
[0125] In one example, after generating the video set, the phone can also display an interface showing other function options, allowing users to edit, add effects, analyze, and perform other operations on the generated video set based on these options. Figure 3As shown in 301, other function options may include, but are not limited to, templates, music, clips, sharing, etc.
[0126] The following example uses a mobile phone as an electronic device, combined with... Figure 5 The hardware structure of electronic devices will be introduced.
[0127] Figure 5 A schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of this application is shown. Figure 5 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a camera 193, a display screen 194, etc.
[0128] The processor 110 may include one or more processing units, such as a controller, application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). The controller may serve as the central nervous system and command center of the mobile phone 100. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The processor 110 may also include memory for storing instructions and data.
[0129] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0130] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0131] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering, such as rendering images... Figures 1-4 The diagram shows the user interface.
[0132] The display screen 194 is used to display the operation interface of the screen mirroring app, the mirrored image, and the mirrored video. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), or a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0133] In this embodiment, the display screen 194 can be used to display, for example, Figure 1 Gallery interfaces 101, 102, and 103; display screen 194 can be used to display, for example, Figure 2 The gallery interface 201; the display screen 194 can be used to display, for example, Figure 3 The image gallery interface 301 and video export interface 303 are shown in the image gallery interface 301; the display screen 194 can be used to display images such as... Figure 4 The image gallery interface in the middle is 401, etc.
[0134] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0135] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0136] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0137] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP performs Fourier transforms on the frequency energy.
[0138] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0139] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0140] The external memory interface 120 can be used to connect an external memory card, thereby expanding the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage.
[0141] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application (APP) required for a function (e.g., camera APP, gallery APP, and third-party video editing software, etc.). The data storage area may store data created during the use of the mobile phone 100 (e.g., photos or videos taken, screenshots, screen recordings, images downloaded from other devices, and video sets generated using the one-click video creation function, etc.). In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, and universal flash storage (UFS, etc.).
[0142] Electronic device 100 can implement audio functions through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0143] For example, after a video set is generated using the one-click video creation function, the audio module 170 decodes the audio signal of the video set, and then the speaker 170A, also known as a "loudspeaker," converts the audio electrical signal into a sound signal. In this way, the user can hear background music synchronized with the video in the highlight clips, as well as added video background music.
[0144] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0145] The following is combined Figure 6 The software architecture of electronic devices will be introduced.
[0146] Figure 6 This is a schematic diagram of the software structure of an electronic device provided in an embodiment of this application.
[0147] like Figure 6As shown, electronic devices can adopt a layered architecture, dividing the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the software layers of the software structure are divided from top to bottom as follows: application (APP) layer, media platform framework layer, application framework (FWK) layer, and hardware abstraction layer (HAL).
[0148] The APP layer, or application layer for short, can include a series of application packages, such as camera, gallery, third-party video editing software, calendar, map, and navigation. When these application packages are run, they can access the various service modules provided by the media platform framework layer and the application framework layer through the application programming interface (API) and execute corresponding intelligent business logic.
[0149] In some embodiments, the camera is used to capture photos, videos, slow-motion images, and panoramic images in response to user actions. After these images are captured by the camera, or after the user triggers a screenshot, or after the user triggers screen recording, or after the electronic device downloads images from other devices, the electronic device can save these images to a gallery, allowing the user to perform video editing operations on the images in the gallery, such as one-click video editing.
[0150] In this embodiment of the application, the image library is divided into three layers from top to bottom: business layer, application function layer, and basic function layer.
[0151] The business layer, also known as the video editing business layer, provides various services (or functions) such as automatic multi-camera recording and editing, AI music videos, one-click video creation, and highlight moments. These services are presented as controls in the gallery's user interface (UI). Users can trigger corresponding video processing actions in the gallery by interacting with these controls. For example, after a user selects image materials (including pictures and / or videos), in response to the user's click on the one-click video creation control in the gallery, the gallery can call the underlying module to automatically analyze and extract highlight segments from the pictures and / or videos using algorithms, and then combine the highlight segments into a pre-edited video set.
[0152] The application functionality layer includes an automatic editing framework. Various business functions in the business layer can call this framework to provide automatic editing services for images and videos. For example, the automatic editing framework may include functional modules such as segment selection, storyline organization, layout splicing, and special effects enhancement. Segment selection is used to call the light segment analysis interface and strategy monitoring interface in the high-media platform framework layer to extract highlight segments from images and / or videos. Storyline organization is used to sequentially splice multiple images and / or videos in the form of a storyline based on their content. Layout splicing is used to adjust the interface layout of images and / or videos. Special effects enhancement is used to adjust the enhancement effects of images and / or videos, such as adjusting screen brightness and beautifying facial features.
[0153] The basic functionality layer is used to perform basic processing on the edited image and / or video clips after the automatic editing framework has edited multiple images and / or videos. For example, the basic functionality layer may include basic functional modules such as video splicing, compositing and saving, video effect rendering, and audio effect processing. Specifically, video splicing is used to splice multiple extracted highlight clips (where highlight clips include images and / or videos). Compositing and saving is used to store the spliced video set. Video effect rendering is used to add video effects to the spliced video set, such as adding style filters and themes. Audio effect processing is used to add sound effects to the spliced video set, such as adding background music.
[0154] The media middleware framework layer is a software layer positioned between the application layer and the application framework. This layer can include an analysis performance query interface, a highlight clip analysis interface, a policy monitoring interface, a pipeline interface, a theme summary interface, and an initialization interface. The analysis performance query interface calculates the total duration of all videos based on the video analysis speed. The highlight clip analysis interface calls the policy monitoring interface to extract highlight clips. The theme summary interface calls underlying algorithms to analyze the content of highlight clips to determine the corresponding theme. The initialization interface initializes the highlight clip algorithm, face detection algorithm, video acceleration algorithm, and image super-resolution algorithm in the HAL layer.
[0155] The policy monitoring interface is used to calculate the initial number of keyframes (I-frames) and their positions within the video (I-frame positions). The channel interface, based on the file descriptors and I-frame positions for each video sent by the policy monitoring interface, reduces the resolution of the video files. It then forwards the data addresses of the reduced-resolution video files to the hardware abstraction layer via the application framework layer. Finally, the analysis results of the I-frames returned by the hardware abstraction layer are reported back to the policy monitoring interface. These analysis results may include aesthetic scores for the image frames.
[0156] The policy monitoring interface is also used to determine the second position of an I-frame based on the analysis results of the I-frame at the first position. When the number of I-frames sent for analysis in the video meets the first requirement for the video, the target region of the video is determined based on the analysis results of each I-frame. The target region is the region in the video corresponding to the I-frame with the highest score. The channel interface is also used to reduce the resolution of the video file based on the file descriptor of each video and the frames (I-frames or image frames) in the target region sent by the policy monitoring interface. The data address of the reduced-resolution video file is forwarded to the hardware abstraction layer through the application framework layer. Then, the analysis results of the frames (I-frames or image frames) in the target region returned by the hardware abstraction layer are reported to the policy monitoring interface. In this embodiment, the image frame represents a regular image frame that is different from the keyframe.
[0157] The policy monitoring interface is also used to determine the highlight segments of each video based on the analysis results of frames (I-frames or image frames) in the target area.
[0158] The theme summary interface is used to call the underlying algorithms to analyze the content of highlight segments and determine the theme corresponding to the content of the highlight segments. The initialization interface is used to initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, and image super-resolution algorithm of the HAL layer.
[0159] It should be noted that this application uses the one-click image creation function provided by an image library as an example for illustration, and does not limit the embodiments of this application. In actual implementation, third-party video editing software can use the video processing method provided in the embodiments of this application to combine multiple images and videos selected by the user into a single video set with one click.
[0160] The FWK layer, or framework layer for short, supports the operation of various modules within the media middleware framework layer. For example, the framework layer may include a one-click video creation interface, a parameter management interface, an image data transmission interface, a theme analysis interface, and a performance analysis interface.
[0161] The Hardware Abstraction Layer (HAL) is a wrapper around Linux kernel drivers, providing interfaces to higher-level systems. It hides the hardware interface details of specific platforms, providing the operating system with a virtual hardware platform that is hardware-independent and portable across multiple platforms. For example, the HAL can include chip analysis speed interfaces, highlight fragment algorithms, face detection algorithms, video acceleration algorithms, and image super-resolution algorithms. The highlight fragment algorithm is an image processing algorithm provided by the image signal processor. This algorithm performs an aesthetic score on each frame of an image based on factors such as image color, texture features, image quality, frame interpolation with preceding and following frames, and edge change rate. The aesthetic score serves as the basis for evaluating whether a frame contains a highlight fragment.
[0162] In this embodiment, the highlight segment algorithm can score each frame (I-frame or ordinary image frame) in the video. For example, if the score of a frame is greater than a preset scoring threshold, then that frame and the frames within its adjacent preset duration can be used as highlight segments of the source video. For example, the preset scoring threshold can be 90. If the score of a frame in the video is greater than 90, then that frame and the frames within its adjacent preset duration can be used as a highlight segment of the video. Alternatively, if the video includes scores for N frames, the frame with the highest score and the frames before and after that frame within a preset duration can be used as highlight segments of the video.
[0163] In some other feasible embodiments, assuming the frame score ranges from 0 to 100, the scores can be divided into different levels based on the different score ranges. For example, a first threshold of 80, where a frame score greater than or equal to the first threshold (score range 80-100), indicates a "high" score. A second threshold of 50, where a frame score greater than or equal to the second threshold (score range 50-79), indicates a "relatively high" score. A third threshold of 20, where a frame score greater than or equal to the third threshold (score range 20-49), indicates a "medium" score. A frame score less than the third threshold (score range 0-19) indicates a "low" score. For example, in one implementation, frames with scores greater than or equal to the first / second / third thresholds can be selected as optional frames for extracting highlight segments from the video. In some other feasible methods, the highest-scoring frame can be selected as the highlight frame for extracting the video's highlights.
[0164] It should be noted that, Figure 6 The layers and components within each layer of the illustrated software architecture do not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more layers than illustrated, such as a system library (FWK LIB) layer and a kernel layer. Each layer may include more or fewer components than illustrated. Furthermore, the aforementioned functional modules may be combined into a single functional module, and the layers may be combined into a single layer; for example, highlight segment analysis may include policy monitoring, and a media middleware framework layer may be located within an application framework layer.
[0165] It is understood that, in order to implement the video processing method in the embodiments of this application, the electronic device includes hardware and / or software modules that perform various functions. Based on the algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments.
[0166] The video processing method provided in this application embodiment includes a strategy monitoring module in the media platform framework layer. This module can determine the first number of keyframes for overview analysis of each video based on the analysis parameters of the image material, and dynamically determine the position of each keyframe. The strategy monitoring module can send the position of each keyframe in each video to the algorithm module in the HAL layer for analysis. From the HAL layer algorithm, the first score of each keyframe is obtained, and the keyframe with the highest first score and its corresponding region are further identified as the target region. Frame-by-frame analysis is performed on the target region to obtain the second score of each frame in the target region. The strategy monitoring module can send each image frame in the target region to the algorithm module in the HAL layer for analysis, obtaining the score result of each frame in the target region from the HAL layer algorithm. This determines the image frame with the highest second score and its corresponding region as a highlight segment. By combining overview analysis and frame-by-frame analysis, instead of performing frame-by-frame analysis on the entire video, the workload of image frame analysis on electronic devices can be reduced, thereby reducing the time required for electronic devices to achieve one-click video creation, improving the efficiency of video processing on electronic devices, and ultimately enhancing the user experience of the one-click video creation function.
[0167] The following section takes the execution entity of the audio and video processing method as an example. Figure 6 Using the modules shown in the software structure diagram as examples, the video processing method provided in this application embodiment will be illustrated by way of example.
[0168] Figure 7 This document provides a flowchart illustrating a method for various modules within an electronic device to perform algorithm initialization and acquire analysis parameters before the policy monitoring module determines the initial number of keyframes for overview analysis of each video. This method can be applied to applications such as... Figures 1-4 In the one-click image generation scenario shown, for example... Figure 7 As shown, taking image materials including pictures and videos as an example, the method may include the following steps S01-S16.
[0169] S01, the business layer receives user input to enable the one-click video creation function.
[0170] In this embodiment, the "business layer" refers to the one-click video creation module within the business layer. That is, the one-click video creation module in the business layer receives user input to enable the one-click video creation function. For example, this operation can specifically be as follows: Figure 1 The click operation of the "One-Click Blockbuster" card is shown in (b) in the image.
[0171] S02, the business layer loads and displays candidate images and candidate videos.
[0172] S03, the business layer receives the user's operation of selecting multiple pictures and videos, and receives the user's confirmation to execute the one-click video creation function.
[0173] For example, the user's selection of multiple images and videos can be done as follows: Figure 1 The click operation on photos and videos shown in (c) allows the user to confirm and execute the one-click video creation function, which can be done as follows: Figure 1 (d) shows the click action for the checkmark option.
[0174] S04, the business layer calls the initialization interface of the media middle platform framework layer through the application function layer to initialize the relevant algorithms of the HAL layer.
[0175] In this embodiment, "related algorithms" refers to the algorithms for the functions that the business layer needs to implement. Here, the business function is a one-click video creation function, so the related algorithms are those involved in the one-click video creation function. For example, related algorithms include highlight segment algorithms, face detection algorithms, video acceleration algorithms, and image super-resolution algorithms, etc.
[0176] S05, the initialization interface of the media middle platform framework layer sends initialization parameters to the algorithm module of the HAL layer through the channel interface of the media middle platform framework layer and the service interface of the FWK layer in sequence.
[0177] In the HAL layer, one algorithm corresponds to one algorithm interface. The FWK layer has multiple service interfaces, and one service interface in the FWK layer corresponds to one algorithm interface in the HAL layer. The various service interfaces in the FWK layer play a role in data pass-through between the algorithm interfaces in the HAL layer and the channel interfaces in the media middleware framework layer.
[0178] Different algorithms involve different initialization parameters, so the algorithm initialization parameters sent to the algorithm interface through each service interface may be different.
[0179] S06, the algorithm module of the HAL layer is initialized according to the initialization parameters.
[0180] S07, the algorithm module of the HAL layer returns an initialization success message to the channel interface of the media middleware framework layer through the service interface of the FWK layer.
[0181] S08, the channel interface of the media middleware framework layer calls the analysis speed interface of the HAL layer through the service interface of the FWK layer to obtain the chip analysis speed.
[0182] Among them, the service interface of the FWK layer can be a performance analysis interface.
[0183] S09, the HAL layer's analysis speed interface returns the chip analysis speed to the media middleware framework layer's channel interface through the FWK layer's service interface.
[0184] S10, the channel interface of the media middleware framework layer returns the chip analysis speed to the initialization interface of the media middleware framework layer.
[0185] Chip analysis speed can characterize the number of image frames analyzed by the image signal processor per unit time; alternatively, it can characterize the single-frame analysis duration of the image signal processor. Single-frame analysis duration refers to the duration of analyzing one image frame in a video. Chip analysis speed can also characterize the duration of analyzing one image.
[0186] Based on this, the single-frame analysis time and the processing time for an image by the image signal processor can be calculated according to the chip analysis speed. The single-frame analysis time and the processing time for an image may differ. For example, the processing time for an image may be 400ms, and the single-frame analysis time may be 200ms.
[0187] It should be understood that because different image signal processors have different performance characteristics, the chip analysis speed corresponding to different image signal processors may vary. For a commercially available electronic device, the image signal processor is fixed, and therefore the chip analysis speed corresponding to that image signal processor is also fixed.
[0188] In some embodiments, the channel interface of the media middleware framework layer can also return other relevant performance parameters of various algorithms to the initialization interface of the media middleware framework layer.
[0189] S11, the initialization interface of the media middleware framework layer returns an initialization success message to the business layer through the application function layer.
[0190] The initialization success message can carry performance parameters of various algorithms, such as chip analysis speed.
[0191] After obtaining the chip analysis speed, for each video in the image materials selected by the user, the following steps S12-S15 can be executed to obtain the analysis parameters for the video and all image materials.
[0192] Taking Video 1 as an example, Video 1 is one of multiple videos in the image material.
[0193] S12, the business layer sends a query message to the analysis performance query interface of the media platform framework layer through the application function layer.
[0194] The query message includes the file descriptor fd1 of the user-selected video 1 and the chip analysis speed. The file descriptor serves as a unique identifier for the video.
[0195] S13, the media middle platform framework layer's analysis performance query interface obtains the estimated analysis duration of video 1 based on video 1's file descriptor fd1.
[0196] The estimated analysis duration can be the time required for the image signal processor to analyze a video. Analyzing a video can refer to analyzing a specified number of keyframes within the video. For example, the specified number can be 1, or it can be greater than 1 but less than or equal to the number of keyframes contained in the video.
[0197] For example, when the specified quantity is 1, that is, when the estimated analysis time represents the time required for the image signal processor to analyze one keyframe in a video, the estimated analysis time can also represent the minimum analysis time corresponding to the keyframe. For example, if the single-frame analysis time of the image signal processor is 200ms, then the estimated analysis time for each video (including video 1) is 200ms.
[0198] In some embodiments, the specified quantity can also be the number of keyframes contained in the video. In this case, the estimated analysis duration can be the duration corresponding to the analysis of all keyframes in the video. For example, if the single-frame analysis duration of the image signal processor is 200ms, and video 1 includes 10 keyframes, then the estimated analysis duration of video 1 is 10 * 200ms = 2000ms. For example, if video 2 includes 12 keyframes, then the estimated analysis duration of video 2 is 12 * 200ms = 2400ms.
[0199] In some embodiments, the specified quantity can also be 'a', where 'a' is greater than 1 and less than the number of image frames contained in the video, and 'a' is a natural number. Therefore, the estimated analysis duration can be the duration corresponding to 'a' keyframes in the video. For example, if the single-frame analysis duration of the image signal processor is 200ms, and video 1 includes 10 keyframes with N = 5, then the estimated analysis duration of video 1 is 5 * 200ms = 1000ms.
[0200] S14, the analysis performance query interface of the media middle platform framework layer returns the estimated analysis duration of video 1 to the business layer through the application function layer.
[0201] After the one-click video generation module in the business layer obtains the estimated analysis duration of video 1, it can return to continue executing S12-S15 to obtain the estimated analysis duration of the next video among multiple image materials, until the estimated analysis duration of all videos among multiple image materials is obtained.
[0202] S15, the business layer obtains analysis parameters based on the estimated analysis time of all image materials.
[0203] The analysis parameters may include total analysis duration, suggested upper limit of total analysis duration, maximum duration of highlight clips, minimum duration of highlight clips, total suggested duration of highlight clips, suggested duration of highlight clips, whether to force each video to output a highlight clip, whether to enable audio analysis, selected highlight clips, etc.
[0204] The option to force each video to output a highlight clip is set to "yes" by default, meaning that in this embodiment, each video needs to output a highlight clip.
[0205] The total analysis time represents the total time required to complete the analysis of all image materials selected by the user (including all pictures and all videos selected by the user). The total analysis time includes the sum of the estimated analysis time of all pictures and the sum of the estimated analysis time of all videos.
[0206] For images in the image source material, the processing time for an image can be determined directly based on the chip speed of the image processor. Therefore, the sum of the estimated analysis times for all images can be determined directly based on the number of images in the image source material.
[0207] For videos in the image material, the estimated analysis duration of each video can be obtained according to S12-15. The estimated analysis durations of each video are summed up to obtain the sum of the estimated analysis durations of all videos in the image material.
[0208] The total analysis time of the image materials can be obtained by summing the estimated analysis times of all images and all videos.
[0209] For example, suppose the user selects 10 images, video 1, and video 2. The estimated analysis time for one image is 400ms, the estimated analysis time for video 1 is 200ms, and the estimated analysis time for video 2 is 300ms. Then the total analysis time is 10 * 400ms + 200ms + 300ms = 4500ms.
[0210] The suggested upper limit for total analysis time represents the maximum recommended time required to complete the analysis of all the images and videos mentioned above. This upper limit can be determined based on the estimated analysis time for each video and image. Generally, the suggested upper limit for total analysis time is greater than the total analysis time. For example, assuming a minimum of one image frame is analyzed per video, the total analysis time would be 4400ms. However, considering that each video may require analyzing multiple image frames, the suggested upper limit for total analysis time can be significantly larger than the total analysis time; for instance, it could be a preset value of 10000ms.
[0211] In scenarios where analysis parameters are input abnormally, the recommended upper limit for the total analysis time may also be set to be less than the total analysis time. When the recommended upper limit for the total analysis time is less than the total analysis time, that is, when the recommended upper limit for the total analysis time is insufficient to analyze all images and all videos (one image frame), a subset of the selected image materials can be selected for analysis. This part is implemented by the strategy monitoring module of the media platform framework layer, which will be described in detail in the following embodiments and will not be elaborated here.
[0212] The maximum duration of a highlight clip indicates the maximum allowed duration of a highlight clip in a video. The minimum duration of a highlight clip indicates the minimum required duration of a highlight clip in a video. Both the maximum and minimum durations of a highlight clip can be preset values. For example, the maximum duration of a highlight clip can be 3000ms, and the minimum duration of a highlight clip can be 1000ms.
[0213] The total suggested duration of highlight clips represents the suggested value of the sum of the suggested durations of all highlight clips across multiple video clips. The total suggested duration of highlight clips can be determined based on the number of videos, the maximum duration of a highlight clip, and the minimum duration of a highlight clip. For example, the maximum duration of a highlight clip can be 3000ms, the minimum duration of a highlight clip can be 1000ms, and when there are 5 videos, the total suggested duration of highlight clips can range from 5000ms to 15000ms; for instance, the total suggested duration of highlight clips could be 8000ms.
[0214] The suggested duration for a highlight clip refers to a recommended value for the duration of a highlight clip in a video. Highlight clips of this suggested duration can effectively showcase the highlight effect. The suggested duration for a highlight clip can be determined based on the maximum and minimum duration of the highlight clip. For example, the maximum duration of a highlight clip can be 3000ms, the minimum duration can be 1000ms, and the suggested duration can range from 1000ms to 3000ms. For instance, the suggested duration for a highlight clip could be 2000ms.
[0215] It is important to understand that the maximum duration of highlight clips, the minimum duration of highlight clips, the total recommended duration of highlight clips, and the recommended duration of highlight clips can all be set according to the actual situation.
[0216] In some embodiments, the analysis parameters may also include the actual duration of each video.
[0217] If the actual duration of all videos in multiple image materials is greater than the sum of their durations, which is greater than a preset duration threshold, the video analysis method provided by S16-S46 of this solution can be executed to reduce the number of video frames analyzed in the image materials, thereby reducing the time required for electronic devices to achieve one-click video creation, improving the efficiency of electronic devices in processing video, and thus enhancing the user experience of the one-click video creation function.
[0218] S16, the business layer sends the file descriptors (fd) and analysis parameters of all materials to be analyzed to the image highlight segment analysis interface of the media platform framework layer through the application function layer.
[0219] After receiving the analysis parameters, the strategy monitoring module of the media middle platform framework layer can determine the material analysis strategy based on the suggested upper limit of the total analysis time and the total analysis time in the analysis parameters. The material analysis strategy refers to analyzing all image materials selected by the user, or selecting a portion of the selected image materials for analysis. In the following embodiments, the image materials include all videos and all images selected by the user.
[0220] After obtaining the total analysis time of the image materials, the policy monitoring module can determine the number of images and videos that can be analyzed based on the total analysis time and the suggested upper limit value of the total analysis time in the analysis parameters. Figure 7 After S16, refer to Figure 8 The given methodology for determining analysis strategies for image materials includes:
[0221] Specifically, in one implementation, it is executed after S16:
[0222] S17, the strategy monitoring module of the media middle platform framework layer determines the analysis strategy based on the total analysis time and the upper limit of the total analysis time suggested in the analysis parameters.
[0223] Specifically, if the total analysis time for image materials is less than or equal to the recommended upper limit for total analysis time, the material analysis strategy can be to analyze all images and all videos. For example, if the total analysis time is 4400ms and the recommended upper limit for total analysis time is 10000ms, the material analysis strategy can be to analyze all images and all videos in the image materials selected by the user.
[0224] If the total analysis time of the image materials exceeds the recommended upper limit, the material analysis strategy can be a random sampling strategy. For example, if the total analysis time is 4400ms and the recommended upper limit is 3000ms, then the material analysis strategy can be a random sampling strategy. Here, the random sampling strategy refers to randomly selecting a portion of the image materials from the user-selected image materials for analysis.
[0225] Assuming the analytical value of a single image is greater than that of a single I-frame from a video, the random sampling strategy could be: For every N images selected, M videos can be selected, until the total analysis time of the selected image materials reaches the suggested upper limit or all images in the selected image materials have been extracted. Here, M < N; for example, N can be a natural number such as 3, 4, or 5, and M can be a natural number less than N such as 1, 2, or 3. The specific values of M and N can be determined based on the number of image materials selected by the user. For example, it could be that for every 4 images selected, 1 video can be selected, until the total analysis time reaches 3000ms; or, all images in the selected image materials have been extracted.
[0226] For example, the total analysis time for 10 images and 2 videos is 4400ms, which exceeds the recommended maximum analysis time of 3000ms. Following a random sampling strategy of allowing 1 video to be selected for every 4 images, selecting 4 images and 1 video results in a total analysis time of 1800ms, which is still below the recommended maximum of 3000ms. Continuing to select image materials, when the third image is selected in this round, the total analysis time reaches 3000ms, at which point image material selection stops. Therefore, the selected image materials are 7 images and 1 video. The remaining 3 images are not analyzed.
[0227] In another embodiment, assume that the analytical value of an image is less than the analytical value of one I-frame of a video. Then, the random sampling strategy could be to allow the selection of Q images for every P videos selected, until the total analysis time of the selected image materials reaches the suggested upper limit of the analysis time or all videos in the selected image materials have been selected. Here, Q < P. For example, P can be a natural number such as 3, 4, or 5, and Q can be a natural number less than P such as 1, 2, or 3. The specific values of P and Q can be determined based on the number of image materials selected by the user. For example, it could be that one image is allowed for every three videos selected, until the total analysis time reaches 3000ms; or, until all videos in the selected image materials have been selected.
[0228] For example, the total analysis time for 10 videos and 3 images is 3200ms, which exceeds the recommended maximum analysis time of 3000ms. Following a random sampling strategy of allowing one image to be selected for every three videos, selecting 3 videos and 1 image results in a total analysis time of 1000ms, which is still below the recommended maximum. Continuing to select image materials, after three rounds of selecting 3 videos and 1 image, a total of 9 videos and 3 images have been selected, with a total analysis time of 3000ms. At this point, image material selection stops. Therefore, the selected image materials are 9 videos and 3 images. The remaining 1 video is not analyzed.
[0229] In some embodiments, the electronic device determines analyzable image materials from the user-selected materials according to the random sampling strategy described above (e.g., a mobile phone). This can be achieved by randomly selecting images and videos from the user-selected image materials in the order they were chosen. Alternatively, random selection can also be used. For example, a random number less than the number of materials can be generated using Java's Random class, and the video or image corresponding to that random number can be selected.
[0230] In some embodiments, the image material includes only pictures, and the total analysis time is the sum of the estimated analysis times of all pictures. When the total analysis time exceeds the suggested upper limit of the analysis time, the number of analyzable pictures is calculated based on the suggested upper limit of the analysis time and the estimated analysis time for analyzing one picture, and a corresponding number of pictures are randomly selected from all pictures for analysis.
[0231] In some embodiments, the image material includes only videos, and the total analysis time is the sum of the estimated analysis times of all videos. When the total analysis time exceeds a suggested upper limit for analysis time, the number of analyzable videos is calculated based on the suggested upper limit, and a corresponding number of videos are randomly selected from all videos for analysis. The time required to reach the number of analyzable videos is less than or equal to the total analysis time.
[0232] After selecting the images and / or videos that can be analyzed within the suggested maximum analysis time from the image materials, the remaining images and videos in the user's selected image materials will not be analyzed.
[0233] The image material includes pictures. After selecting the images that can be analyzed within the recommended maximum analysis time, the electronic device executes the process. Figure 8 S18-S22, as shown, perform highlight segment analysis on the analyzable image.
[0234] The following example illustrates the process of performing highlight segment analysis on image 1 using steps S18-S22. Image 1 is one of the images that can be analyzed within the recommended upper limit of the analysis time.
[0235] S18, the strategy monitoring module of the media middle platform framework layer sends an instruction message to the channel interface of the media middle platform framework layer. The instruction message includes the file descriptor of Figure 1.
[0236] S19, the channel interface of the media middle platform framework layer performs decoding, resolution reduction and format conversion on image 1 according to the file descriptor of image 1, and stores the processed image 1.
[0237] S20, the channel interface of the media middleware framework layer sends the frame data address of image 1 to the highlight fragment algorithm interface of the HAL layer through the service interface of the FWK layer.
[0238] Among them, the service interface can be a one-click production interface.
[0239] S21, the HAL layer's specular fragment algorithm interface obtains image 1 based on the frame data address of image 1, and then analyzes image 1 based on the preset specular fragment algorithm to obtain the analysis results.
[0240] For example, the highlight fragment algorithm interface can perform an aesthetic score on image 1 based on its image color, image texture features, image quality, and edge change rate value, and obtain the analysis result. The aesthetic score can be represented in the analysis result.
[0241] In some embodiments, the analysis results may include a score value, or the analysis results may include both a score value and a scoring result.
[0242] S22, the HAL layer's highlight fragment algorithm interface sequentially returns the analysis results of image 1 to the media platform framework layer's strategy monitoring module through the FWK layer's service interface and the media platform framework layer's channel interface.
[0243] After the strategy monitoring module of the media middle platform framework layer obtains the analysis results of image 1, if there are other images, such as image 2, the electronic device can continue to execute S18-S22 above to obtain the analysis results of the other images.
[0244] Among them, the strategy monitoring module of the media middle platform framework layer can determine the highlight segments of multiple image materials based on the analysis results of each image. The images with scores greater than or equal to the first threshold / second threshold are selected as highlight segments.
[0245] After obtaining the analysis results of all images, if the image material includes video, the electronic device can use steps S23-S36 below to obtain the analysis results of each video. If the image material does not include video, the electronic device outputs a target video set composed of images of highlight segments.
[0246] refer to Figure 9This illustration shows a technical approach for analyzing highlight segments in videos, as provided in this application. During the analysis of each video, the electronic device first acquires a first number of I-frames (keyframes) from the video for analysis, obtaining a first score for each I-frame. The first number of I-frames can be randomly acquired from the video, for example, by randomly acquiring I-frames from different segments based on the video's shot points. Alternatively, the positions of the first number of I-frames can be dynamically determined based on the first score. The target region of each video is determined based on the first I-frame with the highest first score. Then, the target region is analyzed frame-by-frame (frame-by-frame or image-by-image frame) to obtain a second score for each image frame in the target region. The highlight segment of the video is determined based on the second image frame with the highest second score. In this solution, instead of analyzing the entire video frame-by-frame, the workload of the electronic device in frame analysis is reduced, thereby reducing the time required for the electronic device to achieve one-click video creation, improving the efficiency of video processing, and ultimately enhancing the user experience of the one-click video creation function.
[0247] If the image material includes multiple images and videos, since the computational load on images is smaller and takes less time, electronic devices typically perform highlight segment analysis on each image individually first. After completing the highlight segment analysis for all images, they then perform highlight segment analysis on each video individually. In other words, after completing the analysis for all images, the electronic device will proceed with the next step. Figure 8 As shown in S18-S22, continue to execute as follows: Figure 10 S23-S37 are shown.
[0248] If the image material only includes multiple videos, then after executing S17, the electronic device will execute as follows: Figure 10 As shown in S23-S37, there is no need to execute as follows: Figure 8 S19-S22 are shown.
[0249] This embodiment illustrates the case where the image material includes both pictures and videos. The videos mentioned below refer to those that can be analyzed within the recommended upper limit of the total analysis time.
[0250] S23, the strategy monitoring module of the media middle platform framework layer calculates the first number of I-frames for each video.
[0251] The first quantity refers to the number of I-frames that can be analyzed in each video within the recommended upper limit of the total analysis time.
[0252] In some embodiments, the policy monitoring module can determine the first number of I-frames for each video based on the actual duration of each video in multiple image materials, the number of videos in multiple image materials, and the recommended upper limit value of the total analysis duration.
[0253] For example, based on the suggested upper limit value T1 of the total analysis time and the analysis time of a single frame t, determine the number of I-frames T1 / t that can be analyzed within the suggested upper limit value of the total analysis time. Based on the actual duration of each video and the number of videos, allocate the number of analyzable I-frames T1 / t to each video, and obtain the first count for each video.
[0254] In one example, the number of analyzable I-frames, T1 / t, is allocated to each video. This can be done by sorting the videos from longest to shortest according to their actual duration. For the top 25% of videos, T1 / t * 50% of the I-frames are allocated on average. For videos ranked 25%-75%, T1 / t * 40% of the I-frames are allocated on average. For the bottom 25% of videos, T1 / t * 10% of the I-frames are allocated on average.
[0255] In another example, the number of analyzable I-frames T1 / t is allocated to each video, which can be done by allocating a corresponding number to each video according to the proportion of the actual duration of the video.
[0256] Alternatively, in another example, the step of the policy monitoring module determining the first number of image frames for each video may include:
[0257] S231, the strategy monitoring module calculates the basic number and maximum number of I-frames for each video.
[0258] The minimum number of I-frames can be understood as the minimum number of I-frames required to ensure effective video analysis when the analysis time and the analytical capabilities of the electronic device are limited. The maximum number of I-frames can be understood as the maximum number of I-frames that the analysis time allows for. Generally, the maximum number is greater than the minimum number of I-frames.
[0259] In some embodiments, the policy monitoring module can determine the basic and maximum number of I-frames for each video based on the actual duration of each video and a preset correspondence. The preset correspondence represents the maximum and basic number of I-frames corresponding to different threshold ranges in video duration.
[0260] The base number and maximum number of I-frames differ for videos of different durations. For example, electronic devices (such as mobile phones) can pre-store a preset correspondence between video duration and the number of I-frames. In this embodiment, the policy monitoring module can determine the base number and maximum number of I-frames for each video based on this preset correspondence.
[0261] For example, the preset correspondence includes the following (1)-(4):
[0262] (1) If the video duration P is less than the first threshold Q1, the number of I frames in the video is 1.
[0263] (2) The video duration P is equal to the first threshold Q1, and the initial number of I-frames of the video is m.
[0264] (3) If the video duration P is greater than the first threshold Q1 and less than or equal to the second threshold Q2, the number of video I-frames is increased from the initial number. The number of I-frames is updated to... That is, P divided by s, rounded down, and then m. This indicates that from the start of the video to Q1, the number of I-frames in the video is m, and from the start of the video to Q2, there is one I-frame per second.
[0265] (4) The video duration P is greater than the second threshold Q2 and less than the third threshold Q3, and the number of video I-frames is... That is, divide Q2 by s and take the integer part, divide P-Q2 by k and take the integer part, and then add m. This indicates that from the start of the video to Q1, the number of I-frames in the video is m; from the start of the video to Q2, there is one I-frame every s seconds; and from Q2 to Q3, there is one I-frame every k seconds. This is the floor symbol.
[0266] Among them, the first threshold < the second threshold < the third threshold.
[0267] It should be noted that if the video duration is very long, multiple duration thresholds can be set, such as a fourth threshold, a fifth threshold, etc. When determining the base and maximum number of thresholds, the first, second, and third thresholds can be set according to the actual video duration; the interval in seconds can be determined based on the analysis capabilities of the electronic device. The above are merely illustrative examples, and no specific numerical limits are imposed on the parameters.
[0268] In some embodiments, the maximum number of I-frames is greater than the base number, and the sampling density for determining the maximum number of I-frames is greater than the sampling density for determining the base number of I-frames. To obtain a greater number of I-frames, generally, a first threshold for determining the maximum number is less than or equal to a first threshold for determining the base number; or, a preset second threshold for determining the maximum number is equal to or greater than a second threshold for determining the base number; or, a third threshold for determining the maximum number is greater than or equal to a third threshold for determining the base number. In some embodiments, the interval in seconds for determining the maximum number is less than or equal to the interval in seconds for determining the base number.
[0269] The maximum number of I-frames results in a higher density and greater quantity in the video; given sufficient analysis time, more I-frames can be analyzed, potentially increasing the maximum number of I-frames. The above are merely illustrative examples, and no specific numerical limits are imposed on the parameters.
[0270] The following examples illustrate the process of determining the base number and maximum number of I-frames in a video.
[0271] For example, the strategy monitoring module determines the basic number of I-frames in the video. The first threshold can be 3 seconds, s seconds can be 5 seconds, the second threshold can be 30 seconds, k seconds can be 10 seconds, and the third threshold can be 90 seconds.
[0272] For example, taking video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example, see the example below. Figure 11 The diagram shown is a schematic of the I-frame extraction.
[0273] The policy monitoring module receives video 1 and determines at least one I-frame. At this point, the number of I-frames in video 1 is 1. The duration of video 1 (91 seconds) is greater than the first threshold (3 seconds), so the number of I-frames in video 1 is updated to 2 frames.
[0274] Starting from second 0 of video 1's duration, and continuing until second 30 of the video's duration, the number of I-frames increases by one every 5 seconds, such as... Figure 11 As shown, at a video duration of 30 seconds, the basic number of I-frames in video 1 is 8 frames. Starting from the 30th second of the video duration, and continuing until the 90th second of the video duration, the number of I-frames increases by one every 10 seconds, such as... Figure 11 As shown, at 90 seconds into the video, the base number of frames in Video 1 is updated to 14. After the 90th second, the remaining duration of Video 1 is 3 seconds, which does not meet the requirement of extracting one I-frame every 10 seconds after the 90th second. Therefore, the base number of frames is not increased. Thus, the base number of frames in Video 1, with a duration of 93 seconds, is 14.
[0275] In some embodiments, after calculating the base number of each video, if the total base number of all videos does not meet the preset minimum frame rate, the base number of all videos needs to be adjusted. This adjustment can be achieved by increasing the frame rate of videos with a base number less than the preset value, or by increasing the base number of videos with longer durations.
[0276] For example, in order to ensure that the video theme is obtained more accurately, the minimum frame rate can be preset to 5 frames.
[0277] In some embodiments, if there is only one video, for example, if the base number of frames for video 1 is less than 5, the base number of frames for video 1 is directly adjusted to 5. If there are multiple videos, since each video is required to acquire at least one I-frame, the preset minimum frame count should be greater than the number of videos. For example, if there are 3 videos, the preset minimum frame count can be 5 frames. For example, if the base number of frames for video 1 is 1 frame, for video 2 it is 2 frames, and for video 3 it is 1 frame, the total base number of frames for all videos is 4 frames, which is less than the preset minimum frame count of 5 frames. The base number of videos with a base number less than 2 frames can be adjusted to 2 frames. In this case, the base number of frames for video 1 is 2 frames, for video 2 is 2 frames, and for video 3 it is 2 frames, and the total base number of frames for all videos is 6 frames, which is greater than the preset minimum frame count of 5 frames. If the number of frames for a video is 2, and the base number of frames for video 1 and video 2 is 1 frame... If adjusting the base frame count of videos with less than 2 frames to 2 frames, the total base frame count of all videos, at 4 frames, still doesn't meet the preset minimum frame count of 5. In this case, the base frame count of the longest video can be adjusted to 3 frames, ensuring the total base frame count of all videos meets the minimum frame count of 5. For example, if video 1 is 2 seconds long and video 2 is 1 second long, then video 1's base frame count can be adjusted to 3 frames, and video 2's base frame count to 2 frames. Now, the total base frame count of all videos is 5, meeting the preset minimum frame count of 5 frames.
[0278] For example, the policy monitoring module determines the maximum number of I-frames in the video. The first threshold can be 2 seconds, s seconds can be 3 seconds, the second threshold can be 30 seconds, k seconds can be 10 seconds, and the third threshold can be 120 seconds.
[0279] For example, taking video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example, see the example below. Figure 12 The diagram shown is a schematic of the I-frame extraction.
[0280] The policy monitoring module receives video 1 and extracts at least one I-frame. At this point, the maximum number of frames in video 1 is 1. The duration of video 1 (93 seconds) is greater than the first threshold (2 seconds), so the maximum number of frames in video 1 is updated to 2.
[0281] Starting from second 0 of video 1's duration, and continuing until second 30 of the video's duration, the number of I-frames increases by one every 3 seconds, such as... Figure 12 As shown, at a video length of 30 seconds, the maximum number of frames in video 1 is 12. Starting from the 30th second of the video duration, and continuing until the 120th second, the number of I-frames increases by one every 10 seconds, such as... Figure 12 As shown, at 93 seconds into the video, the maximum number of frames in video 1 is updated to 18. After the 90th second of video 1, the remaining video duration is 3 seconds, which does not meet the requirement of extracting one I-frame every 10 seconds after 120 seconds. Therefore, after adding one I-frame at the 90th second, the number of I-frames is not increased further. So, the maximum number of I-frames in video 1 with a duration of 93 seconds is 18.
[0282] S232, the strategy monitoring module determines the total number of I-frames to be analyzed for all videos based on the recommended upper limit of the analysis duration.
[0283] The upper limit of analysis time is a suggested maximum time for completing the analysis of all images and videos. Considering the time required for processing images and videos as described in the above embodiments, the total number of I-frames for all videos is determined by multiplying the upper limit of analysis time by the first factor of T1.
[0284] The strategy monitoring module determines the total number of I-frames to be analyzed across all videos based on the first multiplier of the recommended upper limit of analysis duration T1. This first multiplier can be 30%. That is, the number of I-frames that can be analyzed is determined based on 30% (T1*0.3) of the recommended upper limit of analysis duration.
[0285] If T1*0.3 is less than the time required to analyze the base number of I frames for all videos, meaning T1*0.3 is insufficient to analyze the base number of I frames for all videos, then the policy monitoring module can allocate an analysis time T2 for all videos. In this case, the total number of videos analyzed is the sum of the base number of all videos. Specifically, the analysis time T2 is greater than T*0.3, less than T1, and greater than or equal to the time required to analyze the base number of all videos.
[0286] Alternatively, if the duration of T*0.3 is insufficient to analyze the basic number of I-frames of all videos, the policy monitoring module can allocate the entire upper limit of the analysis duration, T1, to the basic number of I-frames of all videos. The actual number of I-frames processed by the policy monitoring module (total number of analyses) is the sum of the basic number of I-frames of all videos.
[0287] If the upper limit of the analysis time T1 is less than the time taken to analyze the basic number of I-frames of all videos, assuming the analysis time per frame is t, then the actual number of I-frames processed by the policy monitoring module (total number of analyses) is T1 / t (frames).
[0288] If T1*0.3 is greater than the time required to analyze the maximum number of I-frames across all videos, meaning T1*0.3 is sufficient to analyze the maximum number of I-frames across all videos, then the policy monitoring module will only allocate T1*0.3 for analyzing the maximum number of I-frames across all videos. The actual number of I-frames processed by the policy monitoring module (the total number of analyses) will be the sum of the maximum number of I-frames across all videos.
[0289] If T1*0.3 is greater than the time taken to analyze the basic number of I frames of all videos, and T1*0.3 is less than the time taken to analyze the basic number of I frames of all videos, assuming the analysis time for a single frame is t, the actual number of I frames processed by the policy monitoring module (total number of analyses) is T1*0.3 / t (frames).
[0290] For example, suppose the input video durations are 15s, 40s, 100s, and 150s, respectively, and the basic number of I-frames for each video is 5, 9, 14, and 14 frames, respectively, totaling 42 frames. The maximum number of I-frames for each video is 7, 13, 19, and 21 frames, respectively, totaling 60 frames.
[0291] Assuming a single-frame analysis time of 200ms, the time required to analyze the basic number of I-frames across all videos is the time required for 42 frames, which is 8400ms; the time required to analyze the maximum number of I-frames across all videos is the time required for 60 frames, which is 12000ms.
[0292] If T1*0.3 is less than 8400ms, meaning T1*0.3 is less than the time required to analyze the base number of I-frames for all videos, the policy monitoring module allocates 8400ms to process the 42 I-frames of all videos, based on the time required to process 42 frames. Alternatively, the policy monitoring module allocates the entire recommended upper limit value T1 of the analysis duration to the analysis of I-frames of all videos. In this case, the actual number of I-frames processed by the policy monitoring module (total number of analyses) is 42 frames. If T1 is less than 8400ms, the actual number of I-frames processed by the policy monitoring module (total number of analyses) is T1 / 200 (frames).
[0293] If T1*0.3 is greater than 12000ms, that is, T1*0.3 is greater than the time required to analyze the maximum number of I-frames of all videos, the policy monitoring module allocates T1*0.3 for I-frame analysis of all videos, and the actual number of I-frames processed (total number of analyses) is 60 frames.
[0294] If T1*0.3 is greater than 8400ms and T1*0.3 is less than 12000ms, the actual number of I-frames (total number of analyses) processed by the policy monitoring module within T1*0.3 is T1*0.3 / 200ms (frames).
[0295] In this embodiment, the upper limit of analysis duration is only a suggested value for planning purposes, but no strong validation is performed. In this embodiment, the estimated analysis duration for video analysis is allowed to exceed the suggested upper limit of analysis duration.
[0296] S233, the strategy monitoring module of the media middle platform framework layer determines the first number of I-frames for each video based on the total number of analyses of all videos.
[0297] After obtaining the total number of image frames to be analyzed for all videos, the policy monitoring module can sequentially allocate the corresponding number of I-frames to be analyzed (the first number) for each video.
[0298] For example, the videos can be sorted from longest to shortest according to their actual duration, and each video can be traversed sequentially. In each iteration, the first value and the second value of the video are updated.
[0299] The initial value of the first value is the total number of analyses; the initial value of the second value of the video is 0.
[0300] During the video traversal, for each video, the second value of that video is incremented by 1, and the first value is decremented by 1, until the first value is 0, that is, until the total number of analyses has been allocated. When the total number of analyses has been allocated, the second value of each video is the corresponding first quantity.
[0301] It should be noted that when assigning the first number of I-frames to each video, the first number of I-frames allocated to each video should be less than or equal to the maximum number of I-frames for that video, and greater than or equal to the base number of I-frames for that video. If the second value of a video is equal to the maximum number before iterating to that video, then that video is skipped in the current and subsequent iterations, and the next video is traversed. Skipped videos do not increment the second value. In this way, the first number corresponding to the actual duration of each video can be obtained.
[0302] For example, refer to Figure 13 , Figure 13This document provides a diagram illustrating the allocation of I-frames to multiple videos. Assume the input videos are Video 1 (15s), Video 2 (20s), Video 3 (30s), and Video 4 (35s), with a maximum number of I-frames of 7, 8, 12, and 12 frames respectively for each video. Assume the total number of frames analyzed for all videos is 38. The policy monitoring module iterates through these four videos, allocating the 38 frames sequentially. Starting from round 1, one frame is allocated to each video per round; that is, when each video is iterated over, its corresponding second value is incremented by 1. After round 7, the second value of Video 1's I-frames has reached its maximum number of I-frames (7 frames), and no further I-frames are allocated to it; that is, this video is skipped in this and subsequent iterations. Round 8 only iterates through Videos 2, 3, and 4. After round 8, the second value of Video 2 has reached its maximum number of I-frames (8 frames), and no further I-frames are allocated to it. Round 9 only iterates through Videos 3 and 4. By the end of the 11th round, the allocated frames for each video were 7, 8, 11, and 11 frames respectively. Videos 1 and 2 had reached their maximum number of frames, while videos 3 and 4 could still be allocated. At this point, 37 frames of the total analysis quantity had been allocated, leaving 1 frame, which was insufficient to allocate to all videos (videos 3 and 4). In the 12th round, the policy monitoring module allocated the remaining 1 frame to the longer video 4, based on the video's duration. After all 38 frames were allocated, the initial number of frames for video 1 was 7, for video 2 it was 8, for video 3 it was 11, and for video 4 it was 12.
[0303] After obtaining the initial number of I-frames for each video, the policy monitoring module can perform highlight segment analysis on each I-frame. This highlight segment analysis can include two phases: a first phase involves an overview analysis based on the initial number of I-frames, and a second phase involves a frame-by-frame analysis based on the target region.
[0304] The first analysis phase refers to the overview analysis of a first number of I-frames for each video. Specifically, for each video, the strategy monitoring module sends the location of one I-frame to the channel interface of the media platform framework layer each time, enabling it to analyze the I-frame at that location and return the analysis results; this continues until the number of I-frames sent for analysis reaches the first number of videos. After the first analysis phase, the strategy monitoring module can obtain the analysis results of the first number of I-frames for each video. This analysis result includes the first score of the I-frame. Therefore, the target region for each video can be determined based on the first I-frame with the highest score. Optionally, for example, upon receiving the first score of each I-frame, the strategy monitoring module can determine the score of the corresponding region based on the first score of the I-frame. Different weights are assigned to regions with different scores. In regions with lower scores, sparse frame extraction is used to select the location of I-frames for analysis; or, in regions with lower scores, no I-frames are extracted for analysis. In regions with higher scores, dense frame extraction is used to select the location of I-frames for analysis. In this way, the region with the highest score is finally determined as the target region.
[0305] The second analysis phase refers to frame-by-frame analysis of the target region. Specifically, for all videos, a frame-by-frame analysis strategy for the target region is determined based on the duration of the target region. For example, the frame-by-frame analysis strategy could be to perform I-frame analysis for the target region of each video; or, the frame-by-frame analysis strategy could be to perform ordinary image frame analysis for the target region of each video. The strategy monitoring module sequentially sends each frame of the target region to the channel interface of the media platform framework layer to obtain the second score of each frame. Thus, based on the second score of each frame of the target region of each video, the second image frame (I-frame or image frame) with the highest second score and the corresponding region are determined as the highlight segment of the video.
[0306] Taking video 1 as an example, the following example illustrates the process of analyzing highlight segments in the video using steps S24-S37. Steps S24-S29 constitute the first analysis stage, and steps S30-S37 constitute the second analysis stage.
[0307] S24, the strategy monitoring module of the media middle platform framework layer sends the file descriptor fd1 of video 1 and the corresponding first position in video 1 to the channel interface of the media middle platform framework layer.
[0308] Here, the first position refers to the position of the first I-frame in video 1. The first position can include one or more preset positions, for example, the first position can include the position at 1 / 3 of the total video duration. The first position also includes a first preset position and a second preset position, where one preset threshold is the position at 1 / 3 of the total video duration, and the second preset position is the position at 2 / 3 of the total video duration. The preset position is generally the position of the I-frame that is first sent for analysis. Upon receiving the first score of the first I-frame sent for analysis, the I-frame at the second position is dynamically determined for analysis.
[0309] For example, refer to Figure 14 As shown, Figure 14 A schematic diagram is given of the first preset position (1 / 3 of the video duration) and the second preset position (2 / 3 of the video duration) in a 15-second video.
[0310] S25, the channel interface of the media middle platform framework layer performs decoding, resolution reduction and format conversion on video 1 according to the file descriptor of video 1, and stores the processed video 1.
[0311] S26, the channel interface of the media middle platform framework layer sends the frame data address of video 1 and the corresponding first position in video 1 to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one-click video creation interface).
[0312] S27, the HAL layer's highlight fragment algorithm interface obtains the I-frame of the first position in video 1 based on the frame data address of video 1 and the corresponding first position in video 1, and then analyzes the I-frame of the first position based on the preset highlight fragment algorithm to obtain the analysis result of the I-frame of the first position.
[0313] In this embodiment, the HAL layer's specular highlight algorithm interface obtains a first position indicating a time point. In reality, the first position may not necessarily contain an I-frame. In this case, the HAL layer's specular highlight algorithm interface can obtain the I-frame closest to the first position at a time point near the first position, and analyze it as the I-frame of that first position to obtain the analysis result of the I-frame at that first position.
[0314] S28, the HAL layer's highlight fragment algorithm interface sequentially returns the analysis results of the first I-frame to the media platform framework layer's strategy monitoring module through the FWK layer's service interface and the media platform framework layer's channel interface.
[0315] The analysis result of the I-frame at the first position can represent the second score of the I-frame.
[0316] S29, the strategy monitoring module of the media middle platform framework layer determines the second position based on the analysis results of the I-frame at the first position, and returns to execute S24 until the number of analyzed I-frames meets the first number of videos.
[0317] In this way, the strategy monitoring module can obtain the analysis results of the first number of I-frames from all videos. These analysis results include the first score of each I-frame. Based on the first score of the I-frame, the plan for the second analysis phase is executed.
[0318] In some embodiments, determining the second location based on the analysis results of the I-frame at the first location by the policy monitoring module in S29 further includes:
[0319] If the first position is a preset position, and the preset position includes the first preset position, the first preset position divides the video into region 1 and region 2, and takes the midpoint of region 1 or region 2 as the second position, and performs I-frame analysis of the second position.
[0320] If the first position is a preset position, the preset position includes a first preset position and a second preset position. The region formed by the midpoint between the start position of the video and the first and second preset positions is defined as the first region. The first region includes the first preset position, and the score of the first region is the score of the I-frame at the first preset position.
[0321] The region formed by the midpoint between the first and second preset positions and the end position of the video is designated as the second region. The second region includes the second preset position, and the score of the second region is the score of the I-frame at the second preset position.
[0322] Calculate the product of the score and duration of the first region, and the product of the score and duration of the second region. Use the midpoint of the region with the higher product as the second position, and perform I-frame analysis for that second position.
[0323] For example, let's take the first preset position as 1 / 3 of the video duration (corresponding to I-frame 1) and the second preset position as 2 / 3 of the video duration (corresponding to I-frame 2) as an example. After obtaining the first scores of the I-frames at the 1 / 3 and 2 / 3 time points of the video, the second position of the video is determined.
[0324] For example, refer to Figure 15 , Figure 15 Another schematic diagram of the first position is given.
[0325] After obtaining the analysis results of I-frame 1 and I-frame 2, the policy monitoring module can divide the region by taking the midpoint between the locations of I-frame 1 and I-frame 2. (Reference) Figure 15Using the midpoint between the 1 / 3 and 2 / 3 time points as the dividing point, the area from the start of the video to the midpoint is designated as Region 1 for I-frame 1, and the area from the midpoint to the end of the video is designated as Region 2 for I-frame 2. For example, if the first score of I-frame 1 is 70, then Region 1 has a score of 70; if the first score of I-frame 2 is 90, then Region 2 has a score of 90.
[0326] In this embodiment, regions with the same score and consecutive time are considered connected regions, and there is no overlap between connected regions. Combined with... Figure 15 In the given example, the connected regions of the video include Region 1 and Region 2. The product of the score and region duration for Region 1 is 70*(7.5-0) = 525; the product of the score and region duration for Region 2 is 90*(15-7.5) = 675. The policy monitoring module determines that Region 2 has the largest product and selects a position within Region 2 (e.g., the midpoint of Region 2) as the second position for analysis.
[0327] After obtaining the first score of an I-frame, if the first score is less than the third threshold, then the score of the region formed within a preset duration before and after the time point of that I-frame is 0. The preset duration can be 500ms. That is, the region within one second before and after the I-frame as the midpoint has a score of 0. When determining the position of the next I-frame, it will not be determined from the region with a score of 0.
[0328] If the first score of an I-frame is greater than the third threshold, meaning the I-frame's score is "Medium," "High," or "High," the policy monitoring module then obtains the adjacent boundaries before and after the I-frame's location. These adjacent boundaries include one or two of the following: the positions of already analyzed adjacent I-frames, the boundaries of regions with a score of 0, the start position of the video, and the end position of the video.
[0329] If an adjacent boundary is the location of an already analyzed adjacent I-frame, the score of that boundary is the score of that I-frame; if an adjacent boundary is the boundary of a region with a score of 0, the score of that boundary is 0; if an adjacent boundary is the start or end position of the video, the score of that boundary is the score of the original region where that position is located. The score of each connected region can include the average of the first scores of the included I-frames. Connected regions do not include regions with a score of 0.
[0330] In this embodiment, different weights can be assigned to scores within different value ranges. For example, a score greater than or equal to a first threshold corresponds to weight 1; a score greater than or equal to a second threshold but less than the first threshold corresponds to weight 2; and a score greater than a third threshold corresponds to weight 3. Weight 1 is greater than weight 2, and weight 2 is greater than weight 3. For example, weight 1 can be 3, weight 2 can be 2, and weight 1 can be 1.
[0331] The policy monitoring module can determine the first and second boundaries of the third region corresponding to the I-frame based on the scores and corresponding weights of the adjacent boundaries of the I-frame, as well as the first score and corresponding weight of the I-frame.
[0332] For example, the adjacent boundaries of I-frame 3 include the position of the adjacent analyzed I-frame 2 before I-frame 3 (P1, such as the position at 10 seconds of the video) and the end position of the video after I-frame 3. Among them, the first score of I-frame 2 is 90, and the first score is greater than the first threshold, corresponding to a weight of 1 (w1, such as 3).
[0333] The boundary position determined after the time point corresponding to I-frame 3 is the position of the end time of the video (P2, for example, at time point 15s). There is no analyzed I-frame at time point 15s, but the original region 2 where time point 15s is located has a score of 90, which is greater than the first threshold, corresponding to weight 1 (w2, for example, 3).
[0334] If the first position is Figure 15 The position of I-frame 3 shown (P3, for example, the 11.75s position of the video) has a first score of 60, which is greater than the second threshold and less than the first threshold, corresponding to a weight of 2 (w3, for example, 2).
[0335] The policy monitoring module determines the boundary position of the third region of I-frame 3 based on the scores and corresponding weights of the adjacent boundaries of I-frame 3, as well as the first score and corresponding weight of I-frame 3, including:
[0336] The first boundary P11 of the third region (region 3) corresponding to I-frame 3 can be calculated using the following formula:
[0337] P11=(P3-P1) / (w1+w3)*w1+P1.
[0338] For example, refer to Figure 16 , Figure 16 A schematic diagram of region 3 is provided. Here, P1 is 10s, w1 is 3, P3 is 11.75s, and w3 is 2. The position of the first boundary P11 is 10.75s.
[0339] Alternatively, the first boundary P11 can be calculated using P11 = P3 - (P3 - P1) / (w3 + w1) * w3.
[0340] It is understandable that P11 can be calculated using the above formula when w1 is greater than w3. If the score of w1 is less than w3, P11 is calculated as follows:
[0341] P11 = (P3 - P1) / (w3 + w1) * w3 + P1; or, P11 = P3 - (P3 - P1) / (w3 + w1) * w1.
[0342] The second boundary P22 of the third region (region 3) corresponding to I-frame 3 can be calculated using the following formula:
[0343] P22=(P2-P3) / (w3+w2)*w3+P3.
[0344] For example, refer to Figure 16 P2 is at 15s, w2 is at 3s, P3 is at 11.75s, and w3 is at 2s. The position of the second boundary P22 is at 12.75s.
[0345] Alternatively, the second boundary P22 can be calculated using P22 = P2 - (P2 - P3) / (w3 + w2) * w2.
[0346] It is understandable that when w3 is less than w2, P11 can be calculated using the above formula. If the score of w3 is greater than w2, P11 is calculated as follows:
[0347] P22 = (P2 - P3) / (w3 + w1) * w2 + P3; or, P22 = P2 - (P2 - P3) / (w3 + w1) * w3.
[0348] Therefore, the positions of the first boundary of region 3 corresponding to I-frame 3 at 10.75s and the second boundary at 12.75s can be determined. That is, referring to... Figure 16 The region from 10.75s to 12.75s in the video corresponds to region 3 of frame 3.
[0349] After determining region 3 of I-frame 3, the second position is determined based on the scores of all connected regions in the video.
[0350] After determining the region 3 corresponding to I-frame 3 based on the first score of I-frame 3, the existing connected regions in the video are determined.
[0351] For example, refer to Figure 17 , Figure 17 A schematic diagram of the connected regions of a video is given. Figure 17 In the video, the original region 2, where I-frame 3 is located, is divided into two connected regions, region 4 and region 5, by region 3. The existing connected regions in the video include region 1, region 4, region 3, and region 5.
[0352] Among them, region 1 includes only one I-frame 1 sent for analysis (position 5s, first score 70), and region 1 is the original region determined in the first round, so the score of region 1 is the first score of I-frame 1; the score of region 4 is 90; region 5 has no analyzed I-frames for the time being, so the score of region 5 is still the original score of region 2, 90.
[0353] The score for Region 3 can be calculated using the following formula:
[0354] The score for region 3 = (the first score of I-frame 3 + the score of the original region where region 3 is located) / 2.
[0355] The original region where region 3 (10.75s-12.75s) is located is region 2, and region 2 has a score of 90. I-frame 3 has a score of 60. Therefore, region 3 has a score of 75.
[0356] Obtain the product of the score and the duration of each connected region, and select the second position from the connected region with the largest product for analysis. For example, the midpoint of the connected region with the largest product can be used as the second position.
[0357] Figure 17 In the calculation, the product corresponding to region 1 is 70*(7.5-0)=525, the product corresponding to region 3 is 75*(12.75-10.75)=150, the product corresponding to region 4 is 90*(10.75-7.5)=292.5, and the product corresponding to region 5 is 90*(15-12.75)=202.5.
[0358] Among them, the connected region with the largest product is region 1. The policy monitoring module can determine the midpoint of region 1 as the second position for analysis, referencing... Figure 17 The position at 3.75s.
[0359] The strategy monitoring module sends the I-frame at the second position for analysis. Based on the first score of I-frame 4 at the second position, it determines region 4 corresponding to I-frame 4, thereby updating the connected components of the video. The third position is determined from the connected components with the largest product, and this process is repeated until the number of I-frames sent for analysis reaches a first threshold.
[0360] After the number of videos 1 issued by the strategy monitoring module of the media platform framework layer meets the first requirement for video 1, if there are other videos, such as video 2, the electronic device can continue to execute S23-S29 above to obtain the analysis results of the first number of I-frames of the other videos. In this way, by repeatedly executing S23-S29 for each video, the strategy monitoring module can obtain the analysis results of the first number of I-frames in each video among multiple image materials.
[0361] After obtaining the analysis results of the first number of I-frames in each video from the image materials, the strategy monitoring module can determine the target region of each video based on the analysis results of the first number of I-frames in each video, and perform frame-by-frame analysis on the target region of each video to obtain the first highlight segment of each video. Before the strategy monitoring module performs frame-by-frame analysis on the target region of each video, the strategy monitoring module of the media platform framework layer can also determine the frame-by-frame analysis strategy for the target region based on the duration of the target region in each video.
[0362] Using video 1 as an example again, combined with Figure 10 The steps S30-S37 shown illustrate the second analysis stage of a video processing method in this embodiment.
[0363] S30, the strategy monitoring module of the media middle platform framework layer determines the target area of each video based on the analysis results of the I-frames of all videos.
[0364] The analysis results include the first score of the first number of I-frames in each video.
[0365] In this embodiment, for each video, after obtaining the first score of a first number of I-frames, the policy monitoring module can determine the first I-frame with the highest first score. Based on this first I-frame and the suggested duration of the highlight segment, candidate highlight segments of the video are determined, and the candidate highlight segments and the regions formed within a third preset duration before and after them are determined as target regions. For example, the duration of the target region can be twice the duration of the candidate highlight segments.
[0366] In some embodiments, the first I-frame may include at least one I-frame. If the first I-frame includes one I-frame, the I-frames before and after the first I-frame within a first preset duration can be directly determined as the target region of video 1. For example, as... Figure 18 (a) provides a schematic of a target region. The first I-frame of video 1 includes I-frame 1 (e.g., ...). Figure 18 In (a) 1), the segment covered by the first preset duration before and after I-frame 1 is the target area.
[0367] Alternatively, if the first I-frame includes a single I-frame, based on the suggested duration of the highlight segment, segments with a duration of half the suggested highlight segment before and after the first I-frame can be identified as candidate highlight segments. I-frames within a third preset duration before and after the candidate highlight segments are identified as the target region of video 1. The duration of the target region can be b times the length of the candidate highlight segment. For example, b can be a number greater than 1 and less than or equal to 2.
[0368] For example, such as Figure 18(b) provides an illustration of another target region. The first I-frame of video 1 includes I-frame 1 (e.g., Figure 18 In (b) 1), the segments with a suggested duration of 1 / 2 of the highlight segments before and after I-frame 1 are candidate highlight segments, and the segments covered by the third preset duration before and after the candidate highlight segments are the target area.
[0369] If the first I-frame includes multiple consecutive I-frames, based on the suggested duration of the highlight segment, the segment covered by the fourth preset duration before the first I-frame of the multiple consecutive first I-frames to the fourth preset duration after the last I-frame of the multiple consecutive second I-frames is identified as a candidate highlight segment.
[0370] For example, such as Figure 19 , Figure 19 This is a schematic diagram of another target region. The first I-frame of video 1 includes consecutive I-frames 1 (e.g., ...). Figure 19 1) I-frame 2 (e.g.) Figure 19 2) I-frame 3 (e.g.) Figure 19 (3) Therefore, the segment covered by the fourth preset duration before I-frame 1 to the fourth preset duration after I-frame 3 is the candidate highlight segment of video 1. The segment covered by the fourth preset duration before and after the candidate highlight segment is the target area.
[0371] After determining the candidate highlight segments for each video, if the sum of the durations of the candidate highlight segments for all videos is greater than the second multiple of the total recommended duration of the highlight segments, the strategy monitoring module needs to adjust the duration of the candidate highlight segments for each video so that the sum of the durations of the candidate highlight segments for all videos is less than or equal to the second multiple of the total recommended duration of the highlight segments.
[0372] The second multiplier can be a number greater than 1, for example, 1.05. For instance, the duration of candidate highlight segments can be reduced based on the length by which the duration of all candidate highlight segments exceeds the second multiplier of the total recommended duration of highlight segments (the length to be adjusted), according to the actual duration of each video. For example, the shorter the video, the more seconds its candidate highlight segments will be reduced.
[0373] Assume we have two input videos: Video 1 has an actual duration of 10 seconds, and Video 2 has an actual duration of 20 seconds. The suggested duration for highlight clips is 6 seconds, and the total suggested duration for highlight clips is 10 seconds. The duration of candidate highlight clips for both Video 1 and Video 2 is 6 seconds; the duration of target regions for both Video 1 and Video 2 is 12 seconds.
[0374] The combined duration of the candidate highlight segments in Videos 1 and 2 is 12 seconds, exceeding the recommended total duration of highlight segments (10 seconds) by 1.05 times (10.5 seconds). The strategy monitoring module needs to adjust the duration of the candidate highlight segments for each video. The required adjustment is 12 seconds - 10 seconds * 1.05 = 1.5 seconds. The strategy monitoring module uses a weighted allocation based on the reciprocal of the video duration, reducing the duration of the candidate highlight segment in Video 1 (10 seconds) by 1 second; updating the duration of the candidate highlight segment in Video 1 from 6 seconds to 5 seconds; reducing the duration of the candidate highlight segment in Video 2 (20 seconds) by 0.5 seconds; and updating the duration of the candidate highlight segment in Video 2 from 6 seconds to 5.5 seconds. Therefore, the combined duration of the candidate highlight segments in Videos 1 and 2, 10.5 seconds, does not exceed 10 seconds * 1.05, and no further adjustment is needed.
[0375] Therefore, the duration of the candidate highlight segment in video 1 is determined to be 5s, and the duration of the target area corresponding to the candidate highlight segment in video 1 is 10s; the duration of the candidate highlight segment in video 2 is 5.5s, and the duration of the target area corresponding to the candidate highlight segment in video 2 is 11s.
[0376] The target region for each video can be determined using the method provided in S30. After determining the target region for each video, S31-S36 can be executed to perform frame-by-frame analysis of the target region.
[0377] S31, the strategy monitoring module of the media middle platform framework layer determines the frame-by-frame analysis strategy for the target area of each video.
[0378] The frame-by-frame analysis strategy for the target region includes either analyzing all image frames of the target region frame by frame (high-precision analysis) or analyzing all I-frames of the target region frame by frame (low-precision analysis).
[0379] If the remaining analysis time is sufficient to analyze the image frames in the target region of all videos, then all image frames in the target region of each video are analyzed; if the remaining analysis time is insufficient to analyze the image frames in the target region of all videos, but sufficient to analyze the I-frames in the target region of all videos, then all image frames in the target region of each video are analyzed; if the remaining analysis time is insufficient to analyze the I-frames in the target region of all videos, then the I-frames in the target region of each video are analyzed.
[0380] The time taken to analyze I-frames within the target regions of all videos can be determined based on the number of I-frames in the target regions of all videos and the analysis time per frame. The time taken to analyze all image frames within the target regions of each video can be determined based on the number of image frames in the target regions of all videos and the analysis time per frame. The remaining analysis time refers to the time used for video analysis minus the total time spent on the overview analysis of keyframes.
[0381] For example, assume that the input video consists entirely of 10 image frames per second (10 FPS) and 1 I-frame per second; the duration of the target region for each video is 5s, 10s, and 15s, respectively. Assume that processing one frame takes 200ms, and the remaining analysis time is 25s.
[0382] The time taken to analyze all image frames in the target area of the three videos were: 5s*10FPS*200ms=10s, 10s*10FPS*200ms=20s, and 15s*10FPS*200ms=30s, for a total of 60s.
[0383] The time taken to analyze all I-frames in the target area of the three videos were: 5s*1FPS*200ms=1s, 10s*1FPS*200ms=2s, and 15s*1FPS*200ms=3s, for a total of 6s.
[0384] The remaining analysis time of 25 seconds does not meet the time required to analyze all image frames in the target area of the three videos, but it does meet the time required to analyze all I-frames in the target area of the three videos. Therefore, the policy monitoring module analyzes all image frames in the target area of the three videos.
[0385] In some embodiments, after determining the frame-by-frame analysis strategy, the strategy monitoring module can also allocate time according to the scores of the target regions of each video from highest to lowest. For example, target region 1 of video 1 has a score of 80, target region 2 of video 2 has a score of 90, and target region 3 of video 3 has a score of 90. The frame-by-frame analysis strategy is per image frame, with a remaining analysis time of 25 seconds. Time is allocated to target regions with higher scores first. Target region 2 has the highest score, and the analysis time for its image frames is 20 seconds, so 20 seconds of time is allocated to video 2. At this point, the remaining analysis time is 5 seconds. Target region 1 has a score of 80, and the analysis time for its image frames is 10 seconds, so the remaining 5 seconds of analysis time is allocated entirely to video 1. The analysis time is exhausted. Video 3 has not received any allocated time, so frame-by-frame analysis is not performed on target region 3.
[0386] After the strategy monitoring module of the media middle platform framework layer determines the frame-by-frame analysis strategy for the target region of each video, for example, the frame-by-frame analysis strategy instructs the analysis of all I-frames of the target region of each video. Taking video 1 as an example, the process of performing I-frame-by-frame analysis of the target region of video 1 may include:
[0387] S32, the strategy monitoring module of the media middle platform framework layer sends the file descriptor fd1 of video 1 and the position of the I-frame of the target area of video 1 to the channel interface of the media middle platform framework layer.
[0388] If the frame-by-frame analysis strategy indicates that I-frames in the target region of all videos are analyzed, the position of the I-frame can be the position of an I-frame in the target region of video 1.
[0389] It is understandable that if the frame-by-frame analysis strategy indicates that all image frames in the target area of the video are analyzed, the position of the I-frame can be the position of an image frame in the target area of video 1.
[0390] S33, the channel interface of the media middle platform framework layer performs decoding, resolution reduction and format conversion on video 1 according to the file descriptor of video 1, and stores the processed video 1.
[0391] S34, the channel interface of the media middleware framework layer sends the frame data address of video 1 and the position of the I-frame to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one-click video creation interface).
[0392] S35, the HAL layer's highlight fragment algorithm interface obtains the I-frame of video 1 based on the frame data address and I-frame position of video 1, and then analyzes the I-frame based on the preset highlight fragment algorithm to obtain the analysis results of the I-frame.
[0393] The S36 and HAL layer specular clip algorithm interfaces sequentially return I-frame analysis results to the media middleware framework layer's strategy monitoring module via the FWK layer's service interface and the media middleware framework layer's channel interface.
[0394] The strategy monitoring module of the media middle platform framework layer continues to send the position of the next I-frame of the target area of video 1 to the channel interface of the media middle platform framework layer for analysis, until all I-frames in the target area have been analyzed.
[0395] S37, the strategy monitoring module of the media middle platform framework layer determines the highlight segment of video 1 from the analysis results of the I-frame of the target area of video 1.
[0396] The analysis results include the second score of the I-frame of the target region.
[0397] In this embodiment, the strategy monitoring module obtains the second I-frame with the highest second score based on the second score of the I-frame. The second I-frame may include at least one I-frame. If the second I-frame includes only one I-frame, the I-frames before and after the second I-frame within a second preset time period can be directly identified as highlight segments of video 1.
[0398] For example, such as Figure 20 (a) provides a schematic diagram of a highlight segment. The second I-frame of video 1 includes I-frame 1 (e.g., ...). Figure 20In (a) 1), the segment covered by the second preset duration before and after I-frame 1 is the highlight segment of video 1. The duration of the highlight segment is less than or equal to the suggested duration of the highlight segment.
[0399] Alternatively, if the second I-frame includes another I-frame, based on the suggested duration of the highlight segment, the segment before and after the second I-frame, which is half the suggested duration of the highlight segment, can be determined as the highlight segment of Video 1. The duration of the highlight segment is equal to the suggested duration of the highlight segment.
[0400] For example, such as Figure 20 (b) provides a schematic diagram of another highlight segment. In this diagram, the second I-frame of video 1 includes I-frame 1 (as shown in Figure 1). Figure 20 In (b) 1), the recommended duration of the 1 / 2 highlight clip before and after I-frame 1 is the highlight clip of video 1.
[0401] If the second I-frame includes multiple consecutive I-frames, the segment covered by the fifth preset duration before the first I-frame of the multiple consecutive second I-frames and the fifth preset duration after the last I-frame of the multiple consecutive second I-frames is determined as the highlight segment of Video 1, based on the suggested duration of the highlight segment.
[0402] For example, such as Figure 21 , Figure 21 Another schematic diagram of a highlight segment is given. The first I-frame of video 1 includes consecutive I-frames 1 (such as...). Figure 21 1) I-frame 2 (e.g.) Figure 21 2) I-frame 3 (e.g.) Figure 21 (3) Therefore, the segment covered by the fifth preset duration before I-frame 1 to the fifth preset duration after I-frame 3 is the highlight segment of video 1.
[0403] If other videos exist, such as video 2, the electronic device can continue to execute S32-S37 above to obtain the analysis results of the highlight segments of the other videos.
[0404] After obtaining the analysis results of the highlight segments of all videos, in some embodiments, reference is made to... Figure 22 This paper presents a flowchart illustrating the post-processing of highlight segments in a video processing method. After executing S37, the electronic device can report all image and video analysis results using the following S38.
[0405] S38, the strategy monitoring module of the media middle platform framework layer reports the analysis results of all images and videos to the application function layer through the image highlight segment analysis interface of the media middle platform framework layer.
[0406] The analysis results can include the location of all highlight segments. For example, the start and end times of each highlight segment in the video.
[0407] S39, the application function layer, based on the analysis results of all images and videos, edits and filters the user-selected materials to obtain all highlight segments.
[0408] S40, the application layer's application function layer calls the media middle platform framework's theme summary interface to request and obtain the theme template.
[0409] The request is transmitted to the HAL layer through the FKW layer via the theme summary interface and channel interface of the media middle platform framework.
[0410] S41, the HAL layer determines a theme template that matches the scene based on the scene in the highlight clip.
[0411] In some embodiments, electronic devices may be configured with multiple theme templates (style templates). The theme algorithm may recommend theme templates that match the scene based on highlight segments, such as people, landscapes, food, children, pets, sports, or travel.
[0412] S42, the HAL layer returns the determined theme template to the theme summary interface of the media middle platform framework layer, and the theme summary interface returns the theme template to the application function layer of the application layer.
[0413] For example, assuming that the highlights are mostly scenes of parents and children, then the theme template that matches the scene can be identified as the parent-child theme class.
[0414] S43, the application function layer sends the obtained subject and all highlight fragments to the basic capability layer.
[0415] The application function layer will send the theme obtained from S42 and all the highlight fragments obtained from S39 to the basic capability layer.
[0416] S44, the basic capability layer generates the target video set based on the theme and all highlight clips.
[0417] That is, the target video set is a set of videos that conforms to the recommended theme, generated from all the selected highlight clips.
[0418] S45, the basic capability layer sends an instruction message to the video editing business layer to display the target video set.
[0419] S46, the video editing business layer displays the target video set in the gallery interface.
[0420] In this embodiment, the media platform framework layer is used for tasks such as decoding video and image files, converting data formats to a unified format, monitoring remaining time and adjusting computational strategies, sending data, controlling algorithm execution and termination, obtaining results, and returning them to the application layer. The FKW layer is used to package data and provide data and program execution services. After receiving commands from the media platform framework layer, the HAL layer performs highlight analysis according to the commands and returns the parameter calculation results of the highlight analysis to the media platform framework layer. The final algorithm results are collected and organized by the media platform framework layer before being sent to the application layer for processing. The application layer can present the editing application interface, video and image file options, and the final algorithm results.
[0421] After the user activates the one-click video creation function, they select the video and image files to be edited (for example, a maximum of 30 files are supported). After waiting for a moment, the "one-click video creation" application automatically edits the highlight segments of the video and combines the highlight segments and images together according to the algorithm results to generate the edited short video, which can then be previewed and played.
[0422] The video processing method provided in this application embodiment allows an electronic device to first perform an overview analysis on a first number of keyframes in each video from multiple image materials, obtaining a first score for each keyframe. Then, based on the first keyframe with the highest first score, and keyframes before and after the first keyframe within a first preset time period, the electronic device can determine the target region of the video. Subsequently, the electronic device can perform frame-by-frame analysis on the target region in the video, obtaining a second score for each frame. Based on the second image frame with the highest second score, and frames before and after the second image frame within a second preset time period, the highlight segments of the video are determined. Using this scheme, the electronic device first performs an overview analysis on the video, thus locating the target region that needs frame-by-frame processing. When extracting highlight segments, the electronic device only performs frame-by-frame analysis on the target region, rather than analyzing the entire video frame-by-frame, reducing the workload of image frame analysis and thus reducing the time required for the electronic device to achieve one-click video creation, improving the efficiency of video processing, and ultimately enhancing the user experience of the one-click video creation function.
[0423] It should also be noted that in the embodiments of this application, "greater than" can be replaced with "greater than or equal to", "less than or equal to" can be replaced with "less than", or "greater than or equal to" can be replaced with "greater than", and "less than" can be replaced with "less than or equal to".
[0424] The various embodiments described herein can be independent solutions or combinations thereof based on their inherent logic, and all such solutions fall within the protection scope of this application.
[0425] It is understood that the methods and operations implemented by electronic devices in the above-described method embodiments can also be implemented by components (such as chips or circuits) that can be used in electronic devices.
[0426] It should be noted that the personal information used in the technical solution of this application is limited to information for which separate consent has been obtained, including but not limited to notifying and reminding users to read the relevant user agreement (notification) and sign the agreement (authorization) which includes authorization of relevant user information before users use the function. Personal information includes images, videos, and other information stored by users.
[0427] The technical solutions disclosed in this application involve the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, all of which comply with relevant laws and regulations and do not violate public order and good morals.
[0428] The method embodiments provided in this application have been described above. The apparatus embodiments provided in this application will be described below. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments. Therefore, any content not described in detail can be referred to the method embodiments above. For the sake of brevity, it will not be repeated here.
[0429] The foregoing mainly describes the solutions provided by the embodiments of this application from the perspective of method steps. It is understood that, in order to achieve the above functions, the electronic device implementing this method includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of protection of this application.
[0430] This application embodiment can divide an electronic device into functional modules based on the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, other feasible division methods may exist. The following description uses the division of functional modules according to each function as an example.
[0431] This application also provides a chip coupled to a memory, which is used to read and execute computer programs or instructions stored in the memory to perform the methods in the above embodiments.
[0432] This application also provides an electronic device including a chip for reading and executing computer programs or instructions stored in a memory, causing the methods in the various embodiments to be performed.
[0433] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.
[0434] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the electronic device in the above method embodiments. For example, the computer may be the aforementioned electronic device.
[0435] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0436] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0437] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0438] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0439] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video processing method, characterized in that, The method includes: The system receives a user's selection of multiple image materials from the gallery; wherein, the multiple image materials include videos. In response to the selection operation, a first number of I-frames are used to perform an overview analysis on each video in the plurality of image materials to obtain a first score; the first score is an aesthetic score of the corresponding image frame, and the first number corresponds to the duration of the video; the I-frame is an image frame that includes complete image information. Based on the first score of each I-frame in the video, the target region of the corresponding video is determined; the target region includes the first I-frame, and I-frames before and after the first I-frame within a first preset time period; the first I-frame is the I-frame with the highest first score. Frame-by-frame analysis is performed on the target region of each video in the multiple image materials to obtain a second score for each image frame in the target region; the frame-by-frame analysis includes I-frame analysis or frame-by-frame analysis of each image frame, and the second score is an aesthetic score for the corresponding image frame; Based on the second score corresponding to each video, the highlight segments of the corresponding video are determined; wherein, the highlight segments include the second image frame, and frames within a second preset time period before and after the second image frame; the second image frame is the image frame with the highest second score in the target region; the highlight segments of each video in the multiple image materials are used to splice together to obtain the target video set.
2. The method according to claim 1, characterized in that, The first score is obtained by performing a first number of I-frame overview analysis on each video in the plurality of image materials, including: Obtain the first position of the analyzed I-frame in the video; Based on the first position and the first score of the analyzed I-frame, the second position of the I-frame to be analyzed in the video is determined, and the first score of the I-frame at the second position is obtained, until the number of analyzed I-frames reaches the first number.
3. The method according to claim 2, characterized in that, The first position includes a first preset position and a second preset position of the video; Determining the second position of the I-frame to be analyzed in the video based on the first position of the analyzed I-frame and the first score of the analyzed I-frame includes: The region between the midpoint between the first preset position and the second preset position and the starting position of the video is defined as the first region; the first region includes the I-frame of the first preset position, and the score of the first region is the first score of the I-frame of the first preset position. The region between the midpoint and the end of the video is defined as the second region; the second region includes the I-frame at the second preset position, and the score of the second region is the first score of the I-frame at the second preset position; Obtain the product of the score for each region and the duration of the region, and take the midpoint of the region with the highest product as the second position.
4. The method according to claim 3, characterized in that, The first preset position is the 1 / 3 mark of the video duration, and the second preset position is the 2 / 3 mark of the video duration.
5. The method according to claim 2, characterized in that, The first position is a non-preset position; Determining the second position of the I-frame to be analyzed in the video based on the first position of the analyzed I-frame and the first score of the analyzed I-frame includes: Based on the adjacent boundaries before and after the I-frame at the first position, a third region containing the I-frame at the first position is determined; the adjacent boundaries include one of the positions of the I-frames adjacent to the I-frame at the first position that have been analyzed, the start position of the video, the end position of the video, and the boundaries of connected regions; wherein, the connected regions are regions with the same score and are time-continuous, and there is no region overlap between the connected regions, and the third region is one of the connected regions of the video. Multiple connected regions of the video are obtained, and a score is calculated for each connected region; the score of the connected region is the average of the first score of the included I-frame and the score of the connected region. Obtain the product of the score for each region and the duration of the region, and take the midpoint of the connected region with the highest product as the second position.
6. The method according to claim 5, characterized in that, Determining the third region corresponding to the I-frame at the first position based on the adjacent boundaries before and after the I-frame at the first position includes: The first boundary of the third region is determined based on the first score and corresponding first weight of the I-frame at the first position, and the scores and corresponding second weights of the adjacent boundaries of the I-frame at the first position. The second boundary of the third region is determined based on the first score and the first weight of the I-frame at the first position, and the scores and corresponding third weights of the adjacent boundaries of the I-frame at the first position. The scoring and weighting are related.
7. The method according to any one of claims 1-6, characterized in that, The process of determining the target region of a corresponding video based on the first score of the I-frame in each video includes: For each video, the first I-frame is obtained based on the first score of the I-frame in the video; The video contains the first I-frame and has a duration that is the preset recommended duration of the highlight segment, which is then used to determine the candidate highlight segments of the video. The candidate highlight segment and the I-frames within a third preset duration before and after the candidate highlight segment are determined as the target area of the video; the third preset duration is less than the first preset duration.
8. The method according to any one of claims 1-6, characterized in that, The method further includes: In response to the selection operation, analysis parameters corresponding to the plurality of image materials are obtained; wherein, the analysis parameters include the actual duration of the corresponding video and a suggested upper limit value for the total analysis duration, the suggested upper limit value for the total analysis duration representing a suggested maximum duration required for analyzing the plurality of image materials; The first number of I-frames for each video is determined based on the actual duration of each video in the plurality of image materials, the number of videos in the plurality of image materials, and the suggested upper limit value of the total analysis duration.
9. The method according to claim 8, characterized in that, The step of determining the first number of I-frames for each video based on the actual duration of each video in the plurality of image materials, the number of videos in the plurality of image materials, and the suggested upper limit value of the total analysis duration includes: Based on the actual duration of each video and a preset correspondence, the basic number and maximum number of corresponding videos are determined; wherein, the basic number is the minimum number of I-frames required to analyze the video to ensure the analysis effect, and the maximum number is the maximum number of I-frames allowed to be analyzed for the duration; the preset correspondence represents the maximum number and basic number of I-frames corresponding to different threshold ranges of video duration. The total number of analyses is determined based on the base number and maximum number of analyses for each video, as well as the suggested upper limit for the total analysis time; the total number of analyses is the total number of I-frames allowed to be analyzed from all videos in the multiple image materials. Based on the actual duration of each video in the plurality of image materials and the number of videos in the plurality of image materials, the total number of analyses is allocated to each video to obtain the first number of I-frames of each video.
10. The method according to claim 9, characterized in that, The analysis parameters also include single-frame analysis duration; the single-frame analysis duration is the time required to analyze one image frame. The determination of the total number of analyses based on the base and maximum number of each video, and the suggested upper limit for the total analysis duration, includes: If the upper limit of the total analysis time is less than the time taken to analyze the sum of the basic number of I-frames of all videos in the multiple image materials, the total number of analyses is the ratio of the upper limit of the total analysis time to the analysis time of a single frame. If the upper limit of the total analysis time is greater than the time taken to analyze the sum of the base number of I-frames of all videos in the image material, and the second multiple of the upper limit of the total analysis time is less than the time taken to analyze the sum of the base number of I-frames of all videos in the multiple image materials, then the total number of analyses is the sum of the base number of I-frames of all videos in the multiple image materials. If the duration of the second multiple of the suggested upper limit of the total analysis time is greater than the time taken to analyze the sum of the basic number of I-frames of all videos in the multiple image materials, and the duration of the second multiple of the suggested upper limit of the total analysis time is less than the time taken to analyze the sum of the maximum number of I-frames of all videos in the multiple image materials, then the total number of analyses is the ratio of the duration of the second multiple of the suggested upper limit of the total analysis time to the duration of the single-frame analysis. If the duration of the second multiple of the suggested upper limit of the total analysis time is greater than the time taken to analyze the sum of the maximum number of I-frames of all videos in the multiple image materials, then the total number of analyses is the sum of the maximum number of I-frames of all videos in the multiple image materials. Among them, the second multiplier is greater than 0 and less than 1.
11. The method according to claim 9 or 10, characterized in that, The step of allocating the total number of analyses to each video according to the actual duration of each video in the plurality of image materials and the number of videos in the plurality of image materials, and obtaining the first number of I-frames of each video, includes: Iterate through each video in the plurality of image materials, updating the first value and the second value of each video until the first value is 0; wherein, the initial value of the first value is equal to the total number of analyses, and the initial value of the second value is 0; for each video traversed, the second value of the video is incremented by 1, and the first value is decremented by 1. The second value of each video is used as the first number of I-frames of the video.
12. The method according to claim 11, characterized in that, The step of traversing each video in the plurality of image materials, updating the first value and the second value of each video, until the first value is 0, includes: Before traversing to the first video among the multiple image materials, if the second value of the first video is equal to the maximum number of I-frames of the first video, then skip the first video and traverse the next video of the first video. Skipping the first video means that the second value of the first video is not incremented by 1.
13. The method according to any one of claims 1-6, characterized in that, The step of performing frame-by-frame analysis on the target region of each video in the plurality of image materials to obtain a second score for each image frame in the target region includes: Based on the preset single-frame analysis duration and the number of all I-frames in the target area of all videos in the image material, the sum of the first durations for analyzing I-frames is obtained; the single-frame analysis duration is the duration required to analyze one image frame; Based on the single-frame analysis duration and the number of all image frames in the target area of all videos in the image material, obtain the sum of the second duration of the analyzed image frames; If the remaining analysis time is less than the sum of the first time, the I-frames of the target region of each video in the plurality of image materials are analyzed frame by frame to obtain the second score of the I-frames in the target region; the remaining analysis time is equal to the time used for video analysis minus the total time spent performing the overview analysis; If the remaining analysis time is greater than or equal to the sum of the first time, or if the remaining analysis time is greater than the sum of the second time, the image frames of the target region in each of the multiple image materials are analyzed frame by frame to obtain the second score of the image frames in the target region.
14. The method according to claim 13, characterized in that, The method further includes: Based on the target regions' scores from highest to lowest, and according to the time taken to analyze all image frames within the target regions, the remaining analysis time is allocated to each target region until the remaining analysis time allocation is complete. If there are target regions that have not been allocated analysis time, then these target regions will not be analyzed frame by frame.
15. The method according to any one of claims 1-6, characterized in that, The process of determining the highlight segments of each video based on the second score includes: For each video, the second image frame is obtained based on the second score of the image frame of the target region; The second image frame and the image frames before and after the second image frame within the second preset duration are determined as the highlight segments of the video; the duration of the highlight segments is less than or equal to the suggested duration of the highlight segments.
16. The method according to any one of claims 1-6, characterized in that, Before performing a first number of I-frame overview analysis on each video in the plurality of image materials to obtain a first score, the method further includes: If the total analysis time of the multiple image materials exceeds the preset upper limit of the total analysis time, M videos are randomly selected from the multiple image materials; wherein, the total analysis time represents the total time required to analyze the multiple image materials, and the preset upper limit of the total analysis time represents the maximum recommended time required to analyze the multiple image materials; the time required to analyze M videos is less than or equal to the total analysis time, and M is less than the number of videos in the image materials.
17. The method according to any one of claims 1-6, characterized in that, The first score is obtained by performing a first number of I-frame overview analysis on each video in the plurality of image materials, including: If the sum of the actual durations of all videos in the plurality of image materials is greater than a preset duration threshold, an overview analysis of the first number of I-frames is performed on each video in the plurality of image materials to obtain the first score.
18. An electronic device, characterized in that, The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1-17.
19. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-17.
20. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method of any one of claims 1-17.
Citation Information
Patent Citations
Video editing method and device, equipment and storage medium
CN113709560A
Video processing method and electronic equipment
CN115567660A