Video segment processing method, terminal device, and storage medium

Through intensive or sparse key segment analysis strategies, the problem of long-term video analysis is solved, and the rapid generation of highlight segment video sets is achieved, which improves the user experience.

CN118075513BActive Publication Date: 2025-08-05HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211466666.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-08-05
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

When using the one-click filming function in the prior art, it takes too long to analyze long material videos, affecting the user experience.

Method used

The intensive or sparse key segment analysis strategy is adopted to quickly select highlight segments at different locations in long videos by calculating the duration, number and interval duration of the analysis segments to generate a highlight segment video set.

Benefits of technology

Quickly analyze highlight clips within a limited time, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118075513B_ABST
    Figure CN118075513B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method for processing video segments, a terminal device, and a storage medium, which relate to the technical field of video data processing and are used to solve the technical problem that a large amount of time will be consumed if the entire source video is analyzed when using the one-click video generation function. In the solution of the present application, if the selected source video is too long and the highlight segments are distributed at different positions in the source video, then the terminal device can adopt a dense key segment analysis strategy or a sparse key segment analysis strategy to select as many key segments as possible covering different positions in the long video from the long video, and select the highlight segments from the multiple key segments. In this way, by analyzing the positions of the key segments of the long video within a limited time consumption, the highlight segments can be quickly analyzed and a video set composed of the highlight segments can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video data processing, and particularly to a method for processing video segments, a terminal device, and a storage medium. Background Art

[0002] With the rapid development of the short video industry, the demand for synthesizing highlight segments of multiple material videos into a short creative video is increasing.

[0003] Currently, some terminal devices provide a function of generating a video in one click. If a user uses the function of generating a video in one click after selecting videos and pictures, the terminal device can automatically analyze and extract the highlight segments in the material videos through an algorithm, and combine the highlight segments into a clipped video. However, when the selected material video is too long, analyzing the entire material video will consume a large amount of time, thus affecting the user experience. Summary of the Invention

[0004] This application provides a method for processing video segments, a terminal device, and a storage medium, which solves the technical problem that analyzing the entire material video will consume a large amount of time.

[0005] To achieve the above object, this application adopts the following technical solutions:

[0006] In a first aspect, an embodiment of this application provides a method for processing video segments. The method includes:

[0007] In response to a user's selection operation on multiple videos, determine the analysis time consumption and available analysis duration of each video in the multiple videos. The analysis time consumption of a video is the duration required to complete the image analysis of all frames of the video, and the available analysis duration of a video is the expected duration allocated for the image analysis of the video. In the case where the available analysis duration of the first video in the multiple videos is less than the analysis time consumption of the first video, and the available analysis duration of the first video is greater than the analysis duration of P video frames, where P is an integer greater than or equal to 2, determine the duration of the analysis segments corresponding to the first video, the number of analysis segments, and the interval duration between adjacent segments. According to the duration of the analysis segments corresponding to the first video, the number of analysis segments, and the interval duration between adjacent segments, determine multiple analysis segments in the first video. Extract highlight segments from the multiple analysis segments. After extracting all highlight segments from the multiple videos, generate a video set from all the highlight segments.

[0008] Through the above solution, when the user uses the function of generating a video in one click, if the available analysis duration of a certain video among multiple videos is less than the analysis time consumption of this video and greater than the analysis duration of several video frames, then by calculating the duration of the analysis segment corresponding to this video, the number of analysis segments, and the interval duration between adjacent segments, multiple key segments covering different positions in the long video can be selected in this video, and then the highlight segments are selected from multiple key segments. In this way, by analyzing the positions of the key segments of the long video under limited time consumption, the highlight segments can be quickly analyzed and a video set composed of highlight segments can be generated.

[0009] This application provides two analysis strategies for selecting key segments: the intensive key segment analysis strategy or the sparse key segment analysis strategy. Among them, the intensive key segment analysis strategy is characterized in that multiple key segments are evenly distributed throughout the video; the sparse key segment analysis strategy is characterized in that multiple key segments are mainly concentrated in the first and middle segments of the video, and are appropriately distributed in the tail segment.

[0010] Suppose the available analysis duration allocated for the first video is represented by T 预期分析 and the analysis time consumption of the first video is represented by T 分析耗时 and the chip analysis speed is represented by T 芯片分析速度 . The ratio of the available analysis duration T 预期分析 allocated for the first video to the analysis time consumption T 分析耗时 of the first video is represented by R, then the following relationship exists:

[0011]

[0012] Among them, T 总 represents the total duration of the first video. T 芯片分析速度 represents the chip analysis speed. That is, for each video among multiple videos, the analysis time consumption of a video is equal to the total duration of a video divided by the analysis speed of a video. Among them, the analysis speed of a video is determined according to the resolution, frame rate of a video and the chip analysis speed.

[0013] The intensive key segment analysis strategy is adopted when the following relationship is satisfied:

[0014]

[0015] The sparse key segment analysis strategy is adopted when the following relationship is satisfied:

[0016]

[0017] Among them, R1 is less than 1, for example, R1 = 0.6.

[0018] The common point of these two analysis strategies is that they are based on the calculation of three target parameters: the duration T of the analysis segment corresponding to the first video 分析 , the number of analyzed fragments N 分析 , the interval length T between adjacent segments 间隔 .

[0019] The following is an example of how to determine these three target parameters.

[0020] In one possible implementation, the duration T of the analysis segment corresponding to the first video is determined. 分析 ,include:

[0021] According to the total duration T of the first video 总 , the available analysis time of the first video T 预期分析 , and the minimum interval T of highlight fragments 最小间隔 , determine the maximum number N of analysis segments in the first video 上限 The available analysis time of the first video is T 预期分析 The maximum number of analyzed segments N in the first video 上限 The ratio of the first video is used as the minimum length T of the analysis segment in the first video. 下限 The minimum length T of the fragment to be analyzed 下限 and the maximum duration T of the highlight clip 长高光 The maximum of the two is taken as the duration T of the analysis segment corresponding to the first video. 分析 .

[0022] For example, the maximum number N of analysis segments in the first video can be determined using the following relationship: 上限 :

[0023]

[0024] in, N is the floor symbol. 上限 Indicates the maximum number of analysis segments in the first video. 总 Indicates the total duration of the first video. 预期分析 Indicates the available analysis time of the first video. 最小间隔 Indicates the minimum interval of highlight segments.

[0025] It should be understood that the total duration of the first video T 总 The available analysis time T allocated for the first video 预期分析 The difference is the total duration of all interval segments. Divide the total duration of all interval segments by the minimum interval T of the highlight segment. 最小间隔 , we can get the upper limit of the spacer segment. Since a spacer segment is set between every two analysis segments, the number of analysis segments is one more than the number of spacer segments.总 , T 预期分析 and T 最小间隔 can obtain the upper limit value N of the analysis segment 上限 . Additionally, taking the maximum value between the minimum duration T 下限 of the analysis segment and the maximum duration T 长高光 of the highlight segment as the duration T 分析 of the analysis segment can appropriately increase the duration of the analysis segment, so as to cover the positions of more highlight segments as much as possible, and then improve the accuracy of the finally extracted highlight segments.

[0026] In a possible implementation, determining the number N 分析 of analysis segments corresponding to the first video includes:

[0027] According to the ratio of the available analysis duration T 预期分析 of the first video to the duration T 分析 of the analysis segment, determining the number N 分析 of analysis segments corresponding to the first video.

[0028] Exemplarily, if T 剩余 < T 短高光 , the following relational expression can be used to determine the number N 分析 of analysis segments corresponding to the first video:

[0029]

[0030] Exemplarily, if T 剩余 ≥ T 短高光 , the following relational expression can be used to determine the number N 分析 of analysis segments corresponding to the first video:

[0031]

[0032] Among them, T 预期分析 represents the available analysis duration of the first video; T 分析 represents the duration of the analysis segment corresponding to the first video; T 短高光 represents the minimum duration of the highlight segment; N 分析 represents the number of analysis segments corresponding to the first video.

[0033] It should be understood that when the remaining analysis duration T 剩余 is greater than the minimum duration T 短高光 of the highlight segment, by adding one analysis segment, all the expected analysis duration can be effectively utilized, covering the positions of more highlight segments, so as to improve the accuracy of the finally extracted highlight segments.

[0034] In a possible implementation, determine the interval duration T of adjacent segments corresponding to the first video 间隔 , including:

[0035] Based on the number N of analysis segments corresponding to the first video 分析 , the total duration T of the first video 总 , and the available analysis duration T of the first video 预期分析 , determine the interval duration T of adjacent segments corresponding to the first video 间隔 .

[0036] Exemplarily, the following relational expression can be used to determine the interval duration T of adjacent segments corresponding to the first video 间隔 :

[0037]

[0038] where the max() function is used to obtain the maximum value; T 间隔 represents the interval duration of adjacent segments corresponding to the first video; T 总 represents the total duration of the first video; T 预期分析 represents the available analysis duration of the first video; N 分析 represents the number of analysis segments corresponding to the first video; T 最小间隔 represents the minimum interval of highlight segments.

[0039] It should be understood that the minimum interval T of highlight segments 最小间隔 is a minimum interval calculated by the application layer. To avoid the problem of overly concentrated analysis segments due to too small intervals, when T 间隔 is less than the minimum interval T of highlight segments 最小间隔 , T 间隔 should be set to T 最小间隔 .

[0040] In a possible implementation, when the ratio (R) of the available analysis duration of the first video to the analysis time consumption of the first video is greater than or equal to a preset ratio (R1), it can be determined that the first video adopts an intensive key segment analysis strategy, and the following method is used to equally allocate each analysis segment of the first video:

[0041] Allocate analysis segments starting from the first frame of the first video until the last frame of the first video, and finally obtain multiple analysis segments. Among them, the duration of each analysis segment in the first video is equal to the analysis segment duration T 分析 , the interval between any two adjacent analysis segments in the first video is equal to the interval duration T of adjacent segments 间隔 , and the number of multiple analysis segments in the first video is equal to the number of analysis segments N 分析 .

[0042] It should be understood that compared with decoding, downscaling, and scoring all video frames, the intensive analysis segment analysis strategy only needs to decode, downscale, and score the video frames of several key segments. Therefore, the intensive analysis segment analysis strategy can obtain highlight segments faster and generate a video set composed of highlight segments.

[0043] In a possible implementation, when the ratio (R) of the available analysis duration of the first video to the analysis duration of the first video is less than a preset ratio (R1), it can be determined that the first video adopts a sparse key segment analysis strategy, and the following method is used to allocate each analysis segment of the first video at unequal intervals:

[0044] Select one video frame from Q video frames as the starting point of the first analysis segment in the first video. According to the interval duration between adjacent segments and the minimum interval of highlight segments, re-determine the intervals of each analysis segment. Allocate analysis segments starting from the starting point of the first analysis segment until the last frame of the first video, and finally obtain multiple analysis segments. Here, Q is an integer greater than or equal to 2. The duration of an analysis segment in the first video is equal to the duration T of the analysis segment 分析 ; the interval between adjacent analysis segments in the first video is less than or equal to the interval duration T of adjacent segments 间隔 , and the interval before an analysis segment is less than or equal to the interval after an analysis segment; the number of multiple analysis segments in the first video is less than or equal to the number N of analysis segments 分析 .

[0045] It should be understood that compared with the intensive analysis segment analysis strategy, the sparse key segment analysis strategy focuses on analyzing the first and middle segments of the video, and the starting point of the first analysis segment may not be the first video frame. Therefore, the sparse key segment analysis strategy may take less time, so that highlight segments can be obtained faster and a video set composed of highlight segments can be generated.

[0046] In a possible implementation, Q = 3. The above Q video frames may include: the first video frame at the starting point of the first video, the video frame at the 2-second position of the first video, and the video frame at the 1 / 5 position of the first video.

[0047] Correspondingly, determining a video frame from Q video frames as the starting point of the first analysis segment in the first video may include: obtaining the score of the first video frame located at the starting point of the first video. In the case where the score of the first video frame located at the starting point of the first video is greater than or equal to a preset value, the first video frame located at the starting point of the first video is taken as the starting point of the first analysis segment; or, in the case where the score of the first video frame located at the starting point of the first video is less than the preset value, obtaining the score of the video frame located at the 2 - second position of the first video. In the case where the score of the video frame located at the 2 - second position of the first video is greater than or equal to the preset value, the video frame located at the 2 - second position of the first video is taken as the starting point of the first analysis segment; or, in the case where the score of the video frame located at the 2 - second position of the first video is less than the preset value, obtaining the score of the video frame located at the 1 / 5 position of the first video. In the case where the score of the video frame located at the 1 / 5 position of the first video is greater than or equal to the preset value, the video frame located at the 1 / 5 position of the first video is taken as the starting point of the first analysis segment; or, in the case where the score of the video frame located at the 1 / 5 position of the first video is less than the preset value, the video frame with the highest score among the first video frame located at the starting point of the first video, the video frame located at the 2 - second position of the first video, and the video frame located at the 1 / 5 position of the first video is taken as the starting point of the first analysis segment.

[0048] It should be understood that different from the third key frame of the simple I - frame - based analysis strategy being "the I - frame located at the 1 / 3 position of the first video", the third key frame of the sparse key segment analysis strategy is "the I - frame located at the 1 / 5 position of the first video". This is because the simple I - frame - based analysis strategy directly determines a highlight segment, while the sparse key segment analysis strategy needs to first determine multiple analysis segments and then screen out the highlight segment from multiple analysis segments. In order to make multiple analysis segments cover the first and middle parts of the first video as much as possible, the third key frame of the sparse key segment analysis strategy is closer to the starting position of the video.

[0049] In a possible implementation manner, re - determining the intervals of each analysis segment according to the interval duration between adjacent segments and the minimum interval of the highlight segment includes:

[0050] If the number of analysis segments is greater than or equal to 2, the following relational formula is used to determine the first interval duration set between the first analysis segment and the second analysis segment:

[0051] T1 = max(T 间隔 / n, T 最小间隔 )

[0052] If the number of analysis segments is greater than or equal to 3, the following relational expression is used to determine the i-th interval duration set between the i-th analysis segment and the (i + 1)-th analysis segment:

[0053] T i = min(m * T i-1 , T 间隔 ).

[0054] Where, T i represents the i-th interval duration; T i-1 represents the (i - 1)-th interval duration set between the (i - 1)-th analysis segment and the i-th analysis segment; both n and m are greater than 1, and i is an integer greater than or equal to 2.

[0055] It should be understood that by setting a smaller interval for the first half of the video and a larger interval for the middle or end of the video, multiple key segments can be mainly concentrated in the first half of the video, and the end segment can be appropriately allocated.

[0056] In a possible implementation, for the intensive key segment analysis strategy or the sparse key segment analysis strategy, if the remaining duration of the first video is less than the duration of the analysis segment after allocating the last interval, the remaining duration of the first video is used as the duration of the last analysis segment in the first video.<***

[0057] In a possible implementation, the present application also provides another analysis strategy: the full - volume analysis strategy. After determining the analysis time consumption and available analysis duration of each video in multiple videos, before extracting all highlight segments from multiple videos, the method may further include: when the available analysis duration of the second video in multiple videos is greater than or equal to the analysis time consumption of the second video, obtaining the score of each video frame in the second video; then using the video frames in the second video with scores greater than or equal to a preset value as the highlight segments in the second video.

[0058] It should be understood that when the analysis strategy set for a certain video is the full - volume analysis strategy, since the terminal device analyzes each frame of the video file one by one, the final analysis result can completely include all the highlight segments of the video, but the time and resource consumption are relatively large.

[0059] In a possible implementation, the present application also provides another analysis strategy: the simple I - frame - based analysis strategy. After determining the analysis time consumption and available analysis duration of each video in multiple videos, before extracting all highlight segments from multiple videos, the method may further include:

[0060] When the available analysis duration of the third video among multiple videos is less than or equal to the analysis duration of 3 video frames, obtain the score of the first video frame located at the starting point of the third video. When the score of the first video frame located at the starting point of the third video is greater than or equal to the preset value, take the first video frame located at the starting point of the third video as the starting point of the highlight segment, and select a video segment with the target duration as the highlight segment of the third video; or, when the score of the first video frame located at the starting point of the third video is less than the preset value, obtain the score of the video frame at the 2-second position of the third video. When the score of the video frame at the 2-second position of the third video is greater than or equal to the preset value, take the score of the video frame at the 2-second position of the third video as the starting point of the highlight segment, and select a video segment with the target duration as the highlight segment of the third video; or, when the score of the video frame at the 2-second position of the third video is less than the preset value, obtain the score of the video frame at the 1 / 3 position of the third video. When the score of the video frame at the 1 / 3 position of the third video is greater than or equal to the preset value, take the score of the video frame at the 1 / 3 position of the third video as the starting point of the highlight segment, and select a video segment with the target duration as the highlight segment of the third video; or, when the score of the video frame at the 1 / 3 position of the third video is less than the preset value, take the highest score among the first video frame located at the starting point of the third video, the video frame at the 2-second position of the third video, and the video frame at the 1 / 3 position of the third video as the starting point of the highlight segment, and select a video segment with the target duration as the highlight segment of the third video.

[0061] Exemplarily, the target duration can be determined by the following relational expression:

[0062] T 高光 = min(1.2 * T 短高光 , T 长高光 ).

[0063] Where, the min() function is used to obtain the minimum value; T 高光 represents the target duration; T 短高光 represents the minimum duration of the highlight segment; T 长高光 represents the maximum duration of the highlight segment.

[0064] It should be understood that when the available analysis duration allocated for a certain video is less than or equal to the analysis time of three key frames, it indicates that the available analysis duration allocated for this video is very short, and at this time, it is not sufficient to analyze a segment of this video. At this time, a simple I-frame based analysis strategy can be adopted to analyze three pictures located in the middle front section of this video, and the one with the highest score among these three pictures is used as the starting point of the highlight segment result, and a video segment with a certain duration is intercepted as the only highlight segment. On the one hand, since the simple I-frame based analysis strategy requires analyzing at most three pictures, compared with several other analysis strategies, the simple I-frame based analysis strategy can obtain the highlight segment more quickly. On the other hand, since users have a relatively high probability of being interested in the content in the middle front section of the video, using the one with the highest score among the three pictures in the middle front section as the starting point of the highlight segment result can ensure the reliability of the extracted highlight segment to a certain extent.

[0065] In a possible implementation manner, for each of multiple videos, the analysis time consumption of each video can be determined in the following way: Determine the analysis speed of each video according to the chip analysis speed, the resolution and frame rate of each video. Determine the target total analysis duration according to the analysis speed of each video, and the target total analysis duration is the expected total duration for completing the analysis of multiple videos. Determine the analysis time consumption of each video according to the target total analysis duration and the analysis time consumption of each video.

[0066] In the second aspect, the present application provides a processing device for video segments, and the device includes units / modules for executing the method in the first aspect above. The device can correspond to executing the method described in the first aspect above, and for the relevant descriptions of the units / modules in the device, please refer to the description in the first aspect above. For the sake of brevity, it will not be repeated here.

[0067] In the third aspect, a terminal device is provided, including a processor, and the processor is coupled with a memory. The processor is used to execute the computer program or instruction stored in the memory, so that the terminal device implements the video segment processing method in any one of the first aspect.

[0068] In the fourth aspect, a chip is provided, and the chip is coupled with a memory. The chip is used to read and execute the computer program stored in the memory to implement the video segment processing method in any one of the first aspect.

[0069] In the fifth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program. When the computer program runs on a terminal device, the terminal device is enabled to execute the video segment processing method in any one of the first aspect.

[0070] In a sixth aspect, a computer program product is provided. When the computer program product runs on a computer, it causes the computer to execute the video segment processing method according to any one of the first aspect.

[0071] It can be understood that for the beneficial effects of the above second aspect to the sixth aspect, reference can be made to the relevant descriptions in the first aspect above, and details are not repeated here. Description of the Drawings

[0072] Figure 1 FIG. 1 is one of the schematic diagrams of the application scenario of the video segment processing method provided by the embodiment of the present application;

[0073] Figure 2 FIG. 2 is another schematic diagram of the application scenario of the video segment processing method provided by the embodiment of the present application;

[0074] Figure 3 FIG. 3 is a third schematic diagram of the application scenario of the video segment processing method provided by the embodiment of the present application;

[0075] Figure 4 FIG. 4 is a fourth schematic diagram of the application scenario of the video segment processing method provided by the embodiment of the present application;

[0076] Figure 5 FIG. 5 is a schematic diagram of the hardware structure of the terminal device provided by the embodiment of the present application;

[0077] Figure 6 FIG. 6 is a schematic diagram of the software structure of the terminal device provided by the embodiment of the present application;

[0078] Figure 7 FIG. 7 is a schematic diagram of the overall process of the video segment processing method provided by the embodiment of the present application;

[0079] Figure 8 FIG. 8 is a schematic diagram of the process of setting an analysis strategy for each video in sequence provided by the embodiment of the present application;

[0080] Figure 9 FIG. 9 is a schematic diagram of the time consumption of four analysis strategies provided by the embodiment of the present application;

[0081] Figure 10 FIG. 10 is a schematic diagram of the analysis segments of the full - scale analysis strategy provided by the embodiment of the present application;

[0082] Figure 11 FIG. 11 is a schematic diagram of the process of the full - scale analysis strategy provided by the embodiment of the present application;

[0083] Figure 12 FIG. 12 is a schematic diagram of the highlight segments when the I - frame - based simple analysis strategy is adopted provided by the embodiment of the present application;

[0084] Figure 13Schematic diagram of the process based on the simple I-frame analysis strategy provided by the embodiments of this application;

[0085] Figure 14 Schematic diagram of two methods for determining the duration of highlight segments provided by the embodiments of this application;

[0086] Figure 15 One of the schematic diagrams of the analysis segments using the intensive key segment analysis strategy provided by the embodiments of this application;

[0087] Figure 16 Schematic diagram of five initial parameters provided by the embodiments of this application;

[0088] Figure 17 Schematic diagram of the process for calculating three target parameters provided by the embodiments of this application;

[0089] Figure 18 Another schematic diagram of the analysis segments using the intensive key segment analysis strategy provided by the embodiments of this application;

[0090] Figure 19 Another schematic diagram of the analysis segments using the intensive key segment analysis strategy provided by the embodiments of this application;

[0091] Figure 20 Another schematic diagram of the analysis segments using the intensive key segment analysis strategy provided by the embodiments of this application;

[0092] Figure 21 Schematic diagram of the process for obtaining highlight segments according to the intensive key segment analysis strategy provided by the embodiments of this application;

[0093] Figure 22 Schematic diagram of the process for determining the starting point of the first analysis segment using the sparse key segment analysis strategy provided by the embodiments of this application;

[0094] Figure 23 One of the schematic diagrams of the analysis segments using the sparse key segment analysis strategy provided by the embodiments of this application;

[0095] Figure 24 Schematic diagram of the process for determining the interval duration using the sparse key segment analysis strategy provided by the embodiments of this application;

[0096] Figure 25 Another schematic diagram of the analysis segments using the sparse key segment analysis strategy provided by the embodiments of this application;

[0097] Figure 26 Another schematic diagram of the analysis segments using the sparse key segment analysis strategy provided by the embodiments of this application;

[0098] Figure 27 A flowchart for obtaining highlight segments according to the sparse key segment analysis strategy provided by an embodiment of this application;

[0099] Figure 28 A schematic structural diagram of a processing device for video segments provided by an embodiment of this application. Detailed implementation manners

[0100] First, some nouns or terms involved in this application are explained.

[0101] Highlight segments, also known as wonderful segments, refer to images extracted from videos or pictures that record wonderful moments such as people's smiling faces, championship moments, aircraft landings, etc.

[0102] The one - click video generation function means that after a user selects one or more materials, the algorithm automatically analyzes the highlight segments in the materials and combines the highlight segments into a clipped video.

[0103] In an embodiment of this application, the terminal device can obtain the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate values of each frame image in a video or picture, and score each frame image based on these data. If the score of a frame image is greater than or equal to a preset score, then this frame image can be used as a highlight segment. Then, the terminal device can synthesize all the highlight segments extracted from multiple materials into a video set and present it to the user.

[0104] Currently, the one - click video generation function supports the editing of materials such as pictures and videos. When a user uses the one - click video generation function to select a relatively large number of materials or a long video, if the terminal device analyzes each frame image of each material and extracts highlight segments from it, it will cause a long time consumption, so that the user needs to wait for a long time to view the video set synthesized by the highlight segments, thereby reducing the user experience when using the one - click video generation function.

[0105] In view of the above problems, an embodiment of this application provides a method for processing video segments. When a user uses the one - click video generation function, if the selected material video is too long and the highlight segments are distributed at different positions in the material video, then the intensive key segment analysis strategy or the sparse key segment analysis strategy can be adopted to select multiple key segments that can cover different positions (such as the front, middle, and back segments of the video) in the long video as much as possible from the long video, and select highlight segments from the multiple key segments. In this way, by analyzing the positions of the key segments of the long video within a limited time consumption, the highlight segments can be quickly analyzed and a video set composed of the highlight segments can be generated.

[0106] The video segment processing method provided by the embodiments of this application can be applied to terminal devices with image shooting and image processing capabilities. It should be noted that the materials used for generating a video in one click in the embodiments of this application can include pictures and videos. Users can use the one-click video generation function to generate a video set from the highlight segments of multiple pictures, or use the one-click video generation function to generate a video set from the highlight segments of multiple videos, or use the one-click video generation function to generate a video set from the highlight segments of multiple pictures and videos.

[0107] The above terminal device can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city or wireless terminal in smart home, etc. The embodiments of this application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0108] Next, taking the terminal device as a mobile phone as an example, combined with Figures 1 - 4 an example is given to illustrate the application scenario of the video segment processing method provided by the embodiments of this application.

[0109] In an exemplary application scenario, a user can use a mobile phone to pre-shoot multiple materials, and the mobile phone stores the multiple materials in the picture library. As shown in (a) of Figure 1 , an icon of the picture library is displayed on the desktop of the mobile phone. When the user wants to use the mobile phone to generate a clipped video set based on multiple materials, the user can click the icon of the picture library on the mobile phone desktop. In response to the user's click operation on the icon, the mobile phone displays a picture library interface UI1 as shown in (b) of Figure 1 , and the "One-Click Blockbuster" option is provided in the picture library interface UI1. As shown in Figure 1As shown in (c) therein, the user can click on the "One-Click Blockbusters" option. In response to the user's click operation on the "One-Click Blockbusters" option, the mobile phone displays a gallery interface UI2 as shown in Figure 1 (d) therein. The gallery interface UI2 may include multiple recently captured materials. As shown in Figure 1 (e) therein, the user clicks on any one of the multiple materials. In response to the user's click operation on the material, the mobile phone displays a gallery interface UI3 as shown in Figure 1 (f) therein. The gallery interface UI3 includes all the materials in the gallery. In this way, the user can select the materials required to generate video clips in the gallery interface UI3. For example, as shown in Figure 1 (f) therein, assume that the user selects 5 materials. The gallery interface UI3 may also include a video generation option, such as the checkmark option as shown in Figure 1 (f) therein.

[0110] Continuing to refer to Figure 2 (a) therein, after the user selects the materials, the user can click on the video generation option, such as clicking on the checkmark option. As shown in Figure 2 (b) therein, in response to the user's click operation on the checkmark option, the mobile phone starts to analyze the 5 materials selected by the user in the gallery interface UI3, selects the highlight segments from each material, and generates a video set based on the selected highlight segments. During this process, the mobile phone can display the material analysis progress in the gallery interface UI3 so that the user can intuitively view the analysis progress.

[0111] In one example, as shown in Figure 3 (a) therein, after generating the video set, the mobile phone displays a gallery interface UI4. The gallery interface UI4 includes the generated video set, and the mobile phone can automatically play the video set. In addition, a video export option may be provided in the gallery interface UI4. As shown in Figure 3 (b) therein, the user can click on the video export option. As shown in Figure 3 (c) therein, in response to the user's click operation on the video export option, the mobile phone exports the video set and stores the video set in the gallery, so that the user can view the video set from the gallery.

[0112] In another exemplary application scenario, after the user uses the mobile phone to pre-capture multiple videos and multiple images, the mobile phone stores the multiple videos and multiple images in the gallery. In this way, the user can select videos and images in the gallery, and the mobile phone generates a video set based on the videos and images. In one example, as shown in Figure 4As shown in (a) of , after entering the gallery interface UI2, the gallery interface UI2 includes multiple recently taken videos and multiple images. The user can click on any one of the multiple videos and multiple images, for example, click on the first piece of material. In response to the user's click operation on any one of the materials, the mobile phone displays the gallery interface UI3, and the gallery interface UI3 includes all the materials in the gallery. In this way, the user can select the materials required to generate a video set in the gallery interface UI3, for example, as shown in Figure 4 In (b) of , assume that the user selects 4 videos and 2 pictures. The gallery interface UI3 also includes a video generation option, for example, as shown in Figure 4 In (b) of , the video generation option is a checkmark option. As shown in Figure 4 In (c) of , after the user selects the materials, the user can click on the video generation option, for example, click on the checkmark option. As shown in Figure 4 In (d) of , in response to the user's click operation on the checkmark option, the mobile phone starts to analyze the 4 videos and 2 pictures selected by the user in the gallery interface UI3, selects the highlight segments from each material, and selects the highlight pictures from the 2 pictures. During this process, the mobile phone can display the progress of analyzing the materials in the gallery interface UI3 so that the user can intuitively view the analysis progress.

[0113] In one example, as shown in Figure 3 In (a) of , after generating the video set, the mobile phone displays the gallery interface UI4, and the gallery interface UI4 includes the generated video set. The mobile phone can automatically play the video set. In addition, a video export option can be provided in the gallery interface UI4. As shown in Figure 3 In (b) of , the user can click on the video export option. In response to the user's click operation on the video export option, the mobile phone exports the video set, as shown in Figure 3 In (c) of , and stores the video set in the gallery. In this way, the user can view the video set from the gallery.

[0114] In one example, if the number of materials selected by the user is small, during the process of the user selecting materials, the mobile phone can prompt the user so that the user can know how many materials are more appropriate to select. For example, as shown in Figure 1 In (f) of , the prompt message "Better results with more than 6 materials" is displayed in the gallery interface UI3, so that the user can know at least how many materials to select to generate a video set with better results.

[0115] In one example, if the number of materials selected by the user is large, during the process of the user selecting materials, the mobile phone can prompt the user so that the user can know the maximum number of materials that can be selected. For example, as shown in Figure 4As shown in (b) thereof, a prompt message "Up to 30 materials can be selected" is displayed in the gallery interface UI3, so that it is convenient for the user to know how many materials can be selected.

[0116] In one example, after generating a video set, the mobile phone can also provide other function options in the application interface UI4 where the video set is displayed, so that the user can perform operations such as editing, adding special effects, and analyzing the generated video set based on these function options. For example, Figure 3 As shown in (a) thereof, the other function options can include but are not limited to options such as templates, music, clips, and sharing.

[0117] The following combines Figure 5 to introduce the hardware structure of the terminal device.

[0118] Figure 5 Fig. shows a schematic diagram of the hardware structure of the terminal device provided by an embodiment of the present application.

[0119] Taking the terminal device as the mobile phone 100 as an example. As Figure 5 shown, the mobile phone 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, and a bone conduction sensor 180M, etc.

[0120] The processor 110 may include one or more processing units. For example, the processor 110 may include a controller, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, the controller may be the nerve center and command center of the mobile phone 100. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions. A memory may also be provided in the processor 110 for storing instructions and data.

[0121] In the embodiment of the present application, the processor 110 is configured to: in response to the user's one-click video compilation function, analyze multiple selected material videos by the user, and respectively allocate available analysis durations for each material video according to parameters such as the chip analysis speed (chipestSpeed) and video decoding speed (decodeSpeed). Then, according to the proportion of each material video content that can be analyzed within the available analysis duration, a dynamic analysis strategy is set for each material video.

[0122] The embodiment of the present application provides 4 analysis strategies: a full-scale analysis strategy, a dense key segment analysis strategy, a sparse key segment analysis strategy, and a simple I-frame-based analysis strategy. Among them, the full-scale analysis strategy: includes every frame of the video in the analysis content, including all highlight segments, but relatively more time and resources are consumed. The simple I-frame-based analysis strategy: means that when the allocated analysis duration can only meet the time for analyzing three pictures, analyze three pictures at different positions in the video, and take the one with the highest score among these three pictures as the starting point of the highlight segment result, and return a video segment of a certain duration as the result. The dense key segment analysis strategy: is an algorithm strategy that calculates the number of analysis segments and the analysis segment duration of the material video to obtain a reasonable interval duration between adjacent segments, and then evenly distributes each analysis segment to make the analysis segments cover the highlight segments at different positions as much as possible. The sparse key segment analysis strategy: is an algorithm strategy with the same number of analyzed video segments and segment duration as the dense strategy. The difference is that the sparse strategy is a strategy algorithm that unevenly distributes the interval duration between adjacent segments. The analysis segments of this algorithm are mainly distributed in the first and middle segments of the video, and are appropriately allocated in the tail segment. These 4 analysis strategies will be described in detail in the following embodiments and will not be elaborated here.

[0123] For each source video: If the entire source video can be analyzed within the available analysis duration, a full-scale analysis strategy is adopted. If the available analysis duration is less than or equal to the analysis time of three key frames, a simple I-frame-based analysis strategy is adopted. If the available analysis duration is not sufficient to analyze the entire source video but is sufficient to analyze more than 60% of the content of the source video, an intensive key segment analysis strategy can be adopted. If the available analysis duration is greater than the analysis time of three key frames but is not sufficient to analyze more than 60% of the content of the source video, a sparse key segment analysis strategy can be adopted. After determining the analysis strategy corresponding to each video, the image signal processor scores the video frames to screen out the high-scoring highlight segments.

[0124] The wireless communication function of the mobile phone 100 can be implemented by the antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0125] The mobile phone 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and is used for graphics rendering, such as rendering Figures 1 - 4 the schematic diagram of the operation interface as shown, etc.

[0126] The display screen 194 is used to display the operation interface of the screen mirroring APP, the screen mirroring image, the screen mirroring video, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), and a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 100 may include 1 or N display screens 194, and N is a positive integer greater than 1.

[0127] The mobile phone 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0128] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the algorithms for image noise, brightness, and skin color. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0129] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0130] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the mobile phone 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0131] The video codec is used to compress or decompress digital videos. The mobile phone 100 can support one or more video codecs. In this way, the mobile phone 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0132] The NPU is a neural-network (NN) computing processor. By emulating the structure of a biological neural network, such as emulating the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn by itself. Through the NPU, applications such as the intelligent cognition of the mobile phone 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0133] The external memory interface 120 can be used to connect an external memory card to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function.

[0134] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs (APPs) required for at least one function (such as a camera APP, a gallery APP, and third-party video editing software, etc.). The data storage area can store the data created during the use of the mobile phone 100 (such as photos or videos taken, mobile phone screenshots, mobile phone screen recording content, images downloaded from other devices, and video sets generated using the one-click video creation function, etc.). In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, and a universal flash storage (UFS), etc.

[0135] The terminal device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone jack 170D, and the application processor, etc.

[0136] For example, after generating a video set using the one-click video creation function, the audio module 170 decodes the audio signal of the video set, and then the speaker 170A, also called the "loudspeaker", converts the audio electrical signal into a sound signal. In this way, the user can hear the background sound synchronized with the video in the highlight segment and the added video background music, etc.

[0137] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0138] The following will be combined with Figure 6 to introduce the software structure of the terminal device.

[0139] Figure 6 It is a schematic diagram of the software structure of the terminal device provided by the embodiments of the present application.

[0140] As Figure 6As shown in the figure, the terminal device can adopt a hierarchical architecture, dividing the software into several layers, each layer having a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software layers of the software structure are sequentially divided from top to bottom into: application (APP) layer, media middle platform framework layer, application framework (FWK) layer, and hardware abstraction layer (HAL).

[0141] The APP layer, simply referred to as the application layer, can include a series of application program packages, such as cameras, galleries, third-party video editing software, calendars, maps, and navigation. When these application program packages are run, they can access various service modules provided by the media middle platform framework layer and the application framework layer through the application programming interface (API), and execute corresponding intelligent services.

[0142] In some embodiments, the camera is used to take photos, videos, slow-motion images, panoramic images, etc. in response to user operations. After these images are taken by the camera, or after the user triggers a mobile phone screenshot, or after the user triggers a mobile phone screen recording, or after the terminal device downloads images from other devices, the terminal device can save these images in the gallery, so that the user can perform video editing operations on the images in the gallery, such as one-click video creation operations.

[0143] In the embodiments of this application, the gallery is sequentially divided from top to bottom into: business layer, application function layer, and basic function layer.

[0144] Among them, the business layer, also known as the video editing business layer, provides multiple services such as multi-shot video automatic video creation, one-shot multi-gain AI music shorts, one-click video creation, and wonderful moments. These services are presented in the form of controls in the user interface (UI) of the gallery. By operating a certain control, the user can trigger the camera to perform corresponding video processing actions. For example, after the user selects one or more material videos and clicks the one-click video creation control in the gallery, the gallery will automatically analyze and extract the highlight segments from the material videos through algorithms, and then combine the highlight segments into a clipped video.

[0145] The application function layer includes an automatic editing framework. Each service in the service layer can call the automatic editing framework to provide automatic editing services for pictures and videos. Exemplarily, the automatic editing framework may include function modules such as segment optimization, storyline organization, layout splicing, and special effect beautification. The segment optimization is used to call the highlight segment analysis interface and the policy monitoring interface to extract highlight segments from the source video. The storyline organization is used to sequentially splice multiple source videos in the form of a storyline based on the content of the source video. The layout splicing is used to adjust the interface layout of the source video. The special effect beautification is used to adjust the beautification effect of the video, such as adjusting the picture brightness and beautifying the human face, etc.

[0146] The basic function layer is used to perform basic function processing on the edited video segments after the automatic editing framework edits multiple source videos. Exemplarily, the basic function layer may include basic function modules such as video splicing, synthesis and saving, video effect rendering, and audio effect processing. Among them, the video splicing is used to splice the extracted multiple highlight segments. The synthesis and saving is used to store the video set obtained after splicing. The video effect rendering is used to add video effects to the video set obtained after splicing, such as adding style filters and themes to the video. The audio effect processing is used to add sound effects to the video set obtained after splicing, such as adding background music, etc.

[0147] The media middle platform framework layer is a software layer set between the application layer and the application framework. The media middle platform framework layer may include an analysis performance query interface, a highlight segment analysis interface, a policy monitoring interface, a pipeline interface, a theme summary interface, and an initialization interface. Among them, the analysis performance query interface is used to calculate the total duration of all source videos according to the video analysis speed. The highlight segment analysis interface is used to call the policy monitoring interface to extract highlight segments. The policy monitoring interface is used to configure the available analysis duration for each source video according to the total duration of all source videos, and set the analysis policy for each source video dynamically based on parameters such as the available analysis duration and the duration of each video. The pipeline interface is used to reduce the resolution of the video file according to the file descriptor and the analysis policy issued by the policy monitoring interface, forward the data address of the video file with reduced resolution to the hardware abstraction layer through the application framework layer, and then report the analysis result of the highlight segment returned by the hardware abstraction layer to the policy monitoring interface. The theme summary interface is used to call the underlying algorithm to analyze the picture content of the highlight segment to determine the theme corresponding to the content of the highlight segment. The initialization interface is used to initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, and image super-resolution algorithm, etc. in the HAL layer.

[0148] It should be noted that this application is described by taking the one - click video generation function provided by the picture library as an example, which does not limit the embodiments of this application. In actual implementation, third - party video editing software can adopt the video segment processing method provided by the embodiments of this application to synthesize multiple pictures and videos selected by the user into a video set at one click.

[0149] The FWK layer, simply referred to as the framework layer, can be used to support the operation of each module in the media middle - platform framework layer. For example, the framework layer can include a one - click video generation interface, a parameter management interface, an image data transmission interface, a theme analysis interface, a performance analysis interface, etc.

[0150] The HAL layer is an encapsulation of the Linux kernel driver, providing interfaces upward. It hides the hardware interface details of a specific platform, provides a virtual hardware platform for the operating system, making it hardware - independent and portable across multiple platforms. For example, the hardware abstraction layer can include a highlight segment algorithm, a face detection algorithm, a video acceleration algorithm, and an image super - resolution algorithm. Among them, the highlight segment algorithm is an image - processing algorithm provided for the image signal processor. This algorithm can score each frame of an image based on the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate value of each frame. The scoring result can be used as a basis for evaluating whether a frame of an image is a highlight segment. For example, when the scoring result of a frame of an image is greater than or equal to 60 points, this frame of an image can be used as a highlight segment.

[0151] It should be noted that Figure 6 The layers shown in the software structure and the components included in each layer do not constitute specific limitations on the terminal device. In other embodiments, the terminal device may include more layers than shown, such as a system library (FWK LIB) layer and a kernel layer. Each layer may include more or fewer components than shown. In addition, the above - mentioned various functional modules may also be combined into one functional module, and each layer may also be combined into one layer. For example, highlight segment analysis may include policy monitoring. Another example is that the media middle - platform framework layer may be set in the application framework layer.

[0152] It can be understood that in order to implement the video segment processing method in the embodiments of this application, the terminal device includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application in combination with the embodiments.

[0153] The above embodiments introduce that when a user uses the function of generating a video with one click, the user can select multiple pictures, or multiple videos, or multiple pictures and videos as the materials for extracting highlight segments. In the embodiments of the present application, the terminal device can set an upper limit on the number of materials supported by the function of generating a video with one click. For example, it supports up to 30 materials at most. Generally, a video consists of video frames of several seconds, and the number of video frames per second is greater than or equal to 24 frames. When the user selects multiple videos or selects a video with a relatively long length, if all the video frames of the video are analyzed, it may take a long time. In this case, the video segment processing method provided by the embodiments of the present application can be adopted to set the corresponding analysis strategy.

[0154] Next, taking each module in the software structure diagram as shown below Figure 6 as an example, the video segment processing method provided by the embodiments of the present application will be described by way of example.

[0155] Figure 7 The following is a schematic diagram of the overall process of the video segment processing method provided by the embodiments of the present application. This method can be applied to a Figures 1 - 4 scene of generating a video with one click as shown below. As Figure 7 shown, this method may include the following S01-S38.

[0156] S01, the one-click video generation module in the service layer receives the operation of enabling the one-click video generation function input by the user. For example, this operation may specifically be a click operation on the "One-Click Blockbuster" card as shown in (c) below Figure 1 .

[0157] S02, the one-click video generation module in the service layer loads and displays candidate pictures and candidate videos.

[0158] S03, the one-click video generation module in the service layer receives the operation of the user selecting multiple pictures and videos, and receives the operation of the user inputting to determine to execute the one-click video generation function. For example, the operation of the user selecting multiple pictures and videos may be a click operation on photos and videos as shown in (f) below Figure 1 , and the operation of the user inputting to determine to execute the one-click video generation function may be a click operation on the checkmark option as shown in (a) below Figure 2 .

[0159] S04, the one-click video generation module in the service layer calls the initialization interface of the media middle platform framework layer through the application function layer (such as the segment optimization module in the application function layer) to initialize each algorithm in the HAL layer.

[0160] S05, the initialization interface of the media middle platform framework layer sequentially issues initialization parameters to each algorithm interface in the HAL layer through the channel interface of the media middle platform framework layer and the service interface of the FWK layer.

[0161] Among them, multiple service interfaces are set in the FWK layer, and each service interface in the FWK layer plays a role in data transparent transmission between the algorithm interface and the channel interface. The HAL layer may include multiple interfaces such as a high-light segment algorithm, a face detection algorithm, a video acceleration algorithm, and an image super-resolution algorithm. One service interface in the FWK layer corresponds to one algorithm interface in the HAL layer. In addition, the initialization parameters corresponding to different interfaces may be different, so the initialization parameters sent to the service interface through each service interface may be different.

[0162] S06. Initialize each algorithm interface in the HAL layer according to the initialization parameters.

[0163] S07. Each algorithm interface in the HAL layer returns an initialization success message to the channel interface of the media middleware framework layer through the service interface of the FWK layer.

[0164] S08. The channel interface of the media middleware framework layer calls the analysis speed interface of the HAL layer through the service interface of the FWK layer (such as the performance analysis interface) to obtain the chip analysis speed.

[0165] S09. The analysis speed interface of the HAL layer returns the chip analysis speed to the channel interface of the media middleware framework layer through the service interface of the FWK layer (such as the performance analysis interface).

[0166] S10. The channel interface of the media middleware framework layer returns the performance parameters of various algorithms, such as the chip analysis speed, to the initialization interface of the media middleware framework layer.

[0167] Among them, the chip analysis speed represents the processing speed of the image signal processor for images.

[0168] In some embodiments, the chip analysis speed can be a multiple.

[0169] Exemplarily, take a video with a length of 60 seconds as an example. If the multiple is 1, the duration consumed to process this video is 60 seconds. If the multiple is 2, the duration consumed to process this video is 30 seconds. If the multiple is 3, the duration consumed to process this video is 20 seconds. If the multiple is 4, the duration consumed to process this video is 15 seconds. If the multiple is 5, the duration consumed to process this video is 12 seconds. Therefore, the greater the multiple, the faster the processing speed of the image. It should be understood that since the performances of different image signal processors are different, the chip analysis speeds corresponding to different image signal processors may be different. For a terminal device put into the market, the image signal processor is fixed, so the chip analysis speed corresponding to this image signal processor is also fixed.

[0170] S11. The initialization interface of the media middle platform framework layer returns an initialization success message to the one-click video compilation module in the service layer through the application function layer. Among them, the initialization success message carries the performance parameters of various algorithms, such as the chip analysis speed.

[0171] After obtaining the chip analysis speed, for each video selected by the user, the following S12 - S15 can be executed.

[0172] S12. The one-click video compilation module in the service layer sends a query message to the analysis performance query interface of the media middle platform framework layer through the application function layer. The query message includes the file descriptor fd1 of video 1 selected by the user and the chip analysis speed.

[0173] S13. The analysis performance query interface of the media middle platform framework layer obtains the analysis speed of video 1 according to the file descriptor fd1 of video 1. Among them, the file descriptor fd1 of video 1 can be used to uniquely identify the video.

[0174] Exemplarily, the terminal device supports a maximum resolution of 1080p@30fps, with aspect ratios of 16:9 and 9:16.

[0175] S14. The analysis performance query interface of the media middle platform framework layer calculates the analysis speed of video 1 according to the resolution of video 1, the frame rate of video 1, and the chip analysis speed.

[0176] The analysis speed of a single video ultimately depends on the video decoding speed and the chip analysis speed. Among them, the video decoding speed of video 1 can be calculated according to the resolution and frame rate of video 1. It can be understood that since the resolutions or frame rates of each video may be different, the video decoding speeds of each video may be different.

[0177] S15. The analysis performance query interface of the media middle platform framework layer returns the analysis speed of video 1 to the one-click video compilation module in the service layer through the application function layer.

[0178] After the one-click video compilation module in the service layer obtains the analysis speed of video 1, it can continue to execute S12 - S15 to obtain the analysis speed of the next video until the analysis speeds of all videos are obtained.

[0179] S16. The one-click video compilation module in the service layer calculates the target total analysis duration, the maximum duration of the highlight segment, the minimum duration of the highlight segment, and the minimum interval of the highlight segment according to the analysis speeds of all videos returned by the analysis performance query interface of the media middle platform framework layer.

[0180] Among them, the total target analysis duration represents the expected total duration for completing the analysis of all videos. The maximum duration of a highlight segment represents the expected maximum duration of a highlight segment. The minimum duration of a highlight segment represents the expected minimum duration of a highlight segment. The minimum interval of highlight segments represents the minimum duration interval between highlight segments of the same sub-shot.

[0181] For example, assume that the video duration L1 of Video 1 is 120 seconds, the video duration L2 of Video 2 is 40 seconds, and the video duration L3 of Video 3 is 15 seconds; the analysis speed S1 of Video 1 is 5 times speed, the analysis speed S2 of Video 2 is 4 times speed, and the analysis speed S3 of Video 3 is 3 times speed.

[0182] First, combined with Equation 2-2 in the above text, it can be known that:

[0183] The expected analysis duration of Video 1 is min(L1 / S1, 20s), that is, min(120 / 5, 20s) = 20s.

[0184] The expected analysis duration of Video 2 is min(L2 / S2, 20s), that is, min(40 / 4, 20s) = 10s.

[0185] The expected analysis duration of Video 3 is min(L3 / S3, 20s), that is, min(15 / 3, 20s) = 3s.

[0186] Then, combined with Equation 2-1 in the above text, it can be known that: the sum of the expected analysis durations of all videos is 20s + 10s + 3s = 33s.

[0187] In S17, the one-click video creation module in the business layer sends the file descriptor fd of all materials to be analyzed and related parameters to the image highlight segment analysis interface of the media middle platform framework layer through the application function layer. The related parameters at least include the total target analysis duration, the maximum duration of highlight segments, the minimum duration of highlight segments, and the minimum interval of highlight segments. Then, the image highlight segment of the media middle platform framework layer calls the policy monitoring module of the media middle platform framework layer.

[0188] Combined with the description of the above embodiments, the multiple materials selected by the user can only include multiple pictures, only include multiple videos, or also include multiple pictures and videos.

[0189] If only multiple pictures are included, then the terminal device will only execute the following S18-S22 and does not need to execute the following S23-S29.

[0190] If only multiple videos are included, then the terminal device will only execute the following S23-S29 and does not need to execute the following S18-S22.

[0191] If there are multiple pictures and videos, the terminal device will execute the following S18 - S29.

[0192] Take the example of including both pictures and videos as shown below. Since the computational load of pictures is small and the time consumed is less, the terminal device usually first analyzes the highlight segments of pictures one by one. After completing the analysis of the highlight segments of all pictures, it then analyzes the highlight segments of videos one by one. Figure 7 Take the example of including both pictures and videos as shown below. Since the computational load of pictures is small and the time consumed is less, the terminal device usually first analyzes the highlight segments of pictures one by one. After completing the analysis of the highlight segments of all pictures, it then analyzes the highlight segments of videos one by one.

[0193] The following is an example to illustrate the process of analyzing the highlight segments of Picture 1 in combination with S18 - S22.

[0194] S18, The policy monitoring module in the media middle - tier framework layer sends an indication message to the channel interface in the media middle - tier framework layer. The indication message includes the file descriptor of Picture 1.

[0195] S19, The channel interface in the media middle - tier framework layer decodes, reduces the resolution, and converts the format of Picture 1 according to the file descriptor of Picture 1, and stores the processed Picture 1.

[0196] S20, The channel interface in the media middle - tier framework layer sends the frame data address of Picture 1 to the highlight segment algorithm interface in the HAL layer through the service interface (such as the one - click video creation interface) in the FWK layer.

[0197] S21, The highlight segment algorithm interface in the HAL layer obtains Picture 1 according to the frame data address of Picture 1, and then analyzes Picture 1 based on a preset highlight segment algorithm to obtain an analysis result.

[0198] Exemplarily, the highlight segment algorithm interface can score Picture 1 according to the image color, image texture features, image quality, and edge change rate value of Picture 1 to obtain a scoring result. For example, when the scoring result of Picture 1 is greater than or equal to 60 points, Picture 1 can be regarded as a highlight segment.

[0199] S22, The highlight segment algorithm interface in the HAL layer returns the analysis result of Picture 1 to the policy monitoring module in the media middle - tier framework layer through the service interface in the FWK layer and the channel interface in the media middle - tier framework layer in sequence.

[0200] After the policy monitoring module in the media middle - tier framework layer obtains the analysis result of Picture 1, if there are other pictures, such as Picture 2, the terminal device can continue to execute the above S18 - S22 to obtain the analysis results of other pictures. After obtaining the analysis results of all pictures, the terminal device can use the following S23 - S29 to obtain the analysis results of each video.

[0201] S23. The policy monitoring module in the media middle platform framework layer allocates the available analysis duration for each video according to the total target analysis duration and the analysis duration of each video.

[0202] Among them, the total target analysis duration represents the expected total duration to complete the analysis of all videos selected by the user. The analysis duration of a video is the duration required for the image chip to complete the analysis of all frames of the video.

[0203] Exemplarily, the policy monitoring module can allocate the available analysis duration for each video according to the ratio of the analysis duration of each video to the total target analysis duration. For example, the longer the analysis duration of a video, the longer the available analysis duration allocated for the video; the shorter the analysis duration of a video, the shorter the available analysis duration allocated for the video. Of course, the policy monitoring module can also adopt other methods to allocate the available analysis duration for each video, which is not limited in the embodiments of this application.

[0204] S24. The policy monitoring module in the media middle platform framework layer sets analysis policies for each video respectively according to the available analysis duration of each video and the analysis duration of each video. For example, full-analysis policy, intensive key segment analysis policy, sparse key segment analysis policy, and simple I-frame-based analysis policy.

[0205] Specifically, the policy monitoring module first sets an analysis policy for Video 1, and then analyzes the highlight segments of Video 1 successively using S25-S29 described below. Then it sets an analysis policy for Video 2, and analyzes the highlight segments of Video 2 successively using S25-S29 described below. And so on, until the analysis of the highlight segments of all videos is completed.

[0206] In some embodiments, the order of setting analysis policies for each video can be determined according to the shooting time of the video. For example, analyze the most recently shot video first, and then analyze the earlier shot video. In other embodiments, the order of setting analysis policies for each video can be determined according to the available analysis duration of the video. For example, analyze the video with a shorter available analysis duration first, and then analyze the video with a longer available analysis duration.

[0207] The following uses S25-S29 to give an example of the process of analyzing the highlight segments of Video 1.

[0208] S25. The policy monitoring module in the media middle platform framework layer sends the file descriptor fd1 of Video 1 and the analysis policy set for Video 1 to the channel interface in the media middle platform framework layer.

[0209] S26. The channel interface in the media middle platform framework layer performs processing operations such as decoding, downscaling the resolution, and converting the format on Video 1 according to the file descriptor of Video 1, and stores the processed Video 1.

[0210] S27. The channel interface of the media middle platform framework layer sends the frame data address of Video 1 to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one-click video generation interface).

[0211] S28. The highlight segment algorithm interface of the HAL layer obtains Video 1 based on the frame data address of Video 1, and then analyzes Video 1 based on a preset highlight segment algorithm to obtain an analysis result.

[0212] In some embodiments, since a video consists of multiple video frames, the channel interface can send the data address of each video frame to the highlight segment algorithm interface in the order of the frames. Correspondingly, the highlight segment algorithm interface can analyze the image features of the data frames corresponding to each frame data address in the order of receiving the frame data addresses to obtain the scoring results of each frame.

[0213] S29. The highlight segment algorithm interface of the HAL layer returns the analysis result of Video 1, such as the position of the highlight segment, to the policy monitoring module of the media middle platform framework layer through the service interface of the FWK layer and the channel interface of the media middle platform framework layer in sequence.

[0214] After the policy monitoring module of the media middle platform framework layer obtains the analysis result of Video 1, if there are other videos, such as Video 2, the terminal device can continue to execute the above S25 - S29 to obtain the analysis results of other videos. After obtaining the analysis results of all videos, the terminal device can report all the picture and video analysis results by using the following S30.

[0215] S30. The policy monitoring module of the media middle platform framework layer reports all the picture and video analysis results, such as the positions of all highlight segments, to the application function layer through the image highlight segment analysis interface of the media middle platform framework layer.

[0216] S31. The application function layer clips and filters the user-selected materials according to all the picture and video analysis results to obtain all the highlight segments.

[0217] S32. The application function layer sends a request message for obtaining the themes corresponding to all the highlight segments to the theme summary interface of the media middle platform framework layer.

[0218] S33. The theme summary interface of the media middle platform framework layer calls the service interface of the FWK layer and the algorithm module of the HAL layer (such as the theme algorithm) to obtain the themes corresponding to all the highlight segments.

[0219] In some embodiments, the terminal device may be provided with multiple theme templates (style templates). The theme algorithm may recommend a theme template that matches the scene of the highlight segment, such as people, scenery, food, children, pets, sports or travel.

[0220] S34, the theme summary interface of the media middle platform framework layer returns the recommended theme to the application function layer.

[0221] S35, the application function layer distributes the theme obtained from S33 and all the highlight segments obtained from S31 to the basic capability layer.

[0222] S36, the basic capability layer generates a target video based on the theme and all the highlight segments. That is, the target video is a video set that is generated based on all the screened highlight segments and conforms to the recommended theme.

[0223] S37, the basic capability layer sends an indication message for playing the target video to the service layer.

[0224] S38, the service layer plays the target video.

[0225] S01 - S38 of the above embodiments introduce the overall process of the video segment processing method provided by the embodiments of the present application. After allocating the available analysis duration for each video according to the desired time, the policy monitoring module can set the analysis policy for each video one by one. The following combines Figure 8 An example is given to illustrate the specific process of the policy monitoring module setting the analysis policy for each video in turn according to the available analysis duration of each video and the analysis time consumption of each video.

[0226] Figure 8 It is a schematic flowchart of the process of setting the analysis policy for each video provided by the embodiments of the present application. As Figure 8 shown, the above S24 can be specifically implemented through the following S241 - S243.

[0227] S241, the policy monitoring module determines whether it is possible to analyze the entire i-th video within the available analysis duration allocated to the i-th video. Where the i-th video is any one of the multiple videos selected by the user.

[0228] In some embodiments, the policy monitoring module may compare the available analysis duration allocated to the i-th video with the analysis time consumption of the i-th video. If the available analysis duration allocated to the i-th video is greater than or equal to the analysis time consumption of the i-th video, then it is possible to analyze the entire i-th video within the available analysis duration allocated to the i-th video. At this time, a full-scale analysis policy can be adopted to analyze all the video frames of the i-th video. Otherwise, continue to execute the following S242.

[0229] S242, the policy monitoring module determines whether the available analysis duration allocated for the i-th video is less than or equal to the analysis time of three key frames.

[0230] Generally, the analysis time of one key frame is fixed. For example, the analysis time of one key frame is 50 milliseconds or 60 milliseconds.

[0231] When the available analysis duration allocated for the i-th video is less than or equal to the analysis time of three key frames, it means that the available analysis duration allocated for the i-th video is very short, and at this time, it is not sufficient to analyze a segment of the i-th video. At this time, a simple analysis strategy based on I-frames can be adopted to analyze three pictures at different positions of the i-th video, and the one with the highest score among these three pictures is used as the starting point of the highlight segment result, and a video segment of a certain duration is returned as the result.

[0232] If the available analysis duration allocated for the i-th video is greater than the analysis time of three key frames and less than the analysis time-consuming of the i-th video, at this time, the key segments of the i-th video can be analyzed. The embodiments of the present application provide two methods for analyzing key segments: intensive key segment analysis strategy and sparse key segment analysis strategy. Among them, the intensive key segment analysis strategy is characterized in that multiple key segments are evenly distributed throughout the video; the sparse key segment analysis strategy is characterized in that multiple key segments are mainly concentrated in the first and middle segments of the video, and appropriately distributed in the tail segment.

[0233] As an optional implementation manner, the terminal device can randomly select one algorithm from the two algorithms, but this random manner may have the problem of inaccurate extraction of highlight segments.

[0234] As another optional implementation manner, the terminal device can combine the characteristics of the two algorithms to select a suitable algorithm for the i-th video from the two algorithms, that is, the policy monitoring module can continue to execute the following S243.

[0235] S243, the policy monitoring module determines whether the available analysis duration allocated for the i-th video can analyze more than 60% of the content.

[0236] In some embodiments, the policy monitoring module can calculate the ratio of the available analysis duration allocated for the i-th video to the analysis time-consuming of the i-th video. Among them, the analysis time-consuming of the i-th video is equal to the total duration of the i-th video divided by the analysis speed of the i-th video, and the analysis speed of the i-th video is determined according to the resolution of the i-th video, the frame rate of the i-th video, and the chip analysis speed.

[0237] Assume that the available analysis duration allocated for the i-th video is represented by T 预期分析 and the analysis time-consuming of the i-th video is represented by T 分析耗时It is indicated that the chip analysis speed is represented by T 芯片分析速度 The available analysis duration T allocated to the i-th video 预期分析 And the analysis time consumption T of the i-th video 分析耗时 The ratio is represented by R, and the following relationship exists:

[0238]

[0239] The policy monitoring module can determine whether R is greater than the preset ratio R1. For example, R1 = 0.6.

[0240] If the ratio R is greater than R1 (0.6), then within the available analysis duration allocated to the i-th video, the content of the i-th video that can be analyzed is greater than 60%. At this time, the intensive key segment analysis strategy can be adopted, that is, by calculating the number of analysis segments and the duration of the analysis segments of the source video, the reasonable interval duration between adjacent segments is obtained, and then each analysis segment is evenly distributed so that the analysis segments can cover the highlight segments located at different positions as much as possible.

[0241] If the ratio R is less than or equal to 0.6, then within the available analysis duration allocated to the i-th video, the content of the i-th video that can be analyzed is less than or equal to 60%. At this time, the sparse key segment analysis strategy can be adopted. The difference from the intensive key segment analysis strategy is that the sparse key segment analysis strategy makes unequal distribution of the interval duration between adjacent segments, and the analysis segments of this algorithm are mainly distributed in the first and middle segments of the video, and are appropriately distributed in the tail segment. It should be noted that when the user is interested in certain content, the video recording will start, and when the user is not interested, the video recording will stop. Therefore, the probability that the first and middle segments of a video are the content that the user is interested in is relatively high. The sparse key segment analysis strategy is beneficial to more accurately and quickly extract the highlight segments by setting a smaller segment interval in the first and middle segments of the video.

[0242] It should be noted that the above embodiments are illustrated by taking R1 = 0.6 as an example, and it does not limit the embodiments of the present application. In actual implementation, the size of the preset ratio can be adjusted according to the usage requirements. For example, the preset ratio can also be 0.4, 0.5, 0.7, etc.

[0243] Exemplarily, Figure 9 Shows a schematic diagram of the time consumption of four analysis strategies provided by the embodiments of the present application. As Figure 9As shown in the figure, the time consumption of the full - scale analysis strategy, the intensive key - segment analysis strategy, the sparse key - segment analysis strategy, and the simple I - frame - based analysis strategy decreases in turn. Since the full - scale analysis strategy analyzes all frames of the video, it takes the longest time. The simple I - frame - based analysis strategy analyzes at most three video frames, so it takes the shortest time. The time consumption of the intensive key - segment analysis strategy and the sparse key - segment analysis strategy is between the full - scale analysis strategy and the simple I - frame - based analysis strategy. In addition, compared with the intensive key - segment analysis strategy, the sparse key - segment analysis strategy focuses on analyzing the first and middle segments of the video, so the time consumption may be shorter.

[0244] The following will elaborate on the full - scale analysis strategy, the intensive key - segment analysis strategy, the sparse key - segment analysis strategy, and the simple I - frame - based analysis strategy from four embodiments respectively.

[0245] Embodiment 1

[0246] When the available analysis duration allocated by the strategy monitoring module for a certain video meets the requirement that each frame of the video is included in the analysis content, the full - scale analysis strategy can be selected.

[0247] As Figure 10 shown in the figure, in the full - scale analysis strategy, all frame pictures of the video file are analysis segments. The strategy monitoring module will send all frames of the video file, frame by frame from start to end, to the highlight segment algorithm for highlight segment analysis, so that the highlight segment algorithm can screen out the highlight segments from these frames.

[0248] After setting the corresponding analysis strategy for each of the videos selected for the user, the strategy monitoring module can extract the highlight segments from each video in turn according to the analysis strategy corresponding to each video. Suppose the analysis strategy set for Video 1 is the full - scale analysis strategy. As Figure 11 shown in the figure, the specific implementation method of the full - scale analysis strategy is as described in A1 - A9.

[0249] A1. The strategy monitoring module sends the file descriptor fd1 of Video 1 and the analysis strategy (full - scale analysis strategy) set for Video 1 to the channel interface.

[0250] A2. The channel interface decodes all video frames of Video 1 according to the file descriptor fd1 of Video 1 and the analysis strategy (full - scale analysis strategy) set for Video 1.

[0251] It can be understood that by sending the analysis strategy (full - scale analysis strategy) set for Video 1, the strategy monitoring module enables the channel interface to determine that the video frames to be decoded are all video frames of Video 1 according to this strategy.

[0252] A3. The channel interface converts all decoded video frames from the first format to the second format.

[0253] Exemplarily, the first format is the nv12 format and the second format is the i420 format.

[0254] A4, the channel interface reduces the resolution of all the video frames after format change from the first resolution to the second resolution.

[0255] Exemplarily, the first resolution is 1080p and the second resolution is 480p.

[0256] In some embodiments, the second resolution corresponding to different analysis strategies may be different. Therefore, in A2 above, the policy monitoring module issues the analysis strategy (full - volume analysis strategy) set for Video 1, enabling the channel interface to determine the second resolution corresponding to the full - volume analysis strategy.

[0257] A5, the channel interface converts all the video frames with reduced resolution from the second format to the first format.

[0258] A6, the channel interface stores all the video frames that are converted back to the first format in the memory (buffer).

[0259] It should be noted that in order to improve the processing speed of the highlight segment algorithm interface for video frames, before performing highlight analysis on video frames, the channel interface can perform resolution reduction processing on the video frames. Additionally, since the resolution reduction algorithm only supports the second format, all the video frames need to be first converted from the first format to the second format, then the resolution of the video frames is reduced, and then the video frames with reduced resolution are converted back to the original first format and stored in the memory.

[0260] A7, the channel interface sends the frame data address of Video 1 to the highlight segment algorithm interface through the one - key video compilation interface.

[0261] A8, the highlight segment algorithm interface obtains the frame data of Video 1 according to the frame data address of Video 1, and then analyzes Video 1 based on the preset highlight segment algorithm to obtain the highlight segment position of Video 1. Among them, the highlight segment position can include the start address and the end address of the highlight segment.

[0262] In some embodiments, since Video 1 consists of multiple video frames, the channel interface can send the data addresses of each video frame of Video 1 to the highlight segment algorithm interface in sequence from start to end. Correspondingly, the highlight segment algorithm interface can analyze the image features of the data frames corresponding to each frame data address according to the receiving order of the frame data addresses to obtain the scoring results of each frame.

[0263] Exemplarily, take Video 1 including video frame a, video frame b, and video frame c as an example. The channel interface can first send the data address 0001 of video frame a to the highlight segment algorithm interface through the one-click video production interface. The highlight segment algorithm interface obtains video frame a from the data address 0001 and scores video frame a. Then, the channel interface sends the data address 0002 of video frame b to the highlight segment algorithm interface through the one-click video production interface. The highlight segment algorithm interface obtains video frame b from the data address 0002 and scores video frame b. Then, the channel interface sends the data address 0003 of video frame c to the highlight segment algorithm interface through the one-click video production interface. The highlight segment algorithm interface obtains video frame c from the data address 0003 and scores video frame c. After obtaining the scoring results of all frames, the highlight segment algorithm interface calculates the position of the highlight segment of Video 1 based on the scoring results of all frames. The position of the highlight segment includes the start address and end address of the highlight segment. For example, assume that the scoring results of video frame a and video frame b are both greater than 60 points, and the scoring result of video frame c is less than 60 points. Then the start address of the highlight segment is 0001, and the end address is 0002.

[0264] A9. The highlight segment algorithm interface sequentially returns the analysis result of Video 1, such as the position of the highlight segment, to the policy monitoring module through the one-click video production interface and the channel interface.

[0265] After the policy monitoring module obtains the analysis result of Video 1, if there is still a next video, such as Video 2, then the terminal device can extract the highlight segment from the next video according to the analysis policy corresponding to the next video. The analysis policy corresponding to the next video can be a full-scale analysis policy or other analysis policies. After obtaining the analysis results of all videos, the terminal device can report all the analysis results.

[0266] For ​ the implementation manners of S01 - S24 and S30 - S38 in ​ reference can be made to the specific descriptions of S01 - S24 and S30 - S38 in the above embodiments, which will not be elaborated here.

[0267] It can be understood that when the analysis policy set for a certain video is a full-scale analysis policy, since the terminal device will analyze each frame of the video file one by one, the final analysis result can completely include all the highlight segments of the video, but the time and resource consumption are relatively large.

[0268] Embodiment 2

[0269] In conjunction with the description of the above embodiment, users are generally more interested in content located in the front and middle sections of a video. When the available analysis time allocated for a video is less than or equal to the analysis time of three key frames, it indicates that the available analysis time allocated for the video is very short and insufficient to analyze a single segment of the video. In this case, a simple I-frame-based analysis strategy can be adopted to analyze three images located in three locations in the front and middle sections of the video. The highest-scoring of these three images is used as the starting point for the highlight segment result, and a video segment of a certain length is captured as the only highlight segment.

[0270] In some embodiments, the above three positions may be located at any three positions in the front and middle sections of the video, or may be three designated positions.

[0271] For example, ​ FIG1 shows a schematic diagram of a highlight segment when a simple analysis strategy based on I frame is adopted. Assume that the analysis strategy set for video 1 is a simple analysis strategy based on I frame. The three key frames based on the simple analysis strategy of I frame may include: ​ The first I frame (i.e., frame 1) at the start of the video shown in (a) is located at ​ The I frame at the 2nd second position shown in (b) and the I frame at the 2nd second position shown in (b) ​ The terminal device can use the highest-scoring image among the three images as the starting point of the highlight segment result, starting from the total duration T of video 1. 总 The interception time is T 高光 It should be noted that a 1-second video typically includes at least 24 frames, and the I-frame at the 2nd second position can specifically be the first frame of all frames at the 2nd second of the video. In addition, the third key frame can also be an I-frame at the 1 / 4 position of the video, or an I-frame at the 1 / 2 position of the video, etc., which is not limited in the embodiments of the present application.

[0272] It should be noted that the embodiment of the present application is described by taking three key frames as an example based on a simple analysis strategy of I frames, which does not limit the embodiment of the present application. In actual implementation, it can also be two key frames, or four key frames, etc.

[0273] After setting the corresponding analysis strategy for all videos, the strategy monitoring module can extract the highlight segments from each video in turn according to the analysis strategy for each video. Assume that the analysis strategy set for video 1 is a simple analysis strategy based on I frames. ​ As shown, the specific implementation method based on the simple analysis strategy of I frame is described in B1-B16.

[0274] B1. The policy monitoring module sends the file descriptor fd1 of Video 1 and the analysis policy set for Video 1 (simple analysis policy based on I-frames) to the channel interface.

[0275] B2. The channel interface decodes three key frames of Video 1 according to the file descriptor fd1 of Video 1 and the analysis policy set for Video 1 (simple analysis policy based on I-frames).

[0276] By sending the analysis policy set for Video 1 (simple analysis policy based on I-frames), the policy monitoring module enables the channel interface to determine that the video frames to be decoded are the three key frames of Video 1 according to this policy.

[0277] For example, the three key frames may include the first frame at the start of Video 1, the I-frame at the 2-second position, and the I-frame at the 1 / 3 position of Video 1.

[0278] B3. The channel interface converts the three decoded key frames from the first format to the second format.

[0279] Exemplarily, the first format is the nv12 format and the second format is the i420 format.

[0280] B4. The channel interface reduces the resolution of the three key frames with the changed format from the first resolution to the second resolution.

[0281] Exemplarily, the first resolution is 1080p and the second resolution is 480p.

[0282] In some embodiments, the second resolutions corresponding to different analysis policies may be different or the same. Therefore, in B2 above, the policy monitoring module sends the analysis policy set for Video 1 (simple analysis policy based on I-frames), enabling the channel interface to determine the second resolution corresponding to the simple analysis policy based on I-frames.

[0283] B5. The channel interface converts the three key frames with reduced resolution from the second format to the first format.

[0284] To improve the processing speed of the highlight segment algorithm interface for the three key frames, before performing highlight analysis on the video frames, the channel interface can perform downsampling on the video frames. Additionally, since the downsampling algorithm only supports the second format, it is necessary to first convert the three key frames from the first format to the second format, then reduce the resolution of the video frames, and then convert the video frames with reduced resolution back to the original first format.

[0285] B6. The channel interface stores the three key frames that are converted back to the first format in the memory (buffer).

[0286] B7, the channel interface sends the data address of the first frame at the starting point of Video 1 to the highlight segment algorithm interface through the one-click video creation interface.

[0287] B8, the highlight segment algorithm interface obtains the first frame according to the data address of the first frame, then analyzes the first frame based on the preset highlight segment algorithm to obtain the score of the first frame, and then returns the score of the first frame to the policy monitoring module through the one-click video creation interface and the channel interface in sequence.

[0288] B9, the policy monitoring module determines whether the score of the first frame is greater than or equal to the preset value.

[0289] Taking the preset value of 60 points as an example. If the score of the first frame is greater than or equal to 60 points, then the first frame can be used as the starting point of the highlight segment position, and B16 is executed without executing B10 - B15. If the score of the first frame is less than 60 points, then the following B10 is executed.

[0290] B10, the policy monitoring module notifies the channel interface to send the data address of the I-frame at the 2nd second position to the highlight segment algorithm interface. The channel interface sends the data address of the I-frame at the 2nd second position to the highlight segment algorithm interface through the one-click video creation interface.

[0291] B11, the highlight segment algorithm interface obtains the I-frame at the 2nd second position according to the data address of the I-frame at the 2nd second position, then analyzes the I-frame at the 2nd second position based on the preset highlight segment algorithm to obtain the score of the I-frame at the 2nd second position, and then returns the score of the I-frame at the 2nd second position to the policy monitoring module through the one-click video creation interface and the channel interface in sequence.

[0292] B12, the policy monitoring module determines whether the score of the I-frame at the 2nd second position is greater than or equal to the preset value.

[0293] Still taking the preset value of 60 points as an example. If the score of the I-frame at the 2nd second position is greater than or equal to 60 points, then the I-frame at the 2nd second position can be used as the starting point of the highlight segment position, and B16 is executed without executing B13 - B15. If the score of the I-frame at the 2nd second position is less than 60 points, then the following B13 is executed.

[0294] B13, the policy monitoring module notifies the channel interface to send the data address of the I-frame at the 1 / 3 position of Video 1 to the highlight segment algorithm interface. The channel interface sends the data address of the I-frame at the 1 / 3 position of Video 1 to the highlight segment algorithm interface through the one-click video creation interface.

[0295] B14, the highlight segment algorithm interface obtains the I frame located at the 1 / 3 position of the video according to the data address of the I frame located at the 1 / 3 position of the video, and then analyzes the I frame located at the 1 / 3 position of the video based on the preset highlight segment algorithm to obtain the score of the I frame located at the 1 / 3 position of the video, and then returns the score of the I frame located at the 1 / 3 position of the video to the policy monitoring module through the one-key film interface and the channel interface.

[0296] B15, the policy monitoring module determines whether the score of the I frame located at the 1 / 3 position of the video is greater than or equal to a preset value.

[0297] Let's use the default value of 60 as an example. If the score of the I-frame at the 1 / 3 position in the video is greater than or equal to 60, then the I-frame at the 1 / 3 position in the video can be used as the starting point for the highlight segment, and B16 can be executed. If the score of the I-frame at the 1 / 3 position in the video is also less than 60, then the video frame with the highest score among the three key frames can be used as the starting point for the highlight segment, and B16 can be executed.

[0298] It should be noted that the above description of B9, B12 and B15 is based on the example that the preset values of the three positions are equal. In actual implementation, the three preset values may not be completely equal, and this embodiment of the application does not limit this.

[0299] B16, the strategy monitoring module uses a key frame as the starting point of the highlight segment position and selects a video segment of a certain length as the only highlight segment.

[0300] For example, ​ Two schematic diagrams for determining the duration of highlight segments are shown.

[0301] Assume that the minimum length of the highlight segment is T 短高光 Indicates that the maximum duration of the highlight segment is T 长高光 Indicates that the duration of the highlight segment is expressed in T 高光 If , then the following relationship exists:

[0302] T 高光 =min(1.2*T 短高光 ,T 长高光 ).

[0303] Among them, the min() function is used to find the minimum value. That is, take 1.2*T 短高光 and T 长高光 The smaller value of T 高光 .

[0304] like ​ As shown in (a), due to T 长高光 >1.2*T 短高光 , so T高光 = 1.2 * T 短高光 .

[0305] As ​ shown in (b) of 长高光 < 1.2 * T 短高光 , so T 高光 = T 长高光 .

[0306] After the policy monitoring module obtains the analysis result of Video 1, if there is still a next video, such as Video 2, the terminal device can extract the highlight segments from the next video according to the analysis policy corresponding to the next video. The analysis policy corresponding to the next video can be a simple I-frame based analysis policy or other analysis policies. After obtaining the analysis results of all videos, the terminal device can report all the analysis results.

[0307] For ​ the implementation manners of S01 - S24 and S30 - S38 in ​ , reference can be made to the specific descriptions of S01 - S24 and S30 - S38 in the above embodiments, which will not be elaborated here.

[0308] It can be understood that when the available analysis duration allocated for a certain video is less than or equal to the analysis time of three key frames, it means that the available analysis duration allocated for this video is very short, and at this time, it is not sufficient to analyze a segment of this video. At this time, a simple I-frame based analysis policy can be adopted to analyze three pictures located in the middle front section of this video, and the one with the highest score among these three pictures is used as the starting point of the highlight segment result, and a video segment with a certain duration is intercepted as the only highlight segment. On the one hand, since the simple I-frame based analysis policy only needs to analyze at most three pictures, compared with other several analysis policies, the simple I-frame based analysis policy can obtain the highlight segment more quickly. On the other hand, since the probability that users are interested in the content in the middle front section of the video is relatively high, using the one with the highest score among the three pictures in the middle front section as the starting point of the highlight segment result can ensure the reliability of the extracted highlight segment to a certain extent.

[0309] Embodiment III

[0310] Combined with the description of the above embodiments, the policy monitoring module can calculate the ratio R of the available analysis duration T 预期分析 allocated for the i-th video to the analysis duration T 分析耗时 of the i-th video. When the ratio R is greater than the preset ratio R1 (taking R1 = 0.6 as an example), it means that more than 60% of the video content can be analyzed, and at this time, the analysis policy set for the i-th video can be set as the intensive key segment analysis policy.

[0311] That is, the intensive key segment analysis strategy is adopted when the following relationship is satisfied:

[0312]

[0313] Among them, T 总 is the total duration of the i-th video, and T 芯片分析速度 is the chip analysis speed.

[0314] The idea of the intensive key segment analysis strategy is as follows: First, calculate three target parameters - the number N 分析 of the analysis segments of the material video, the duration T 分析 of the analysis segments, and the interval duration T 间隔 between adjacent segments. Then, according to these three target parameters, evenly distribute each analysis segment so that the analysis segments can cover the highlight segments at different positions as much as possible. Among them, the analysis segments are also called key segments and are used to extract highlight segments from them.

[0315] Exemplarily, as ​ shown, evenly distribute the analysis segments at equal intervals starting from the first frame of the i-th video file. Among them, the durations of analysis segments N1, analysis segments N2, analysis segments N3... are T 分析 , and the duration of the last analysis segment Nn is T 剩余 . Among them, T 剩余 ≤T 分析 .

[0316] In the embodiments of the present application, the above three target parameters can be calculated from five initial parameters as ​ shown. Next, the meanings of these five initial parameters will be introduced. As ​ described, the five initial parameters may include:

[0317] The minimum duration T 短高光 of the highlight segments;

[0318] The maximum duration T 长高光 of the highlight segments;

[0319] The minimum interval T 最小间隔 of the highlight segments;

[0320] The available analysis duration T 预期分析 allocated to the i-th video, which is also called the expected analysis duration;

[0321]

[0321] and the total duration T 总 of the i-th video.

[0322] Among them, the minimum duration T 短高光 of the highlight segments is the expected minimum duration configured for each highlight segment. The maximum duration T 长高光The maximum expected duration of each highlight segment is configured. The minimum interval T of the highlight segment 最小间隔 is the minimum time interval between highlight segments of the same shot. Referring to the description of the above embodiments S16-S17, these three parameters can be configured by the application layer and sent to the policy monitoring module.

[0323] The available analysis time T allocated for the i-th video 预期分析 , is the total expected duration allocated by the policy monitoring module for analyzing the i-th video according to the expected time. 预期分析 The method of obtaining can refer to the description of the above embodiment S23 and will not be repeated here.

[0324] The total length of the i-th video T 总 In the embodiment of the present application, the total duration of a video can be set to no more than 30 seconds, which can reduce the amount of data calculation of the highlight segment interface.

[0325] It should be noted that in ​ The lengths of the line segments corresponding to the various durations are only exemplary. In actual implementation, the specific values of the various durations can be adjusted according to actual needs.

[0326] The following combination ​ , an example is given to illustrate the principle of calculating 3 target parameters based on 5 initial parameters.

[0327] (1) According to the total length T of the i-th video 总 , the available analysis time T allocated to the i-th video 预期分析 , and the minimum interval T of highlight fragments 最小间隔 , determine the maximum number of analyzed fragments N 上限 , which is the upper limit of the analyzed fragment.

[0328] In one example, N 上限 It can be calculated by the following relationship:

[0329]

[0330] in, is the floor symbol.

[0331] For example, N 上限 The maximum value of is in the range of [1 to 5], and N 上限 is an integer. For example, when the terminal device pre-sets N 上限 When the maximum value of is set to 4, if the above relationship is used to calculate N 上限 =5, then N 上限 The final value is set to 4.

[0332] It should be understood that the total duration T of the i-th video 总 and the available analysis duration T allocated for the i-th video 预期分析 The difference between them is the total duration of all interval segments. Divide the total duration of all interval segments by the minimum interval T of the highlight segments 最小间隔 , and the upper limit value of the interval segments can be obtained. Since one interval segment is set between every two analysis segments, the number of analysis segments is 1 more than the number of interval segments. Thus, using T 总 、T 预期分析 and T 最小间隔 The upper limit value N of the analysis segments can be obtained 上限 .

[0333] (2) According to the maximum number N of the analysis segments 上限 , and the available analysis duration T allocated for the i-th video 预期分析 , determine the minimum duration T of the analysis segments 下限 .

[0334] In an example, T 下限 can be calculated through the following relational expression:

[0335]

[0336] It should be understood that when the available analysis duration T allocated for the i-th video 预期分析 is a fixed value, when the number of analysis segments is set to the maximum number N 上限 , the calculated duration of the analysis segments is the minimum duration T 下限 .

[0337] (3) According to the minimum duration T of the analysis segments 下限 , and the maximum duration T of the highlight segments 长高光 , determine the duration T of the analysis segments 分析 .

[0338] In one implementation, without considering the minimum duration T of the analysis segments 下限 , by default, the maximum duration T of the highlight segments 长高光 is used as the duration T of the analysis segments 分析 , and there is the following relational expression at this time:

[0339] T 分析 = T 长高光 .

[0340] In another implementation, it is necessary to consider the minimum duration T of the analysis segments 下限 , and the minimum duration T of the analysis segments 下限 and the maximum duration T of the highlight segments长高光 The largest value among them is used as the duration T of the analysis segment 分析 , and there is the following relational expression at this time:

[0341] T 分析 = max(T 长高光 , T 下限 ).

[0342] Among them, the max() function is used to obtain the maximum value.

[0343] It should be understood that the largest value among the minimum duration T of the analysis segment 下限 and the maximum duration T of the highlight segment 长高光 is used as the duration T of the analysis segment 分析 , which can appropriately increase the duration of the analysis segment, so as to cover as many positions of the highlight segments as possible, and then improve the accuracy of the finally extracted highlight segments.

[0344] (4) According to the duration T of the analysis segment 分析 , and the available analysis duration T 预期分析 assigned to the i-th video, determine the number N 分析 of the analysis segments.

[0345] In one implementation, N 分析 can be calculated through the following relational expression:

[0346]

[0347] Among them, is the floor symbol.

[0348] In another implementation, after calculating N 分析 in the above manner, there may still be a remaining analysis duration T 预期分析 in the available analysis duration T 分析 assigned to the i-th video that is less than the duration T of the analysis segment 剩余 :

[0349]

[0350] The remaining analysis duration T 剩余 may be greater than the minimum duration T of the highlight segment 短高光 . In this case, an additional analysis segment can be added based on N 分析 , and the remaining analysis duration T 剩余 is used as the duration of the last analysis segment.

[0351] Specifically, N 分析 can be calculated through the following relational expression:

[0352]

[0353] It should be understood that when the remaining analysis duration T 剩余 is greater than the minimum duration T 短高光 of the highlight segment, by adding an analysis segment, all the expected analysis duration can be effectively utilized to cover more positions where the highlight segments are located, thereby improving the accuracy of the finally extracted highlight segments.

[0354] (5) According to the number N 分析 of the analysis segments, the total duration T 总 of the i-th video, and the available analysis duration T 预期分析 allocated for the i-th video, determine the interval duration T 间隔 of the analysis segments.

[0355] In one implementation, without considering the minimum interval T 最小间隔 of the highlight segments, T 间隔 can be directly calculated through the following relational expression:

[0356]

[0357] In another implementation, it is necessary to consider the minimum interval T 最小间隔 of the highlight segments. Take the maximum value between T 间隔 calculated in the above manner and the minimum interval T 最小间隔 of the highlight segments as the interval duration T 间隔 of the analysis segments. At this time, there is the following relational expression:

[0358] T 间隔 = max(T 间隔 , T 最小间隔 ).

[0359] It should be understood that the minimum interval T 最小间隔 of the highlight segments is a minimum interval calculated by the application layer. To avoid the problem that the intervals are too small and the analysis segments are too concentrated, when T 间隔 is less than the minimum interval T 最小间隔 of the highlight segments, T 间隔 should be set to T 最小间隔 .

[0360] So far, the policy monitoring module has obtained three target parameters: the number N<s 分析 of the analysis segments of the material video, the duration T 分析 of the analysis segments, and the interval duration T 间隔 between adjacent segments. Then, the policy monitoring module can evenly distribute each analysis segment according to these three target parameters.

[0361] To facilitate understanding of the calculation principles of the above three target parameters, the minimum duration T of the analysis segment that needs to be considered in the calculation process is taken below 下限 , the remaining analysis duration T 剩余 , and the minimum interval T of the highlight segments 最小间隔 as examples. Through three specific examples, the calculation processes of the three target parameters are described.

[0362] Example 1, assume T 短高光 = 2 seconds, T 长高光 = 5 seconds, T 最小间隔 = 2 seconds, T 预期分析 = 15 seconds, T 总 = 20 seconds, then the calculation processes of the three target parameters are as follows.

[0363] 1.

[0364] 2.

[0365] 3. T 分析 = max(T 长高光 , T 下限 ) = max(5, 5) = 5 seconds.

[0366] 4.

[0367] 5. T 剩余 = T 预期分析 - N 分析 * T 分析 = 15 - 3×5 = 0 seconds.

[0368] 6. Since (T 剩余 = 0 seconds) < (T 短高光 = 2 seconds), so N 分析 = 3.

[0369] 7.

[0370] 8. T 间隔 = max(T 间隔 , T 最小间隔 ) = max(2.5, 2) = 2.5 seconds.

[0371] So far, the three target parameters have been obtained: T 分析 = 5 seconds, N 分析 = 3, T 间隔 = 2.5 seconds.

[0372] As ​ shown, the available analysis duration T allocated for the i-th video 预期分析It is divided into three segments of 5 seconds each, leaving no remaining time for analysis. 间隔 = 2.5 seconds greater than T 最小间隔 = 2 seconds, so T 间隔 The final value is 2.5 seconds. When dividing the analysis segments, the analysis segments are allocated at equal intervals starting from the first frame of the i-th video file. The duration of analysis segments N1, N2, and N3 is 5 seconds, the interval between analysis segments N1 and N2 is 2.5 seconds, and the interval between analysis segments N2 and N3 is 2.5 seconds.

[0373] Example 2, assuming T 短高光 = 2 seconds, T 长高光 =5 seconds, T 最小间隔 = 3 seconds, T 预期分析 = 12 seconds, T 总 = 20 seconds, the calculation process of the three target parameters is as follows:

[0374] 1.

[0375] 2.

[0376] 3. T 分析 =max(T 长高光 , T 下限 )=max(5,4)=5 seconds.

[0377] 4.

[0378] 5. T 剩余 =T 预期分析 -N 分析 *T 分析 =12-2×5=2 seconds.

[0379] 6. Due to (T 剩余 =2 seconds) = (T 短高光 = 2 seconds), so N 分析 =N 分析 +1=3.

[0380] 7.

[0381] 8. T 间隔 =max(T 间隔 ,T 最小间隔 )=max(4,3)=4 seconds.

[0382] So far, three target parameters have been obtained: T 分析 =5 seconds, N 分析 =3, T 间隔 =4 seconds.

[0383] As shown ​ in the figure, the available analysis duration T allocated to the i-th video 预期分析 is divided into 2 segments each with a duration of 5 seconds, and there is still 2 seconds of remaining analysis duration. Since T 间隔 = 4 seconds is greater than T 最小间隔 = 3 seconds, so T 间隔 finally takes the value of 4 seconds. When dividing the analysis segments, the analysis segments are allocated at equal intervals starting from the first frame of the i-th video file. Among them, the durations of analysis segment N1 and analysis segment N2 are both 5 seconds, and the duration of analysis segment N3 is equal to the remaining analysis duration (2 seconds). The interval between analysis segment N1 and analysis segment N2 is 4 seconds, and the interval between analysis segment N2 and analysis segment N3 is 4 seconds.

[0384] Example 3, assume T 短高光 = 2 seconds, T 长高光 = 4 seconds, T 最小间隔 = 3 seconds, T 预期分析 = 15 seconds, T 总 = 20 seconds, then the calculation process of the 3 target parameters is as follows:

[0385] 1.

[0386] 2.

[0387] 3. T 分析 = max(T 长高光 , T 下限 ) = max(4, 7.5) = 7.5 seconds.

[0388] 4.

[0389] 5. T 剩余 = T 预期分析 - N 分析 * T 分析 = 15 - 2×7.5 = 0 seconds.

[0390] 6. Since (T 剩余 = 0 seconds) < (T 短高光 = 2 seconds), so N 分析 = 2.

[0391] 7.

[0392] 8. T 间隔 = max(T 间隔 , T 最小间隔 ) = max(5, 3) = 5 seconds.

[0393] So far, the 3 target parameters have been obtained: T 分析= 7.5 seconds, N 分析 = 2, T 间隔 = 5 seconds.

[0394] As ​ shown, the available analysis duration T allocated to the i-th video 预期分析 is divided into 2 segments each with a duration of 7.5 seconds, with no remaining analysis duration. Since T 间隔 = 5 seconds is greater than T 最小间隔 = 3 seconds, so T 间隔 finally takes the value of 5 seconds. When dividing the analysis segments, the analysis segments are allocated at equal intervals starting from the first frame of the i-th video file. Among them, the analysis segment N1 is 7.5 seconds, the first interval segment is 5 seconds, and the analysis segment N2 is 7.5 seconds.

[0395] Next, in combination with ​ , an example is given to illustrate the process of obtaining the highlight segments according to the intensive key segment analysis strategy. As ​ shown, the process of obtaining the highlight segments according to the intensive key segment analysis strategy includes C1 - C10.

[0396] C1, the strategy monitoring module determines the following 3 target parameters according to the minimum duration T 短高光 of the highlight segments, the maximum duration T 长高光 of the highlight segments, the minimum interval T 最小间隔 of the highlight segments, the available analysis duration T 预期分析 allocated to the i-th video, and the total duration T 总 of the i-th video: the duration T 分析 of the analysis segments, the number N 分析 of the analysis segments, and the interval duration T 间隔 of the analysis segments. Then, the strategy monitoring module determines multiple analysis segments in the i-th video according to these 3 target parameters.

[0397] For the implementation method of C1, reference can be made to the specific descriptions in (1) - (5) above, which will not be elaborated here.

[0398] C2, the strategy monitoring module sends the file descriptor fdi of the i-th video to the channel interface, determines multiple analysis segments in the i-th video, and the analysis strategy (intensive key segment analysis strategy) set for the i-th video.

[0399] C3, the channel interface decodes the multiple analysis segments in the i-th video according to the file descriptor fd1 of the i-th video and the analysis strategy (intensive key segment analysis strategy) set for the i-th video.

[0400] C4, the channel interface converts the video frames of each decoded analysis segment from the first format to the second format.

[0401] Exemplarily, the first format is the nv12 format and the second format is the i420 format.

[0402] C5. The channel interface reduces the resolution of the video frame after format change from the first resolution to the second resolution.

[0403] Exemplarily, the first resolution is 1080p and the second resolution is 480p.

[0404] In some embodiments, the second resolution corresponding to different analysis strategies may be different or the same. Therefore, in C3 above, the policy monitoring module issues the analysis strategy (intensive key segment analysis strategy) set for the i-th video, so that the channel interface can determine the second resolution corresponding to the intensive key segment analysis strategy.

[0405] C6. The channel interface converts the video frame after resolution reduction from the second format to the first format.

[0406] To improve the processing speed of the highlight segment algorithm interface for the video frames of each analysis segment, before performing highlight analysis on the video frame, the channel interface can perform resolution reduction processing on the video frame. Additionally, since the resolution reduction algorithm only supports the second format, it is necessary to first convert the video frames of each analysis segment from the first format to the second format, then reduce the resolution of the video frame, and then convert the video frame after resolution reduction back to the original first format.

[0407] C7. The channel interface stores the video frames of each analysis segment that are converted back to the first format in the memory (buffer).

[0408] C8. The channel interface sends the data address of the video frame of the analysis segment to the highlight segment algorithm interface through the one-click video compilation interface.

[0409] C9. The highlight segment algorithm interface obtains the video frame of the analysis segment according to the data address of the video frame of the analysis segment, and then analyzes the video frame of the analysis segment based on the preset highlight segment algorithm to determine the positions of one or more highlight segments in the analysis segment. Among them, the highlight segment position may include the start address and the end address of the highlight segment.

[0410] In some embodiments, since each segment in the analysis segment consists of multiple video frames, the channel interface can send the data address of each video frame of the analysis segment to the highlight segment algorithm interface in the order from start to end. Correspondingly, the highlight segment algorithm interface can analyze the image features of the data frames corresponding to each frame data address according to the receiving order of each frame data address to obtain the scoring results of each frame.

[0411] Exemplarily, as shown in Table 1, assume that the policy monitoring module extracts three analysis segments from the i-th video: analysis segment N1, analysis segment N2, and analysis segment N3. Among them, the segment address of analysis segment N1 is 0000 0100—000001FF, the segment address of analysis segment N2 is 0000 0300—0000 03FF, and the segment address of analysis segment N3 is 00000500—0000 05FF. After the channel interface performs downscaling processing on these 3 analysis segments, the channel interface can first send the data address 0000 0100 of the first frame in analysis segment N1 to the highlight segment algorithm interface through the one-click video generation interface. The highlight segment algorithm interface obtains the first frame of analysis segment N1 from the data address 0000 0100 and scores the first frame of analysis segment N1; then sends the data address 0000 0101 of the second frame in analysis segment N1 to the highlight segment algorithm interface through the one-click video generation interface. The highlight segment algorithm interface obtains the second frame of analysis segment N1 from the data address 0000 0101 and scores the second frame of analysis segment N1... After that, send the data address 0000 01FF of the last frame in analysis segment N1 to the highlight segment algorithm interface through the one-click video generation interface. The highlight segment algorithm interface obtains the last frame of analysis segment N1 from the data address 0000 01FF and scores the last frame of analysis segment N1. Then, the channel interface can score each data frame of analysis segment N2 according to the above method. Then, the channel interface can score each data frame of analysis segment N3 according to the above method. After obtaining the scoring results of all frames, the highlight segment algorithm interface calculates the position of the highlight segment of the i-th video according to the scoring results of all frames. The position of the highlight segment includes the start address and end address of the highlight segment. It can be understood that each data frame of the highlight segment is the data frame extracted from analysis segment N1, analysis segment N2, and analysis segment N3.

[0412] Table 1

[0413] ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​

[0414] C10, the highlight segment algorithm interface sequentially returns the analysis results of the i-th video, such as the position of the highlight segment, to the policy monitoring module through the one-click video generation interface and the channel interface.

[0415] After the policy monitoring module obtains the analysis results of the i-th video, if there is still a next video, then the terminal device can extract the highlight segment from the next video according to the analysis policy corresponding to the next video. The analysis policy corresponding to the next video can be an intensive key segment analysis policy or other analysis policies. After obtaining the analysis results of all videos, the terminal device can report all analysis results.

[0416] for ​ The implementation of S01-S24, S30-S38 in the above embodiment can refer to the implementation of S01-S24, S30-S38 in the above embodiment. ​ The detailed description of S01-S24 and S30-S38 will not be repeated here.

[0417] It can be understood that compared with the full analysis strategy that decodes, reduces the resolution and scores all video frames, the intensive analysis fragment analysis strategy only needs to decode, reduce the resolution and score the video frames of a few key fragments. Therefore, the intensive analysis fragment analysis strategy can obtain highlight fragments faster and generate a video set consisting of highlight fragments.

[0418] Example 4

[0419] In combination with the description of the above embodiment, the policy monitoring module can calculate the available analysis time T allocated to the i-th video. 预期分析 The analysis time of the i-th video is T 分析耗时 When the ratio R is less than or equal to the preset ratio R1 (taking R1=0.6 as an example), it indicates that it is insufficient to analyze 60% of the video content. In this case, the analysis strategy for the i-th video is set to the sparse key segment analysis strategy.

[0420] A sparse key segment analysis strategy is used when the following relationships are met:

[0421]

[0422] Among them, T 总 is the total length of the i-th video, T 芯片分析速度 Analyze the speed of the chip.

[0423] The idea of sparse key segment analysis strategy is: first, according to the minimum length T of the highlight segment, 短高光 , the maximum duration of the highlight segment T 长高光 , the minimum interval T of highlight fragments 最小间隔 , the available analysis time T allocated to the i-th video 预期分析 , and the total duration of the i-th video T 总 , calculate the three target parameters - the number of analysis segments N of the material video 分析 , the duration of the analysis segment T 分析 , the interval length T between adjacent segments 间隔 (like ​ Then, determine the starting point of the first analysis segment (as shown in ​ Then, according to these three target parameters, the interval lengths of adjacent segments are distributed unequally (as shown in ​as shown, and the analysis segments are mainly distributed in the first and middle parts of the video, with appropriate allocation in the end part. Finally, each analysis segment of the i-th video is scored to obtain the analysis result of the highlight segment (such as ​ as shown).

[0424] Exemplarily, as ​ shown, a certain frame that meets the picture quality requirements among the first frame of the i-th video, the I-frame at the 2nd second position, or the I-frame at the 1 / 5 position of the video is used as the starting point of the first analysis segment N1. Then, starting from the starting point, N analysis segments with a duration of T are set at different intervals. The interval duration between analysis segment N1 and analysis segment N2 is T1, and the interval duration between analysis segment N2 and analysis segment N3 is T2 (T2 = 2T1), and so on. The duration of the last analysis segment Nn is set to T 分析 . 剩余 .

[0425] It should be noted that for the number N 分析 of the analysis segments of the material video, the duration T 分析 of the analysis segments, and the interval duration T 间隔 between adjacent segments, the specific acquisition methods of these three target parameters are the same as those of the intensive key segment analysis strategy. The relevant descriptions in the above Embodiment 3 can be referred to and will not be elaborated here.

[0426] Next, in combination with ​ , an example is given to illustrate the process of determining the starting point of the first analysis segment of the sparse key segment analysis strategy.

[0427] D1. The policy monitoring module sends the file descriptor fdi of the i-th video and the analysis policy (sparse key segment analysis policy) set for the i-th video to the channel interface.

[0428] D2. The channel interface decodes the three key frames of the i-th video according to the file descriptor fdi of the i-th video and the analysis policy (sparse key segment analysis policy) set for the i-th video.

[0429] By sending the analysis policy (sparse key segment analysis policy) set for the i-th video, the policy monitoring module enables the channel interface to determine that the video frames to be decoded are the three key frames of the i-th video according to this policy.

[0430] For example, the three key frames may include the first frame at the starting point of the i-th video, the I-frame at the 2nd second position, and the I-frame at the 1 / 5 position of the i-th video.

[0431] It should be noted that different from the third key frame of the simple I-frame analysis strategy, which is the "I-frame at the 1 / 3 position of the i-th video", the third key frame of the sparse key segment analysis strategy is the "I-frame at the 1 / 5 position of the i-th video". This is because the simple I-frame analysis strategy directly determines a highlight segment, while the sparse key segment analysis strategy needs to first determine multiple analysis segments and then screen out the highlight segments from them. To make multiple analysis segments cover the first half of the i-th video as much as possible, the third key frame of the sparse key segment analysis strategy is closer to the start position of the video.

[0432] In addition, the embodiment of the present application takes the third key frame of the sparse key segment analysis strategy as the "I-frame at the 1 / 5 position of the i-th video" as an example for illustration, and it can also be the "I-frame at the 1 / 4 position of the i-th video" or the "I-frame at the 1 / 6 position of the i-th video", etc., which can be adjusted according to actual needs.

[0433] D3. The channel interface converts the three decoded key frames from the first format to the second format.

[0434] Exemplarily, the first format is the nv12 format and the second format is the i420 format.

[0435] D4. The channel interface reduces the three key frames after format change from the first resolution to the second resolution.

[0436] Exemplarily, the first resolution is 1080p and the second resolution is 480p.

[0437] In some embodiments, the second resolution corresponding to different analysis strategies may be different or the same. Therefore, in D2 above, the strategy monitoring module issues the analysis strategy (sparse key segment analysis strategy) set for the i-th video, so that the channel interface can determine the second resolution corresponding to the sparse key segment analysis strategy.

[0438] D5. The channel interface converts the three key frames after resolution reduction from the second format to the first format.

[0439] To improve the processing speed of the highlight segment algorithm interface for the three key frames, before performing highlight analysis on the video frames, the channel interface can perform resolution reduction processing on the video frames. In addition, since the resolution reduction algorithm only supports the second format, it is necessary to first convert the three key frames from the first format to the second format, then reduce the resolution of the video frames, and then convert the video frames after resolution reduction to the original first format.

[0440] D6. The channel interface stores the three key frames that are converted back to the first format in the memory (buffer).

[0441] D7. The channel interface sends the data address of the first frame at the starting point of the i-th video to the highlight segment algorithm interface through the one-click video generation interface.

[0442] D8. The highlight segment algorithm interface obtains the first frame according to the data address of the first frame, then analyzes the first frame based on the preset highlight segment algorithm to obtain the score of the first frame, and then returns the score of the first frame to the policy monitoring module through the one-click video generation interface and the channel interface in sequence.

[0443] D9. The policy monitoring module determines whether the score of the first frame is greater than or equal to the preset value.

[0444] Taking the preset value as 60 points as an example. If the score of the first frame is greater than or equal to 60 points, then the first frame can be used as the starting point of the first analysis segment, and D16 is executed without executing D10 - D15. If the score of the first frame is less than 60 points, then the following D10 is executed.

[0445] D10. The policy monitoring module notifies the channel interface to send the data address of the I-frame at the 2nd second position to the highlight segment algorithm interface. The channel interface sends the data address of the I-frame at the 2nd second position to the highlight segment algorithm interface through the one-click video generation interface.

[0446] D11. The highlight segment algorithm interface obtains the I-frame at the 2nd second position according to the data address of the I-frame at the 2nd second position, then analyzes the I-frame at the 2nd second position based on the preset highlight segment algorithm to obtain the score of the I-frame at the 2nd second position, and then returns the score of the I-frame at the 2nd second position to the policy monitoring module through the one-click video generation interface and the channel interface in sequence.

[0447] D12. The policy monitoring module determines whether the score of the I-frame at the 2nd second position is greater than or equal to the preset value.

[0448] Still taking the preset value as 60 points as an example. If the score of the I-frame at the 2nd second position is greater than or equal to 60 points, then the I-frame at the 2nd second position can be used as the starting point of the first analysis segment, and D16 is executed without executing D13 - D15. If the score of the I-frame at the 2nd second position is less than 60 points, then the following D13 is executed.

[0449] D13. The policy monitoring module notifies the channel interface to send the data address of the I-frame at the 1 / 5 position to the highlight segment algorithm interface. The channel interface sends the data address of the I-frame at the 1 / 5 position to the highlight segment algorithm interface through the one-click video generation interface.

[0450] D14. The high - light segment algorithm interface obtains the I - frame at the 1 / 5 position according to the data address of the I - frame at the 1 / 5 position, then analyzes the I - frame at the 1 / 5 position based on a preset high - light segment algorithm to obtain the score of the I - frame at the 1 / 5 position, and then returns the score of the I - frame at the 1 / 5 position to the policy monitoring module through the one - click video generation interface and the channel interface in sequence.

[0451] D15. The policy monitoring module determines whether the score of the I - frame at the 1 / 5 position is greater than or equal to a preset value.

[0452] Still taking the preset value of 60 points as an example. If the score of the I - frame at the 1 / 5 position is greater than or equal to 60 points, then the I - frame at the 1 / 5 position can be used as the starting point of the first analysis segment, and D16 is executed. If the score of the I - frame at the 1 / 5 position is less than 60 points, then the video frame with the highest score among the three key frames can be used as the starting point of the first analysis segment, and D16 is executed.

[0453] It should be noted that the above D9, D12, and D15 are described by taking the preset values at the three positions as equal. In actual implementation, these three preset values may not be exactly equal, and the embodiments of the present application do not limit this.

[0454] D16. The policy monitoring module determines a certain frame among the three key frames as the starting point of the first analysis segment, that is, the starting position of the first analysis segment.

[0455] The above embodiments D1 - D16 are described by taking the determination of the starting point of the first analysis segment using three key frames as an example, which does not limit the embodiments of the present application. In actual implementation, it can also be two key frames, or four key frames, etc.

[0456] After determining the starting point of the first analysis segment, the policy monitoring module also needs to calculate the intervals of each analysis segment according to the interval duration T of adjacent segments 间隔 and the minimum interval T of high - light segments 最小间隔 .

[0457] Specifically, the first interval duration T1 set between the first analysis segment N1 and the second analysis segment N2 = max(T 间隔 / n, T 最小间隔 ). The interval duration T i after the second analysis segment N2 = min(m * T i-1 , T 间隔 ), where T i-1 is the previous interval duration of T i . Here, both n and m are greater than 1, and i is an integer greater than or equal to 2.

[0458] Taking the maximum number N of analysis segments as an example, 上限 = 5, n = 4, m = 2, the process of the policy monitoring module calculating the intervals of each analysis segment will be illustrated with reference to ​ .

[0459] When the number N of analysis segments 分析 ≥ 2 (for example, N 分析 = 2, 3, 4 or 5), a first interval duration T1 is set between the first analysis segment N1 and the second analysis segment N2.

[0460] The process for obtaining the first interval duration T1 is as shown in (a) of ​ :

[0461] The policy monitoring module determines whether T 间隔 / 4 is greater than T 最小间隔 ;

[0462] If T 间隔 / 4 > T 最小间隔 , then T1 = T 间隔 / 4;

[0463] If T 间隔 / 4 ≤ T 最小间隔 , then T1 = T 最小间隔 .

[0464] It should be understood that the minimum interval T of the highlight segment 最小间隔 is a minimum interval calculated by the application layer. To avoid the problem of overly concentrated analysis segments due to too small intervals, when T 间隔 / 4 is less than or equal to the minimum interval T of the highlight segment 最小间隔 , the first interval duration T1 should be set to T 最小间隔 .

[0465] When the number N of analysis segments 分析 ≥ 3 (for example, N 分析 = 3, 4 or 5), a second interval duration T2 is set between the second analysis segment N2 and the third analysis segment N3.

[0466] The process for obtaining the second interval duration T2 is as shown in (b) of ​ :

[0467] The policy monitoring module determines whether 2 * T1 is less than T 间隔 ;

[0468] If 2 * T1 < T 间隔 , then T2 = 2 * T1;

[0469] If 2 * T1 ≥ T 间隔, then T2 = T 间隔 .

[0470] It should be understood that by taking the smaller value between 2*T1 and T 间隔 as T2, the interval duration between the second analysis segment N2 and the third analysis segment N3 can be shortened, making the analysis segments more concentrated in the first and middle parts of the video. Since the probability of the first and middle parts of the video being the content of interest to users is relatively high, such a setting is more conducive to accurately and quickly extracting the highlight segments.

[0471] When the number of analysis segments N 分析 ≥ 4 (for example, N 分析 = 4 or 5), a third interval duration T3 is set between the third analysis segment N3 and the fourth analysis segment N4.

[0472] The acquisition process of the third interval duration T3 is as shown in ​ (c) of

[0473] The policy monitoring module determines whether 2*T2 is less than T 间隔 ;

[0474] If 2*T2 < T 间隔 , then T3 = 2*T2;

[0475] If 2*T2 ≥ T 间隔 , then T3 = T 间隔 .

[0476] It should be understood that by taking the smaller value between 2*T2 and T 间隔 as T3, the interval duration between the third analysis segment N3 and the fourth analysis segment N4 can be shortened, making the analysis segments more concentrated in the first and middle parts of the video. Since the probability of the first and middle parts of the video being the content of interest to users is relatively high, such a setting is more conducive to accurately and quickly extracting the highlight segments.

[0477] When the number of analysis segments N 分析 ≥ 5 (for example, N 分析 = 5), a fourth interval duration T4 is set between the fourth analysis segment N4 and the fifth analysis segment N5. The acquisition process of the fourth interval duration T4 is as shown in ​ (d) of

[0478] The acquisition process of the fourth interval duration T4 is as shown in ​ (d) of <[

[0479] The policy monitoring module determines whether 2*T3 is less than T 间隔 ;

[0480] If 2*T3 < T 间隔 , then T4 = 2*T3;

[0481] If 2*T3 ≥ T 间隔 , then T4 = T 间隔 .

[0482] It should be understood that by taking the smaller value between 2*T3 and T 间隔 as T4, the interval duration between the fourth analysis segment N4 and the fifth analysis segment N5 can be shortened, making the analysis segments more concentrated in the first and middle parts of the video. Since the probability of the first and middle parts of the video being the content of interest to users is relatively high, such a setting is more conducive to accurately and quickly extracting the highlight segments.

[0483] It should be noted that in some cases, after setting each segment and the segment interval, the remaining duration for the last analysis segment Nn may be less than the duration T of one analysis segment 分析 . For example, as shown in (e) of ​ , when N 分析 = 4 or 5, the segment corresponding to the remaining duration of the i-th video can be used as the last analysis segment Nn.

[0484] To facilitate understanding the principle of calculating the interval of the above analysis segments, the following takes the maximum number of analysis segments N 上限 = 5, n = 4, m = 2, and the I-frame at the 1 / 5 position as the starting point of the first analysis segment as an example, and through two specific examples, the calculation process of the interval of the analysis segments is explained.

[0485] Example 1, assume that the five initial parameters are T 短高光 = 2 seconds, T 长高光 = 5 seconds, T 最小间隔 = 2 seconds, T 预期分析 = 15 seconds, T 总 = 20 seconds. The three target parameters obtained according to the method provided in the third embodiment are T 分析 = 5 seconds, N 分析 = 3, T 间隔 = 2.5 seconds.

[0486] 1. Since (T 间隔 / 4 = 0.625) < (T 最小间隔 = 2), therefore, the first interval duration T1 = T 最小间隔 = 2 seconds is set between the first analysis segment N1 and the second analysis segment N2.

[0487] 2. Since (2*T1 = 4) > (T 间隔 = 2.5), therefore, the second interval duration T2 = T 间隔 = 2.5 seconds is set between the second analysis segment N2 and the third analysis segment N3.

[0488] As ​ shown, the total duration T of the i-th video 总 = 20 seconds. The I-frame at the 1 / 5 position is used as the starting point of the first analysis segment. According to T 分析 = 5 seconds, the duration of the first analysis segment N1 is set to 5 seconds. The first interval duration between the first analysis segment N1 and the second analysis segment N2 is set to T1 = 2 seconds. According to T 分析 = 5 seconds, the duration of the second analysis segment N2 is set to 5 seconds. The second interval duration between the second analysis segment N2 and the third analysis segment N3 is set to T2 = 2.5 seconds. After setting the above segments and intervals, the remaining duration of the third analysis segment N3 is less than 5 seconds. At this time, the segment corresponding to the remaining duration of 1.5 seconds is used as the third analysis segment N3.

[0489] Example 2, assume that the 5 initial parameters are T 短高光 = 2 seconds, T 长高光 = 5 seconds, T 最小间隔 = 3 seconds, T 预期分析 = 12 seconds, T 总 = 20 seconds. The 3 target parameters obtained according to the method provided in Embodiment 3 are T 分析 = 5 seconds, N 分析 = 3, T 间隔 = 4 seconds.

[0490] 1. Since (T 间隔 / 4 = 1) < (T 最小间隔 = 3), therefore, the first interval duration T1 = T 最小间隔 = 3 seconds is set between the first analysis segment N1 and the second analysis segment N2.

[0491] 2. Since (2 * T1 = 6) > (T 间隔 = 4), therefore, the second interval duration T2 = T 间隔 = 4 seconds is set between the second analysis segment N2 and the third analysis segment N3.

[0492] As ​ shown, the total duration T of the i-th video 总 = 20 seconds. The I-frame at the 1 / 5 position is used as the starting point of the first analysis segment. According to T 分析 = 5 seconds, the duration of the first analysis segment N1 is set to 5 seconds. The first interval duration between the first analysis segment N1 and the second analysis segment N2 is set to T1 = 3 seconds. According to T 分析= 5 seconds, set the duration of the second analysis segment N2 to 5 seconds. After setting the above segments and intervals, the remaining duration is 3 seconds, which is less than the second interval duration of 4 seconds. At this time, the third analysis segment N3 is not set.

[0493] The following combines ​ , and gives an example to illustrate the process of obtaining highlight segments according to the sparse key segment analysis strategy. As ​ shown, the process of obtaining highlight segments according to the sparse key segment analysis strategy includes D0 - D27.

[0494] D0, the strategy monitoring module determines the following three target parameters according to the minimum duration T 短高光 of the highlight segment, the maximum duration T 长高光 of the highlight segment, the minimum interval T 最小间隔 of the highlight segment, the available analysis duration T 预期分析 assigned to the i-th video, and the total duration T 总 of the i-th video: the duration T 分析 of the analysis segment, the number N 分析 of the analysis segments, and the interval duration T 间隔 of the analysis segments.

[0495] D1 - D16, the strategy monitoring module determines a certain frame among the three key frames as the starting point of the first analysis segment.

[0496] For example, the three key frames can include the first frame at the starting point of the i-th video, the I-frame at the 2-second position, and the I-frame at the 1 / 5 position of the i-th video.

[0497] D17, the strategy monitoring module calculates the intervals of each analysis segment according to the interval duration T 间隔 between adjacent segments and the minimum interval T 最小间隔 of the highlight segment.

[0498] Specifically, the first interval duration T1 set between the first analysis segment N1 and the second analysis segment N2 = max(T 间隔 / n, T 最小间隔 ). The interval duration T i after the second analysis segment N2 = min(m * T i-1 , T 间隔 ), and T i-1 is the previous interval duration of T i . Here, both n and m are greater than 1, and i is an integer greater than or equal to 2.

[0499] D18, the strategy monitoring module determines the starting point of the first analysis segment, the intervals of each analysis segment, and the duration T 分析, and the number N of analysis segments 分析 , determine multiple analysis segments in the i-th video.

[0500] D19, the policy monitoring module sends a command to the channel interface to determine multiple analysis segments in the i-th video.

[0501] D20, the channel interface decodes the multiple analysis segments in the i-th video.

[0502] D21, the channel interface converts the video frames of each decoded analysis segment from the first format to the second format.

[0503] Exemplarily, the first format is the nv12 format, and the second format is the i420 format.

[0504] D22, the channel interface reduces the resolution of the video frames after format change from the first resolution to the second resolution.

[0505] Exemplarily, the first resolution is 1080p, and the second resolution is 480p.

[0506] D23, the channel interface converts the video frames after resolution reduction from the second format to the first format.

[0507] D24, the channel interface stores the video frames of each analysis segment that are converted back to the first format in the memory (buffer).

[0508] D25, the channel interface sends the data address of the video frames of the analysis segments to the highlight segment algorithm interface through the one-click video creation interface.

[0509] D26, the highlight segment algorithm interface obtains the video frames of the analysis segments according to the data addresses of the video frames of the analysis segments, and then analyzes the video frames of the analysis segments based on a preset highlight segment algorithm to determine the positions of one or more highlight segments in the analysis segments. Among them, the highlight segment position may include the start address and end address of the highlight segment.

[0510] In some embodiments, since each segment in the analysis segments consists of multiple video frames, the channel interface can send the data address of each video frame of the analysis segments to the highlight segment algorithm interface in the order from beginning to end. Correspondingly, the highlight segment algorithm interface can analyze the image features of the data frames corresponding to each frame data address according to the reception order of each frame data address, and obtain the scoring results of each frame.

[0511] D27, the highlight segment algorithm interface sequentially returns the analysis results of the i-th video, such as the positions of the highlight segments, to the policy monitoring module through the one-click video creation interface and the channel interface.

[0512] After the policy monitoring module obtains the analysis result of the i-th video, if there is still a next video, the terminal device can extract highlight segments from the next video according to the analysis policy corresponding to the next video. The analysis policy corresponding to the next video can be a sparse key segment analysis policy or other analysis policies. After obtaining the analysis results of all videos, the terminal device can report all the analysis results.

[0513] For ​ The implementation manners of S1-S24, S30-S38, and D0-D17 in [reference] can refer to the relevant descriptions in the above embodiments, and will not be elaborated here.

[0514] It can be understood that, compared with the intensive analysis segment analysis policy, the sparse key segment analysis policy focuses on analyzing the first half and middle part of the video, and the starting point of the first analysis segment may not be the first video frame. Therefore, the sparse key segment analysis policy may take less time, so that the highlight segments can be obtained faster, and a video set composed of highlight segments can be generated.

[0515] It can be understood that, in order to implement the above functions, the terminal device includes the corresponding hardware structure or software module for each function, or a combination of both. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0516] The embodiments of the present application can divide the terminal device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. There may be other division methods in actual implementation. The following takes the example of dividing each functional module corresponding to each function for illustration.

[0517] ​ It is a schematic structural diagram of a device for processing video segments provided by an embodiment of the present application. As ​ shown, the device 80 may include an input module 81, a determination module 82, an extraction module 83, and a generation module 84.

[0518] An input module 81 for receiving a user's selection operation on multiple videos.

[0519] A determination module 82 for, in response to a user's selection operation on multiple videos, determining the analysis time consumption and available analysis duration of each video among the multiple videos; and in the case where the available analysis duration of the first video among the multiple videos is less than the analysis time consumption of the first video and the available analysis duration of the first video is greater than the analysis duration of P video frames, determining the duration of the analysis segments corresponding to the first video, the number of analysis segments, and the interval duration between adjacent segments; then determining multiple analysis segments in the first video according to the duration of the analysis segments corresponding to the first video, the number of analysis segments, and the interval duration between adjacent segments. Here, the analysis time consumption of a video is the duration required to complete the image analysis of all frames of the video, and the available analysis duration of a video is the expected duration allocated for the image analysis of the video. P is an integer greater than or equal to 2.

[0520] An extraction module 83 for extracting highlight segments from multiple analysis segments.

[0521] A generation module 84 for, after the extraction module 83 extracts all highlight segments from multiple videos, generating all highlight segments into a video set.

[0522] In some embodiments, the determination module 82 is specifically configured to: in the case where the ratio of the available analysis duration of the first video to the analysis time consumption of the first video is greater than or equal to a preset ratio, allocate analysis segments starting from the first frame of the first video until the last frame of the first video, and finally obtain multiple analysis segments. Here, the duration of each analysis segment in the first video is equal to the duration T of the analysis segment 分析 , the interval between any two adjacent analysis segments in the first video is equal to the interval duration T between adjacent segments 间隔 , the number of multiple analysis segments in the first video is equal to the number N of analysis segments 分析 .

[0523] In some other embodiments, the determination module 82 is specifically configured to: in the case where the ratio of the available analysis duration of the first video to the analysis time consumption of the first video is less than the preset ratio, select a video frame from Q video frames as the starting point of the first analysis segment in the first video; and re-determine the intervals of each analysis segment according to the interval duration between adjacent segments and the minimum interval of highlight segments; and allocate analysis segments starting from the starting point of the first analysis segment until the last frame of the first video, and finally obtain multiple analysis segments. Here, Q is an integer greater than or equal to 2. The duration of the analysis segments in the first video is equal to the duration T of the analysis segment 分析 ; the interval between adjacent analysis segments in the first video is less than or equal to the interval duration T between adjacent segments 间隔and the interval before an analysis segment is less than or equal to the interval after an analysis segment; the number of multiple analysis segments in the first video is less than or equal to the number N of analysis segments 分析 .

[0524] An embodiment of this application also provides a terminal device, including a processor, the processor is coupled with a memory, and the processor is configured to execute a computer program or instruction stored in the memory, so that the terminal device implements the methods in the above embodiments.

[0525] An embodiment of this application also provides a computer-readable storage medium, in which computer instructions are stored; when the computer-readable storage medium runs on a terminal device, the terminal device is enabled to execute the method as shown above. The computer instructions may be stored in the computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server or a data center to another website, computer, server or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk or a magnetic tape), an optical medium or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0526] An embodiment of this application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer is enabled to execute the methods in the above embodiments.

[0527] An embodiment of this application also provides a chip, the chip is coupled with a memory, and the chip is configured to read and execute a computer program or instruction stored in the memory to execute the methods in the above embodiments. The chip may be a general-purpose processor or a dedicated processor. It should be noted that the chip may be implemented by the following circuits or devices: one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing various functions described throughout this application.

[0528] The terminal device, apparatus, computer-readable storage medium, computer program product, and chip provided in the embodiments of the present application are all used to execute the methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the methods provided above, and will not be elaborated here.

[0529] It should be noted that the terms "first" and "second" in the description, claims, and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0530] It should be understood that in the present application, "at least one" means one or more, "multiple" means two or more, "at least two" means two or three or more, and "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0531] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. For example, B can be determined according to A. It should also be understood that determining B according to A does not mean determining B only according to A, but B can also be determined according to A and / or other information. In addition, the "connection" that appears in the embodiments of the present application refers to various connection methods such as direct connection or indirect connection to achieve communication between devices, and the embodiments of the present application do not make any limitations on this.

[0532] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0533] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0534] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can also be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0535] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0536] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device, such as a single-chip microcomputer, a chip, etc., or a processor to execute all or part of the steps of the methods provided in each embodiment of the present application. The foregoing storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks or optical discs.

[0537] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A video clip processing method, characterized in that: The method comprises: In response to a user selecting a plurality of videos, determining an analysis time and an available analysis time for each of the plurality of videos, where the analysis time for a video is the time required to complete image analysis of all frames of the video, and the available analysis time for a video is the expected time allocated for image analysis of the video; When the available analysis time of a first video among the multiple videos is less than the analysis time of the first video, and the available analysis time of the first video is greater than the analysis time of P video frames, determining the time of analysis segments corresponding to the first video, the number of analysis segments, and the interval time between adjacent segments, where P is an integer greater than or equal to 2; Determining a plurality of analysis segments in the first video according to the duration of the analysis segments corresponding to the first video, the number of the analysis segments, and the interval durations between adjacent segments; extracting highlight segments from the plurality of analysis segments; After all highlight segments are extracted from the multiple videos, all the highlight segments are generated as a video set.

2. The method according to claim 1, characterized in that Determining the duration of the analysis segment corresponding to the first video includes: Determining a maximum number of analysis segments in the first video based on the total duration of the first video, the available analysis duration of the first video, and the minimum interval between highlight segments; The ratio of the available analysis time of the first video to the maximum number of analysis segments in the first video is used as the minimum length of the analysis segments in the first video; The maximum of the minimum duration of the analysis segment and the maximum duration of the highlight segment is used as the duration of the analysis segment corresponding to the first video.

3. The method according to claim 2, characterized in that Determining the maximum number of analysis segments in the first video according to the total duration of the first video, the available analysis duration of the first video, and the minimum interval between highlight segments includes: The maximum number of analysis segments in the first video is determined using the following relationship: in, is the floor rounding symbol; N 上限 represents the maximum number of analysis segments in the first video; T 总 Indicates the total duration of the first video; T 预期分析 Indicates the available analysis time of the first video; T 最小间隔 Indicates the minimum interval of highlight segments.

4. The method according to claim 1, wherein Determining the number of the analysis segments corresponding to the first video includes: The number of the analysis segments corresponding to the first video is determined according to a ratio of the available analysis duration of the first video to the duration of the analysis segments.

5. The method according to claim 4, characterized in that The number of analysis segments corresponding to the first video according to the ratio of the available analysis time of the first video to the time length of the analysis segments includes: If T 剩余 <T 短高光 , the number of the analysis segments corresponding to the first video is determined using the following relationship: If T 剩余 ≥T 短高光 , the number of the analysis segments corresponding to the first video is determined using the following relationship: in, T 预期分析 Indicates the available analysis time of the first video; T 分析 Indicates the duration of the analysis segment corresponding to the first video; T 短高光 Indicates the minimum length of the highlight segment; N 分析 Indicates the number of the analysis segments corresponding to the first video.

6. The method according to claim 1, characterized in that Determining the interval duration between adjacent segments corresponding to the first video includes: The interval duration of the adjacent segments corresponding to the first video is determined according to the number of the analysis segments corresponding to the first video, the total duration of the first video, and the available analysis duration of the first video.

7. The method according to claim 6, characterized in that Determining the interval duration of adjacent segments corresponding to the first video based on the number of analysis segments corresponding to the first video, the total duration of the first video, and the available analysis duration of the first video includes: The following relationship is used to determine the interval duration between the adjacent segments corresponding to the first video: Among them, the max() function is used to find the maximum value; T 间隔 Indicates the interval length between the adjacent segments corresponding to the first video; T 总 Indicates the total duration of the first video; T 预期分析 Indicates the available analysis time of the first video; N 分析 Indicates the number of the analysis segments corresponding to the first video; T 最小间隔 Indicates the minimum interval of highlight segments.

8. The method according to claim 1, characterized in that The determining of a plurality of analysis segments in the first video according to the duration of the analysis segment corresponding to the first video, the number of the analysis segments, and the interval duration of adjacent segments includes: When the ratio of the available analysis time of the first video to the analysis time of the first video is greater than or equal to a preset ratio, allocating analysis segments starting from the first frame of the first video to the last frame of the first video, to ultimately obtain the multiple analysis segments; Among them, the duration of each analysis segment in the first video is equal to the duration of the analysis segment, the interval between any two adjacent analysis segments in the first video is equal to the interval duration of the adjacent segments, and the number of the multiple analysis segments in the first video is equal to the number of analysis segments.

9. The method according to claim 1, characterized in that The determining of a plurality of analysis segments in the first video according to the duration of the analysis segment corresponding to the first video, the number of the analysis segments, and the interval duration of adjacent segments includes: When a ratio of the available analysis time of the first video to the analysis time of the first video is less than a preset ratio, selecting a video frame from the Q video frames as a starting point of a first analysis segment in the first video; Re-determining the intervals of the analysis segments according to the interval durations of the adjacent segments and the minimum interval of the highlight segments; Allocating analysis segments starting from the starting point of the first analysis segment until the last frame of the first video, and finally obtaining the plurality of analysis segments; Wherein, Q is an integer greater than or equal to 2; the duration of the analysis segment in the first video is equal to the duration of the analysis segment; the interval between adjacent analysis segments in the first video is less than or equal to the interval duration of the adjacent segments, and the interval before an analysis segment is less than or equal to the interval after the analysis segment; the number of the multiple analysis segments in the first video is less than or equal to the number of analysis segments.

10. The method according to claim 9, characterized in that The Q video frames include: a first video frame located at the starting point of the first video, a video frame located at the 2nd second position of the first video, and a video frame located at the 1 / 5 position of the first video; Determining a video frame from the Q frames as a starting point of a first analysis segment in the first video includes: Obtaining a score for a first video frame located at the starting point of the first video; If the score of the first video frame at the starting point of the first video is greater than or equal to a preset value, use the first video frame at the starting point of the first video as the starting point of the first analysis segment; or, if the score of the first video frame at the starting point of the first video is less than the preset value, obtain the score of the video frame at the 2nd second position of the first video; If the score of the video frame at the 2nd second position of the first video is greater than or equal to the preset value, the video frame at the 2nd second position of the first video is used as the starting point of the first analysis segment; or, if the score of the video frame at the 2nd second position of the first video is less than the preset value, the score of the video frame at the 1 / 5 position of the first video is obtained; When the score of the video frame located at the 1 / 5 position of the first video is greater than or equal to the preset value, the video frame located at the 1 / 5 position of the first video is used as the starting point of the first analysis segment; or, when the score of the video frame located at the 1 / 5 position of the first video is less than the preset value, the first video frame located at the starting point of the first video, the video frame located at the 2nd second position of the first video, and the video frame located at the 1 / 5 position of the first video, whichever has the highest score, is used as the starting point of the first analysis segment.

11. The method according to claim 9, characterized in that The re-determining the interval of each analysis segment according to the interval duration of the adjacent segments and the minimum interval of the highlight segments includes: If the number of analysis segments is greater than or equal to 2, the first interval duration set between the first analysis segment and the second analysis segment is determined using the following relationship: T1=max(T 间隔 / n,T 最小间隔 ); If the number of analysis segments is greater than or equal to 3, the following relationship is used to determine the i-th interval duration between the i-th analysis segment and the i+1-th analysis segment: T i =min(m*T i-1 ,T 间隔 ); Among them, the max() function is used to find the maximum value; The min() function is used to find the minimum value; T i represents the duration of the i-th interval; T i-1 Indicates the i-1th interval duration set between the i-1th analysis segment and the i-th analysis segment; Both n and m are greater than 1, and i is an integer greater than or equal to 2.

12. The method according to any one of claims 8 to 11, characterized in that If, after allocating the last interval, the remaining duration of the first video is less than the duration of the analysis segment, the remaining duration of the first video is used as the duration of the last analysis segment in the first video.

13. The method according to any one of claims 1 to 11, characterized in that After determining the analysis time consumption and available analysis time of each of the multiple videos, and before extracting all highlight segments from the multiple videos, the method further includes: When the available analysis time of a second video among the multiple videos is greater than or equal to the analysis time of the second video, obtaining a score for each video frame in the second video; The video frames in the second video whose scores are greater than or equal to a preset value are used as highlight segments in the second video.

14. The method according to any one of claims 1 to 11, characterized in that P=3; after determining the analysis time and available analysis time for each of the multiple videos, and before extracting all highlight segments from the multiple videos, the method further includes: When the available analysis duration of a third video among the multiple videos is less than or equal to the analysis duration of three video frames, obtaining a score of a first video frame located at the starting point of the third video; If the score of the first video frame at the starting point of the third video is greater than or equal to a preset value, the first video frame at the starting point of the third video is used as the starting point of the highlight segment, and a video segment of the target duration is selected as the highlight segment of the third video; or, if the score of the first video frame at the starting point of the third video is less than the preset value, the score of the video frame at the 2nd second of the third video is obtained; If the score of the video frame at the 2nd second position of the third video is greater than or equal to the preset value, the score of the video frame at the 2nd second position of the third video is used as the starting point of the highlight segment, and the video segment of the target duration is selected as the highlight segment of the third video; or, if the score of the video frame at the 2nd second position of the third video is less than the preset value, the score of the video frame at the 1 / 3 position of the third video is obtained; In the case that the score of the video frame located at the 1 / 3 position of the third video is greater than or equal to the preset value, the score of the video frame located at the 1 / 3 position of the third video is used as the starting point of the highlight segment, and the video segment of the target length is selected as the highlight segment of the third video; or, in the case that the score of the video frame located at the 1 / 3 position of the third video is less than the preset value, the first video frame located at the starting point of the third video, the video frame located at the 2nd second position of the third video, and the video frame located at the 1 / 3 position of the third video, whichever has the highest score, is used as the starting point of the highlight segment, and the video segment of the target length is selected as the highlight segment.

15. The method according to claim 14, characterized in that The target duration is determined using the following relationship: T 高光 =min(1.2*T 短高光 ,T 长高光 ); Among them, the min() function is used to find the minimum value; T 高光 Indicates the target duration; T 短高光 Indicates the minimum length of the highlight segment; T 长高光 Indicates the maximum duration of the highlight clip.

16. The method according to any one of claims 1 to 11, characterized in that The analysis time of a video is equal to the total length of the video divided by the analysis speed of the video. The analysis speed of a video is determined by the resolution of the video, the frame rate of the video, and the chip analysis speed.

17. The method according to any one of claims 1 to 11, characterized in that Determining the analysis time of each of the multiple videos includes: Determining an analysis speed for each video based on a chip analysis speed, a resolution of each video, and a frame rate of each video; Determining a target total analysis duration based on the analysis speed of each video, where the target total analysis duration is the expected total duration for completing the analysis of the multiple videos; The analysis time of each video is determined based on the total target analysis time and the analysis time of each video.

18. A terminal device, characterized in that: It includes a processor, a communication interface, and a memory coupled to the processor and the communication interface; wherein the memory stores instructions, and when the processor executes the instructions, the terminal device executes the video clip processing method according to any one of claims 1 to 17.

19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed on a terminal device, the terminal device executes the video segment processing method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Sports event video editing method and system based on artificial intelligence

    CN113490049A

  • Media Processing

    US20190370558A1