Target video generation method and device, computer device, and storage medium
By using clustering algorithms and feature information to process video clips in video editing templates, the problem of duration mismatch is solved, and a target video that matches the style of the editing template is generated, thereby improving the efficiency and effectiveness of video editing.
Patent Information
- Application Number
- CN202411464143.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-10-19
AI Technical Summary
Existing video editing templates cannot accurately process video clips that are too short or too long, resulting in the inability to generate adapted videos.
By determining the adjusted duration of the video clip to be processed, clustering is performed using a clustering algorithm combined with the feature information and time information of the video clip to generate video sub-segments, and the expansion or cropping method is determined according to the image features of the video sub-segments to generate the target video.
It achieves accurate expansion or cropping of video clips, ensures that the generated target video matches the style of the editing template, and improves the efficiency and effect of video editing.
Smart Images

Figure CN119383396B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of video processing, and particularly relates to a target video generation method and device, computer equipment and a storage medium. BACKGROUND
[0002] With the development of various video platforms, users usually use various editing software to edit the videos that are shot. In order to improve the efficiency of video editing and cater to the preferences of various users, various editing templates are added in the editing software.
[0003] However, in general, each video segment in each editing template has a fixed duration, and if the duration of the video segment put into the editing template by the user is too short or too long, the editing template cannot accurately process the video with the duration that is too short or too long, so that the video adapted to the editing template cannot be generated. SUMMARY
[0004] Therefore, it is necessary to provide a target video generation method, device, computer equipment and storage medium to solve the above technical problems.
[0005] In a first aspect, the present disclosure provides a target video generation method. The method comprises:
[0006] According to the positions of the plurality of video segments input into the editing template, determining a to-be-processed video segment in the plurality of video segments, and determining an adjusted duration matched with each to-be-processed video segment;
[0007] In response to the to-be-processed video segment being a to-be-expanded video segment, determining a cluster number according to the adjusted duration matched with the to-be-expanded video segment, and performing clustering according to the feature information of the to-be-expanded video segment, the time information of each frame of video image in the to-be-expanded video segment and the cluster number, to obtain a plurality of video sub-segments;
[0008] According to the image features of the plurality of video sub-segments, determining a video expansion mode between the plurality of video sub-segments, and generating a target segment;
[0009] In response to the to-be-processed video segment being a to-be-cut video segment, determining a first video style of the to-be-cut video segment, and determining a target segment in the to-be-cut video segment according to the relationship between the adjusted duration matched with the to-be-cut video segment, the first video style and a second video style of the editing template;
[0010] Generating a target video at least according to the generated target segment and the editing template.
[0011] In one of the embodiments, the clustering according to the feature information of the video segment to be augmented, the time information of each frame of video image in the video segment to be augmented, and the cluster number, comprises:
[0012] clustering the video images according to the feature information of each frame of video image, the time information corresponding to the frame of video image, and the cluster number, to obtain a clustering result matching the cluster number;
[0013] determining a cluster boundary between the clustering results, and segmenting the video segment to be augmented according to the cluster boundary to obtain a plurality of video sub-segments.
[0014] In one of the embodiments, the determining the cluster boundary between the clustering results comprises:
[0015] in response to the existence of a time overlap region in the clustering results, determining the cluster boundary between the clustering results based on a time point in the time overlap region;
[0016] the determining the cluster boundary between the clustering results based on the time point in the time overlap region comprises:
[0017] determining a time median point in the time overlap region, and determining the cluster boundary between the clustering results according to the time median point;
[0018] or, determining a total number of feature information contained in the time overlap region, and determining a first number of clustering results existing the same time overlap region;
[0019] determining a time point from the time overlap region, which divides the total number by the first number, and determining the cluster boundary between the clustering results based on the time point.
[0020] In one of the embodiments, after the obtaining the clustering result matching the cluster number, the method further comprises:
[0021] calculating a profile coefficient of the clustering result;
[0022] in response to the profile coefficient being less than a preset coefficient threshold, decomposing the feature information by using a principal component analysis method to obtain feature information of multiple dimensions;
[0023] combining the feature information of multiple dimensions and the time information corresponding to each frame of image to obtain a modified clustering feature group;
[0024] clustering according to the modified clustering feature group and the cluster number to obtain a modified clustering result.
[0025] In one of the embodiments, the determining the first video style of the video clip to be cut includes:
[0026] color extraction is performed on each frame of video image in the video clip to be cut to obtain color information;
[0027] a color spectrum of each frame of video image is constructed according to the color information, and a scene style of each frame of video image is determined according to a usage frequency of color and a combination of color indicated by the color spectrum;
[0028] a camera movement style between each frame of video image in the video clip to be cut is determined according to a change between each frame of video image in the video clip to be cut;
[0029] the first video style of the video clip to be cut is determined according to the camera movement style between each frame of video image and the scene style of each frame of video image.
[0030] In one of the embodiments, the determining the target segment in the video clip to be cut according to the relationship between the adjustment duration matched with the video clip to be cut, the first video style and the second video style of the clip template includes:
[0031] in response to the second video style being the same as the first video style, a first material area in each frame of video image of the video clip to be cut having a matching relationship with the first video style is determined;
[0032] a style score of each frame of video image is determined according to the first material area, and a target segment in the video clip to be cut is determined according to the style score of each frame of video image and the duration threshold;
[0033] in response to the second video style being different from the first video style, a second material area in each frame of video image of the video clip to be cut having a matching relationship with the second video style is determined;
[0034] a style score of each frame of video image is determined according to the second material area, and a target segment in the video clip to be cut is determined according to the style score of each frame of video image and the duration threshold.
[0035] In one of the embodiments, when the first video style is a motion style, the determining the target segment in the video clip to be cut according to the relationship between the duration threshold, the first video style and the second video style includes:
[0036] in response to the second video style being the same as the first video style, pose information corresponding to each frame of video image in the video clip to be cut is determined;
[0037] determine a change range of the pose information between each frame of video images according to the pose information corresponding to each frame of video images;
[0038] sort each frame of video images according to the change range, and filter the sorted video images according to the time length threshold to obtain at least one frame of video image satisfying the time length threshold;
[0039] determine a target segment in the video segment to be cropped according to the at least one frame of video image satisfying the time length threshold.
[0040] In a second aspect, the present disclosure further provides a target video generation device. The device comprises:
[0041] a video segment determination module, configured to determine a video segment to be processed in a plurality of video segments according to a position of the video segment to be processed in a clip template, and determine an adjusted time length matched with each of the video segment to be processed;
[0042] a clustering processing module, configured to, in response to the video segment to be processed being a video segment to be expanded, determine a clustering number according to the adjusted time length matched with the video segment to be expanded, and perform clustering according to feature information of the video segment to be expanded, time information of each frame of video images in the video segment to be expanded, and the clustering number to obtain a plurality of video sub-segments;
[0043] an expansion module, configured to determine a video expansion manner between the plurality of video sub-segments according to image features of the plurality of video sub-segments, and generate a target segment;
[0044] a target segment determination module, configured to, in response to the video segment to be processed being a video segment to be cropped, determine a first video style of the video segment to be cropped, and determine a target segment in the video segment to be cropped according to a relationship between the adjusted time length matched with the video segment to be cropped, the first video style, and a second video style of the clip template;
[0045] a target video generation module, configured to generate a target video according to at least the generated target segment and the clip template.
[0046] In a third aspect, the present disclosure further provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements steps in any of the above method embodiments when executing the computer program.
[0047] In a fourth aspect, the present disclosure also provides a computer readable storage medium. The computer readable storage medium has stored thereon a computer program which, when executed by a processor, implements the steps of any of the method embodiments described above.
[0048] In a fifth aspect, the present disclosure also provides a computer program product. The computer program product comprises a computer program which, when executed by a processor, implements the steps of any of the method embodiments described above.
[0049] In the above embodiments, in response to the to-be-processed video segment being a to-be-expanded video segment, the number of clusters is determined according to an adjustment time length matched with the to-be-expanded video segment, and clustering is performed according to the feature information of the to-be-expanded video segment, time information of each frame of video image in the to-be-expanded video segment, and the number of clusters, to obtain a plurality of video sub-segments. The number of clusters and the feature information can be used to accurately cluster the video images, and similar video images are classified into a category. In addition, the clusters formed by traditional clustering are regular in the feature space, but may be disconnected and irregular in time, which can lead to the formation of many small video segments (video images clustered into segments). Therefore, in the clustering process, time information can be used for clustering to ensure that the plurality of video sub-segments obtained after clustering are continuous in time. The video expansion mode between the plurality of video sub-segments is determined according to the image features of the plurality of video sub-segments, and an expanded video segment is generated, which can accurately expand the video segment, and the image features can be used to judge the degree of video change, so as to select a suitable video expansion mode, ensure the effect of video expansion, and obtain a smoother target video. In response to the to-be-processed video segment being a to-be-cut video segment, a first video style of the to-be-cut video segment is determined, and a target segment in the to-be-cut video segment is determined according to a relationship between an adjustment time length matched with the to-be-cut video segment, the first video style, and a second video style of the clip template. The determined first video style can provide information support for subsequent processing, thereby ensuring the accuracy of subsequent target segment determination. Since the style of the clip template also affects the identification of the target segment to some extent, the second video style matched with the clip template can be determined first, and the target segment in the to-be-cut video segment is determined according to the relationship between the time length threshold, the first video style, and the second video style. Thus, the target segment is adapted to the second video style of the clip template. Finally, at least according to the generated target segment and the clip template, a target video is generated, which can ensure the adaptation degree of the target segment and the clip template. BRIEF DESCRIPTION OF DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the specific embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0051] Figure 1 An application environment schematic diagram of the target video generation method in one embodiment;
[0052] Figure 2 A flowchart of the target video generation method in one embodiment;
[0053] Figure 3 A flowchart of the step S204 in one embodiment;
[0054] Figure 4 A schematic diagram of the clustering result in one embodiment;
[0055] Figure 5 A schematic diagram of the time overlap area in one embodiment;
[0056] Figure 6 A flowchart of determining the clustering boundary in one embodiment;
[0057] Figure 7 A flowchart after the step S302 in one embodiment;
[0058] Figure 8 A flowchart of a part of the step S208 in one embodiment;
[0059] Figure 9 A flowchart of another part of the step S208 in one embodiment;
[0060] Figure 10 A flowchart of still another part of the step S208 in one embodiment;
[0061] Figure 11 A structural schematic block diagram of the target video generation apparatus in one embodiment;
[0062] Figure 12 An internal structure schematic diagram of the computer device in one embodiment;
[0063] Figure 13 An internal structure schematic diagram of the computer device in one embodiment. DETAILED DESCRIPTION
[0064] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and not intended to limit the present disclosure.
[0065] It should be noted that the terms "first", "second" and the like in the description of the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product or apparatus including a series of steps or units does not necessarily limit to those clearly listed steps or units, but can include other steps or units not clearly listed or inherent to these processes, methods, products or apparatuses.
[0066] In this paper, the term "and / or" is only a description of the relationship between the associated objects, which means that there can be three relationships. For example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a "or" relationship.
[0067] The embodiments of the present disclosure provide a target video generation method, which can be applied to, for example Figure 1The application environment is shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The server 104 can determine the to-be-processed video segment in the plurality of video segments according to the position of the plurality of video segments input into the clip template in the terminal 102. The plurality of video segments can be pre-stored in the server 104 or the terminal 102. In response to the to-be-processed video segment being a to-be-expanded video segment, the server 104 can determine the cluster number according to the adjustment time length matched by the to-be-expanded video segment, and cluster according to the feature information of the to-be-expanded video segment, the time information of each frame of video image in the to-be-expanded video segment and the cluster number, to obtain a plurality of video sub-segments. The server 104 determines the video expansion mode between the plurality of video sub-segments according to the image features of the plurality of video sub-segments, and generates a target segment. Among them, the terminal 102 can be, but not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, panoramic cameras, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers. It should be noted that the present scheme can also be applied to the terminal 102 or the server 104 alone.
[0068] In one embodiment, as shown in Figure 2 , a target video generation method is provided. Taking the server 104 in Figure 1 as an example, the method includes the following steps:
[0069] S202, according to the position of the plurality of video segments input into the clip template, determine the to-be-processed video segment in the plurality of video segments, and determine the adjustment time length matched by each to-be-processed video segment.
[0070] The clip template refers to a preset format or design style used in the video editing process, helping editors or users quickly produce and arrange videos. The clip template usually includes video clips, transition effects, text styles, sound effects, etc. elements, which can improve the efficiency of creation and the consistency of the video. Many video editing software and platforms (such as Adobe Premiere Pro, Final Cut Pro, Filmora, etc.) provide various clip templates, which users can choose and adjust according to their needs and styles. Using clip templates can effectively save time. Multiple video clips are usually used in clip templates, and transition effects are added among the multiple video clips. The video duration of each video clip in the clip template can be a duration threshold. The video duration of each video clip in the clip template can be the same or different. The video clips to be processed can include video clips to be expanded and video clips to be cropped. The video clip to be cropped is a video clip with a video duration greater than the duration threshold. The video clip to be expanded can be a video clip that needs to expand the video duration in some embodiments of the present disclosure, i.e. expanding the duration of the video clip to the same duration threshold.
[0071] Specifically, according to the position of the multiple video clips input into the clip template, the duration threshold corresponding to each current video clip can be determined. Then, according to the duration threshold and the duration of the corresponding video clip, the video clip with a duration different from the duration threshold is determined. The video clip with a duration different from the duration threshold can be the video clip to be processed. When the duration of the video clip is greater than the duration threshold, the video clip can be the video clip to be cropped. When the duration of the video clip is less than the duration threshold, the video clip can be the video clip to be expanded. The absolute value of the difference between the duration of the video clip to be processed and the duration threshold can be used to determine the adjusted duration.
[0072] S204, in response to the video clip to be processed being a video clip to be expanded, determining the number of clusters according to the adjusted duration matched by the video clip to be expanded, and clustering according to the feature information of the video clip to be expanded, the time information of each frame of video image in the video clip to be expanded, and the number of clusters to obtain multiple video sub-clips.
[0073] The feature information can include color features, texture features, shape features, etc. The number of clusters can be the number of cluster centers. The feature information can be extracted from the video clip to be expanded by feature extraction.
[0074] Specifically, when the video segment to be processed is a video segment to be expanded, the number of clusters can be determined according to the adjustment time length matched with the video segment to be expanded. Generally, the longer the adjustment time length, the more the number of clusters. The time information corresponding to each frame of video image can be determined. Then the feature information and the time information in each frame of image are combined. Then the video images are clustered according to the combined information and the number of clusters, and a clustering result matched with the number of clusters is obtained. The video images are classified according to the clustering result, and thus a plurality of video sub-segments are obtained.
[0075] In some exemplary embodiments, the combined information can be clustered using a clustering algorithm such as K-means, and the value of K can be set according to the number of clusters during the clustering process.
[0076] S206, determining a video expansion mode between the plurality of video sub-segments according to the image features of the plurality of video sub-segments, and generating a target segment.
[0077] The image features can include optical flow, pixel value, brightness, etc. Optical flow is an important concept in computer vision and image processing, which refers to the motion information generated by the change of pixel brightness in two consecutive frames of images due to object motion or camera movement. Optical flow can be used to estimate the speed and direction of object motion in the scene. By analyzing images at different time points, the movement of objects on the image plane can be understood. The video expansion mode can generally be a processing mode between two video sub-segments, which can include, for example, cyclic alignment, transition effect, slow motion extension, and adding static segments. Cyclic alignment is a technique used in video expansion and analysis, mainly for processing video segments with cyclic patterns, i.e. playing the video segment in a loop. Such videos usually contain repeated actions or scenes. The transition effect can include fade-in and fade-out, blur transition, etc. Slow motion extension can be a slow motion processing of a video segment. Adding static segments can be inserting static pictures related to the video sub-segment.
[0078] Specifically, the optical flow between adjacent video sub-clips can be calculated. The degree of picture change between adjacent video sub-clips is determined according to the optical flow. For example, if the optical flow between two adjacent video sub-clips is small, it means that the two video sub-clips are both slow-paced videos, and then the video frames between the two video sub-clips can be supplemented by "loop alignment" to increase the video duration. If the optical flow between two adjacent video sub-clips is large, a transition effect can be added between the two video sub-clips to increase the video duration. It should be noted that the video expansion method can be flexibly selected according to the actual situation and the optical flow, and the specific selection of the video expansion method is not limited in some embodiments of the present disclosure.
[0079] Further, according to the optical flow of the plurality of video sub-clips, a video expansion method between the plurality of video sub-clips is determined, and the plurality of video sub-clips are processed according to the determined video expansion method. The processed plurality of video sub-clips are merged to generate a target clip.
[0080] For example, the optical flow of adjacent frame video images in adjacent video sub-clips can be calculated. Taking two video sub-clips as an example, the optical flow of the last frame video image in the first video sub-clip and the optical flow of the first frame video image in the second video sub-clip can be calculated, and then the optical flow difference between the optical flow of the last frame video image and the optical flow of the first frame video image is calculated. The video expansion method between the two video sub-clips is determined according to the optical flow difference. Then, the two video sub-clips are processed according to the determined video expansion method, and then the processed video sub-clips are merged, thereby increasing the duration of the video clip to be expanded and obtaining a target clip.
[0081] S208, in response to the video clip to be processed being a video clip to be cropped, determining a first video style of the video clip to be cropped, and determining a target clip in the video clip to be cropped according to a relationship between an adjusted duration matched with the video clip to be cropped, the first video style, and a second video style of the clip template.
[0082] Specifically, when the user determines the clip template, each clip template will have a plurality of corresponding labels, such as landscape, sports, happy, and various labels. The second video style matched with the clip template can be determined according to the label corresponding to the clip template, and then the target clip related to the second video style is screened from the video clip to be cropped according to the relationship between the first video style and the second video style and using the duration threshold. In addition, the first video style of the video clip to be cropped can be determined by visual style analysis or using a pre-trained neural network model.
[0083] S210: Generate a target video at least according to the generated target segment and the editing template.
[0084] Specifically, the to-be-processed video segment among the multiple video segments may be replaced with the target segment, thereby generating the target video according to the editing template.
[0085] In the target video generation method described above, in response to the video segment being processed being a segment to be expanded, the number of clusters is determined based on the adjusted duration matching the segment to be expanded. Clustering is then performed based on the feature information of the segment to be expanded, the temporal information of each frame of the video image within the segment to be expanded, and the number of clusters to obtain multiple video sub-segments. Using the number of clusters and the feature information, video images can be accurately clustered, classifying similar video images into categories. Furthermore, while clusters formed by traditional clustering are regular in feature space, they may be disconnected and irregular in time, resulting in the formation of many small video segments (clustered segments of video images). Therefore, temporal information can be used during clustering to ensure temporal continuity among the multiple video sub-segments obtained after clustering. The video expansion method between the multiple video sub-segments is determined based on the image features of the multiple video sub-segments to generate the expanded video segment. This allows for accurate expansion of the video segment. Furthermore, the image features can be used to determine the severity of video changes, thereby selecting an appropriate video expansion method to ensure the video expansion effect and produce a smoother target video. In response to the fact that the video segment to be processed is a video segment to be cropped, the first video style of the video segment to be cropped is determined, and the target segment in the video segment to be cropped is determined based on the relationship between the adjusted duration matching the video segment to be cropped, the first video style, and the second video style of the editing template. The determined first video style can provide information support for subsequent processing, thereby ensuring the accuracy of subsequent target segment determination. Since the style of the editing template will also have a certain degree of influence on the identification of the target segment, the second video style matching the editing template can be determined first, and the target segment in the video segment to be cropped can be determined based on the duration threshold, the relationship between the first video style and the second video style. In this way, the target segment is adapted to the second video style of the editing template. Finally, based on at least the generated target segment and the editing template, a target video is generated, which can ensure the degree of adaptation between the target segment and the editing template.
[0086] In one embodiment, Figure 3 As shown, clustering is performed based on the feature information of the video segment to be expanded, the time information of each frame of the video image in the video segment to be expanded, and the number of clusters to obtain multiple video sub-segments, including:
[0087] S302, cluster the video images according to the feature information in each frame of video image, the time information corresponding to the each frame of video image, and the number of clusters, to obtain a clustering result matching the number of clusters.
[0088] Specifically, the time information corresponding to each frame of video image can be determined. Then the feature information in each frame of image and the time information are combined. Then the video images are clustered according to the combined information and the number of clusters, to obtain a clustering result matching the number of clusters.
[0089] S304, determine the clustering boundary between the clustering results, segment the video segment to be expanded according to the clustering boundary, to obtain a plurality of video sub-segments.
[0090] Wherein, the clustering boundary refers to the boundary or limit used to separate different clusters in clustering analysis. As shown in Figure 4 , the clustering result is two categories, which are A clustering result and B clustering result. Wherein, the clustering boundary in A clustering result can be S1, and the clustering boundary in B clustering result can be S2.
[0091] Specifically, after determining the clustering boundary, the feature information contained in the clustering boundary can be determined, the video image corresponding to the specific information is determined, and the determined video image is divided into a segment. In this way, a plurality of video sub-segments are obtained.
[0092] In this embodiment, by using the clustering boundary, the video segment to be expanded can be accurately segmented, and similar video images can be divided into a segment, so as to ensure the fluency of video expansion.
[0093] In one embodiment, the determination of the clustering boundary between the clustering results comprises:
[0094] In response to the existence of the time overlap region in the clustering result, the clustering boundary between the clustering results is determined based on the time point in the time overlap region.
[0095] Specifically, the time overlap region can be a region that overlaps in time but belongs to different clustering results. As shown in Figure 5 , taking A clustering result and B clustering result as an example, A clustering result and B clustering result overlap in the time X1 to time X2 segment. Therefore, the clustering result located in the X1 and X2 time segment can be a time overlap region.
[0096] Specifically, when there is a time overlap region in the clustering result, the clustering result needs to be divided according to the time point in the time overlap region, so as to determine the clustering boundary between the clustering results.
[0097] As Figure 6 illustrated, the determining the clustering boundary between the clustering results based on the time point in the time overlap region comprises:
[0098] S402, determining a time median point in the time overlap region, and determining the clustering boundary between the clustering results according to the time median point.
[0099] Specifically, continue to take the example as Figure 5 illustrated, the time median point can be determined according to the time period corresponding to the time overlap region (X1-X2 period in Figure 5 ), and the time median point is used to distinguish the A clustering result and the B clustering result. The feature information located before the time median point can be divided into the A clustering result, and the feature information located after the time median point can be divided into the B clustering result.
[0100] Alternatively, S404, determining the total number of feature information contained in the time overlap region, and determining the first number of clustering results existing in the same time overlap region.
[0101] S406, determining a time point from the time overlap region, which divides the total number by the first number, and determining the clustering boundary between the clustering results based on the time point.
[0102] Specifically, the total number of feature information contained in the time overlap region and the first number of clustering results existing in the same time overlap region can be determined. The number of feature information corresponding to each time point corresponding to the time overlap region can also be determined. The total number is divided by the first number, and then the time point dividing the total number is determined according to the number of feature information corresponding to each time point, which can be the clustering boundary.
[0103] In some exemplary embodiments, for example, the total number of feature information contained in the time overlap region is 120, and the clustering results containing the time overlap region are two. There are five time points in the time overlap region, and the number of feature information corresponding to each time point is 10, 30, 20, 50, and 10 respectively, so that the third time point can be determined as the clustering boundary, and there are 60 feature information before the third time point and 60 feature information after the third time point. It can be understood that the above is only used for example.
[0104] In the embodiment, the clustering results can be accurately divided by the time points in the time overlap region, so as to ensure the accuracy of subsequent processing. The time boundary in the time overlap region can be quickly determined by using the time median point for division, so as to improve the processing speed. In addition, the total number of feature information contained in the time overlap region can be determined, and the time point that divides the total number in half can be determined from the time overlap region. Based on the time point, the clustering boundary between the clustering results can be determined. The determined time point can divide the feature information contained in the time overlap region in half, so as to ensure that the clustering features on both sides of the clustering boundary in the time overlap region are basically the same, and the accuracy of the feature information in each clustering result is ensured. The problem that the feature information is too much or too little due to uneven distribution of feature information in multiple clustering results caused by the time overlap region does not occur.
[0105] In one embodiment, as shown in Figure 7 After obtaining the clustering results matching the number of clusters, the method further includes:
[0106] S502, calculating the silhouette coefficient of the clustering results.
[0107] S504, in response to the silhouette coefficient being less than a preset coefficient threshold, decomposing the feature information by using principal component analysis to obtain feature information of multiple dimensions.
[0108] The silhouette coefficient is an index for evaluating the quality of clustering results, reflecting the similarity of data points to their own clusters and the similarity to other clusters. The value of the silhouette coefficient is between -1 and 1, which can help to judge the effect of clustering. Principal component analysis (PCA) is a commonly used dimension reduction technique, which is used to map high-dimensional data to low-dimensional space while preserving as much original information as possible. In general, the coefficient threshold can be 0.5.
[0109] Specifically, the silhouette coefficient of each clustering result can be calculated. Then the silhouette coefficient is compared with the preset coefficient threshold. When the silhouette coefficient is less than the coefficient threshold, it can be determined that the clustering quality is poor, indicating that the feature points are not suitable for clustering into multiple classes in the existing feature space, for example, random noise data cannot be clustered into multiple classes. Therefore, for all feature information, let their feature vector latitude be n, and perform PCA decomposition on them to obtain new feature information of k latitudes.
[0110] S506, combining the feature information of multiple dimensions and the time information corresponding to each frame of image to obtain a modified clustering feature group.
[0111] S508, clustering according to the modified clustering feature group and the number of clusters to obtain a modified clustering result.
[0112] Specifically, the new feature information and the time can be combined to obtain a new clustering feature group, which can be the modified clustering feature group. Then, clustering is performed according to the modified clustering feature group and the number of clusters to obtain a modified clustering result. For the specific embodiments, refer to the above embodiments, which will not be repeated here. Then, the steps mentioned in the above embodiments can be performed again according to the modified clustering result.
[0113] In this embodiment, due to the maximum variance theory of the PCA algorithm and the characteristic of reducing noise interference, the feature vector distance between the new feature points is larger than before, and it is easier to find the internal difference of the data and cluster into multiple categories, thereby ensuring the accuracy of clustering.
[0114] In one embodiment, the clustering of the video images according to the feature information in each frame of video images, the time information corresponding to each frame of video images, and the number of clusters to obtain a clustering result matching the number of clusters comprises:
[0115] Combining the feature information in each frame of video images and the time information corresponding to each frame of video images to obtain a clustering feature group.
[0116] Clustering according to the clustering feature group and the number of clusters to obtain a clustering result.
[0117] Specifically, the feature information in each frame of video images and the corresponding time can be combined to form a new feature group, and the new feature group and the number of clusters are used for clustering calculation to obtain a final clustering result.
[0118] In some exemplary embodiments, the wheel disc algorithm can be used for clustering when clustering, for example, a feature point in the feature group can be randomly selected as the first clustering center c1; the shortest distance between each feature point in the feature group and the existing clustering center (i.e., the distance to the nearest clustering center) is calculated, denoted as D(x); the larger the value, the greater the probability of being selected as a clustering center; the next clustering center is selected by the wheel disc method. Repeat the above steps until the same number of clustering centers as the number of clusters are selected.
[0119] In other exemplary embodiments, the following formula can be used to iteratively adjust the clustering center:
[0120]
[0121] Where V represents the total variance or error measure of clustering, which is usually used to evaluate the tightness of clustering. The smaller the value, the better the clustering effect. K represents the number of clusters, that is, the number of different clusters into which the data is divided. i represents the set of all data points in the i-th clustering result. j represents the j-th data point belonging to the clustering Si. i represents the center (mean) of the i-th cluster, that is, the average value of all data points in the clustering Si. tj is the collection time of the j-th data; tistart is the earliest collection time of all data in the i-th class, that is, the start collection time of the class. j ∈[t istart , t end ] represents that for any feature point xj in the i-th class Si, its collection time must be in the collection time interval of this class, that is, its collection time is not in other classes, that is, the collection time between the data in each class is continuous.
[0122] In this embodiment, by combining feature information and time information, the image collection time corresponding to all feature points in the same class is continuous, and clustering can be better performed.
[0123] In one embodiment, as shown in Figure 8 , the determining the first video style of the to-be-cut video segment comprises:
[0124] S602, color extraction is performed on each frame of video image in the to-be-cut video segment to obtain color information.
[0125] S604, a color atlas of each frame of video image is constructed according to the color information, and a scene style of each frame of video image is determined according to the color usage frequency and color combination indicated by the color atlas.
[0126] Wherein, color extraction is a technology in the field of computer vision and image processing, which aims to identify and extract the color information contained in the image or video. Color information includes: saturation, brightness, color distribution and other information. Color atlas (Color Atlas) is a collection of color samples of each frame of video image.
[0127] Specifically, color extraction is performed on each frame of video image to obtain color information. Color space conversion: each frame of video image is converted from BGR / RGB to a color space more suitable for analysis (such as HSV). The color histogram of each frame of video image is calculated to obtain color information. The extracted color information is organized into a color atlas. The color information of each frame of video image can be stored using a dictionary or a data frame (such as Pandas DataFrame). For the color atlas of each frame, the frequency of use of different color values is calculated. Clustering algorithms (such as K-means) can be used to cluster colors to identify the main colors. Color combination features (such as primary color, secondary color) are extracted and analyzed. Pre-trained models or machine learning methods can also be used to authenticate scene style. Certain color combinations often correspond to a certain style, such as low saturation and high contrast, which may indicate a retro or artistic style.
[0128] S606, according to the changes between each frame of video image in the to-be-cut video segment, determine the camera movement style between each frame of video image in the to-be-cut video segment.
[0129] Specifically, according to the changes between each frame of video image, the moving way of the shot (such as pan, push-pull, shake, etc.) can be recorded and classified. Fast and unstable motion may indicate dynamic or tense style, while slow and smooth shot may belong to narrative or lyrical style.
[0130] S608, according to the camera movement style between each frame of video image and the scene style of each frame of video image, determine the first video style of the to-be-cut video segment.
[0131] Specifically, the camera movement style and scene style of each frame can be encoded for subsequent analysis. Different camera movement ways can be represented by numerical values or category labels. For example: pan: 0, rotation: 1, zoom: 2, still: 3. According to color combination, use the corresponding label to represent. Primary color and secondary color or clustering label can be used. The camera movement style and scene style of each frame are integrated into a data structure, such as a list or a data frame. Analyze the style combination of each frame in the entire video segment, find the most common combination and summarize it. Count the frequency of each camera movement style and scene style combination. Based on the frequency, select the camera movement style and scene style combination with the highest frequency. According to the camera movement style and scene style combination with the highest frequency, determine the first video style of the to-be-cut video segment.
[0132] In this embodiment, through color analysis, the scene style of each frame of video image is determined, and through the changes between each frame of video image, the camera movement style of each frame of video image is determined, and according to the scene style and the camera movement style, the first video style can be accurately determined, thereby ensuring the accuracy of the target segment determination.
[0133] In an embodiment, the determining the first video style of the to-be-cut video segment according to the camera movement style between each frame of video image and the scene style of each frame of video image comprises:
[0134] determining the image style of each frame of video image according to the camera movement style between each frame of video image and the scene style of each frame of video image.
[0135] determining the style proportion in the to-be-cut video segment according to the image style of each frame of video image.
[0136] determining the first video style of the to-be-cut video segment according to the image style with the largest proportion of style proportion.
[0137] The image style refers to the characteristics and artistic effects embodied in the visual performance of an image, including the uniqueness of color, texture, shape, line, composition, etc. The image style can include landscape style, sports style, portrait style, etc.
[0138] Specifically, in order to accurately determine the first video style of the to-be-cut video segment, each frame of video image in the to-be-cut video segment can be processed individually. For each frame of video image, the image style of the current frame of video image can be determined according to its corresponding camera movement style and scene style. Then the image style of each frame of video image in the to-be-cut video segment is counted, the video images with the same image style are classified into a category, then the number of each category of video image is determined, and the style proportion of each category of video image is determined according to the number of each category of video image and the total number of video images. According to the style proportion of each category of video image, the image style with the largest proportion of style proportion is selected as the first video style.
[0139] In this embodiment, the image style of each frame of video image is used to determine the style proportion, and the style proportion is used to determine the first video style, which can accurately find the first video style that best fits the to-be-cut video segment, thereby improving the accuracy of target segment determination.
[0140] In an embodiment, as shown in Figure 9 determining the target segment in the to-be-cut video segment according to the relationship between the adjustment duration matched with the to-be-cut video segment, the first video style and the second video style of the clip template comprises:
[0141] S702, in response to the second video style being the same as the first video style, determining a first material area in each frame of video image of the to-be-cut video segment that has a matching relationship with the first video style.
[0142] The first material region can include a person, various scenery (e.g., mountains, water, trees, buildings, etc.), and an object (e.g., a car, a bicycle, etc.).
[0143] Specifically, when the second video style is the same as the first video style, material recognition can be performed on each frame of the video image in the video clip to be cropped to identify the material region in each frame of the video image. Then, the material region is filtered according to the first video style to determine the first material region.
[0144] In some example embodiments, taking a frame of video image as an example, it is identified that the material region in the current frame of video image includes an A1 person material region, an A2 person material region, a B1 scenery material region, a B2 scenery material region, and a C1 object material region. If the current first video style is a sports style (a person is photographed performing sports), the first material region can be the A1 person material region and the A2 person material region. It can be understood that the above is only used for example illustration.
[0145] S704, determining a style score of each frame of video image according to the first material region, and determining a target segment in the video clip to be cropped according to the style score of each frame of video image and the time threshold.
[0146] Specifically, the style score of each frame of video image can be determined according to the proportion of the first material region, the area of the first material region, and the like. After determining the style score of each frame of video image, each frame of video image can be sorted according to the style score, and the frame of video image with a larger style score can be placed in front. Then, a plurality of frames of video image are selected according to the time threshold to determine the target segment.
[0147] In some example embodiments, generally, the content of a video does not suddenly change every moment during video shooting, and therefore, the frames of video image with similar style scores are generally continuous. If the frame rate of the video is 60 frames / s and the time threshold is 3 s, 180 frames of video image can be selected as the target segment according to the style score. The 180 frames of video image can be one or more continuous frames of video image.
[0148] S706, in response to the second video style being different from the first video style, determining a second material region in each frame of video image in the video clip to be cropped that has a matching relationship with the second video style.
[0149] Specifically, when the second video style is different from the first video style, since the generated video is generated by using the clip template, it is necessary to determine the second material region in each frame of video image which has a matching relationship with the second video style. It should be noted that, generally, the to-be-cut video segment involved in some embodiments of the present disclosure can be a panoramic video, and therefore each frame of video image in the to-be-cut video segment contains more material information, and generally the second material region can be identified. If the second material region is not identified, the second material region of the frame of video image can be determined as 0, and the style score of the frame of video image is also 0.
[0150] S708, determining a style score of each frame of video image according to the second material region, and determining a target segment in the to-be-cut video segment according to the style score of each frame of video image and the time length threshold.
[0151] Specifically, the style score of each frame of video image can be determined according to the second material region in each frame of video image. After determining the style score of each frame of video image, each frame of video image can be sorted according to the style score, and the frame of video image with a larger style score can be arranged in front. Then, a plurality of frames of video image are selected according to the time length threshold, so as to determine the target segment. (For specific implementation, refer to the above S704 step, which will not be repeated here.)
[0152] In this embodiment, by comparing the second video style and the first video style, when they are the same, the style score of each frame of video image is determined according to the first material region, so as to determine the target segment. When they are different, the style score of each frame of video image is determined according to the second material region, so as to determine the target segment. It can be ensured that the final target segment is more suitable for the clip template.
[0153] In one embodiment, the determining of the style score of each frame of video image according to the first material region comprises at least one of the following:
[0154] determining the style score of each frame of video image according to the number of the first material region;
[0155] determining the style score of each frame of video image according to the proportion of the first material region in the video image.
[0156] Specifically, the more the number of the first material region in a frame of video image, the larger the style score of the frame of video image. The larger the proportion of the first material region in the entire video image in a frame of video image, the larger the style score of the frame of video image.
[0157] In one embodiment, as Figure 10As shown, when the first video style is a sports style. When the first video style includes a sports style. Wherein the sports style refers to a visual style and performance method, which is usually used to convey the feeling of dynamic, vitality and passion. This style is common in sports, extreme sports, dance, fitness videos or any video content that emphasizes action and activity. According to the relationship between the time length threshold, the first video style and the second video style, the target segment in the video segment to be cropped is determined, including:
[0158] S802, in response to the second video style being the same as the first video style, determining the pose information corresponding to each frame of video image in the video segment to be cropped.
[0159] Wherein, the pose information can be IMU information. IMU information can be obtained by IMU measurement. IMU (Inertial Measurement Unit) is a device used to measure and report the acceleration and angular velocity of an object. IMU is usually composed of multiple sensors, including accelerometers, gyroscopes, and sometimes magnetometers. Its main functions and components are as follows: Accelerometer: measures the linear acceleration of an object in three axes (X, Y, Z). Through acceleration data, the motion state and direction change of the object can be perceived. Gyroscope: measures the angular velocity of the object (i.e. the rotation speed of the object around its axis), so that it can detect the direction change of the object. Magnetometer (optional): measures the direction of the earth's magnetic field, which can be used to assist in determining the spatial orientation of the object, especially in the orientation of the object.
[0160] Specifically, when the second video style is the same as the first video style, it can be determined that the style of the clip template adaptation is also a sports style. At this time, the IMU information generated by the device in the process of shooting each frame of video image can be obtained.
[0161] S804, according to the pose information corresponding to each frame of video image, determining the change amplitude of the pose information between each frame of video image.
[0162] Specifically, since the IMU information reflects various poses of the device, the change amplitude of the pose information between each frame of video image can be determined according to the IMU information corresponding to each frame of video image, and the change amplitude between each frame of video image can be determined according to the change amplitude of the pose information.
[0163] S806, sorting each frame of video image according to the change amplitude, and selecting at least one frame of video image that meets the time length threshold from the sorted video images according to the time length threshold.
[0164] Specifically, since it is a sports style, the more drastic the change in the picture, the more intense the movement can be determined, and the probability of a highlight scene appearing will be greatly increased. Therefore, each frame of video image can be sorted according to the change amplitude, and the video image with large change amplitude can be arranged in the front position. Then, at least one frame of video image satisfying the time length threshold can be obtained by screening the sorted video image according to the time length threshold.
[0165] In S808, a target segment in the to-be-cut video segment is determined according to at least one frame of video image satisfying the time length threshold.
[0166] In some exemplary embodiments, ten frames of video images are taken as an example for illustration. The change amplitude of the posture information between the fifth frame and the sixth frame of video images is the largest, and the change amplitudes between the remaining frames of video images are the same. Therefore, the fifth frame and the sixth frame of video images can be arranged in the front position. If 4 frames of video images need to be selected according to the time length threshold. However, the change amplitudes between the remaining frames of video images are the same, since the fifth frame of video image and the sixth frame of video image have been determined, in order to ensure the fluency of the determined target segment, two frames of video images can be selected near the fifth frame of video image and the sixth frame of video image, and the fourth frame of video image and the seventh frame of video image can be selected. The fourth frame, the fifth frame, the sixth frame and the seventh frame of video images can be taken as the target segment.
[0167] In the present embodiment, when the first video style is a sports style, since the sports style will move quickly or adjust the direction of the device during shooting, the posture information corresponding to each frame of video image can be determined, and the change amplitude of the video image in the shooting process can be determined by using the posture information, so that the image with intense movement can be accurately determined, and the accuracy of the target segment determination can be ensured.
[0168] It should be understood that, although each step in the flowchart involved in each of the above-described embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each of the above-described embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0169] Based on the same inventive concept, the embodiments of the present disclosure further provide a target video generation apparatus for implementing the target video generation method as described above. The implementation scheme for solving the problem provided by the apparatus is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more target video generation apparatus embodiments provided below can refer to the limitations of the target video generation method described above, which will not be described here again.
[0170] In one embodiment, as shown in Figure 11 A target video generation apparatus 900 is provided, comprising a video segment determination module 802, a clustering processing module 804, an expansion module 806, a target segment determination module 808 and a target video generation module 810, wherein:
[0171] The video segment determination module 802 is configured to determine a to-be-processed video segment from a plurality of video segments according to a position of the plurality of video segments input into a clip template, and determine an adjustment duration matched with each to-be-processed video segment;
[0172] The clustering processing module 804 is configured to, in response to the to-be-processed video segment being a to-be-expanded video segment, determine a clustering number according to the adjustment duration matched with the to-be-expanded video segment, and perform clustering according to feature information of the to-be-expanded video segment, time information of each frame of video image in the to-be-expanded video segment and the clustering number, to obtain a plurality of video sub-segments;
[0173] The expansion module 806 is configured to determine a video expansion mode between the plurality of video sub-segments according to image features of the plurality of video sub-segments, and generate a target segment;
[0174] The target segment determination module 808 is configured to, in response to the to-be-processed video segment being a to-be-cut video segment, determine a first video style of the to-be-cut video segment, and determine a target segment in the to-be-cut video segment according to a relationship between the adjustment duration matched with the to-be-cut video segment, the first video style and a second video style of the clip template;
[0175] The target video generation module 810 is configured to generate a target video according to at least the generated target segment and the clip template.
[0176] In one embodiment of the apparatus, the clustering processing module 804 comprises:
[0177] A clustering module is configured to cluster the video images according to the feature information in each frame of video image, the time information corresponding to each frame of video image and the clustering number, to obtain a clustering result matched with the clustering number;
[0178] a segmenting module configured to determine clustering boundaries between the clustering results, and segment the video segment to be augmented according to the clustering boundaries to obtain a plurality of video sub-segments.
[0179] In one embodiment of the apparatus, the segmenting module comprises:
[0180] a clustering boundary determining module configured to, in response to the existence of a time overlap region in the clustering results, determine the clustering boundaries between the clustering results based on a time point in the time overlap region.
[0181] The clustering boundary determining module is further configured to determine a time median point in the time overlap region, and determine the clustering boundaries between the clustering results according to the time median point; or determine a total number of feature information contained in the time overlap region, and determine a first number of the clustering results that exist in the same time overlap region; determine a time point from the time overlap region that divides the total number by the first number, and determine the clustering boundaries between the clustering results based on the time point.
[0182] In one embodiment of the apparatus, the apparatus further comprises a dimension augmenting module configured to calculate profile coefficients of the clustering results; and in response to the profile coefficients being less than a preset coefficient threshold, decompose the feature information by using a principal component analysis method to obtain feature information of a plurality of dimensions.
[0183] The clustering module is further configured to combine the feature information of the plurality of dimensions and time information corresponding to each frame of video image to obtain a modified clustering feature group; and cluster according to the modified clustering feature group and the clustering number to obtain a modified clustering result.
[0184] In one embodiment of the apparatus, the target segment determining module 808 comprises:
[0185] a color extracting module configured to extract colors of each frame of video image in the video segment to be cropped to obtain color information;
[0186] a scene style determining module configured to construct a color spectrum of each frame of video image according to the color information, and determine a scene style of each frame of video image according to a usage frequency of colors and a combination of colors indicated by the color spectrum;
[0187] a camera movement style determining module configured to determine a camera movement style between each frame of video image in the video segment to be cropped according to a change between each frame of video image in the video segment to be cropped;
[0188] a video style determination module, configured to determine a first video style of the video segment to be cropped according to the camera movement style between each frame of video images and a scene style of each frame of video images.
[0189] In an embodiment of the apparatus, the target segment identification module comprises:
[0190] a first material determination module, configured to, in response to the second video style being the same as the first video style, determine a first material region in each frame of video images of the video segment to be cropped that has a matching relationship with the first video style.
[0191] a target identification module, configured to determine a style score of each frame of video images according to the first material region, and determine a target segment in the video segment to be cropped according to the style score of each frame of video images and the time threshold.
[0192] a second material determination module, configured to, in response to the second video style being different from the first video style, determine a second material region in each frame of video images of the video segment to be cropped that has a matching relationship with the second video style.
[0193] The target identification module is further configured to determine a style score of each frame of video images according to the second material region, and determine a target segment in the video segment to be cropped according to the style score of each frame of video images and the time threshold.
[0194] In an embodiment of the apparatus, when the first video style is a motion style, the target segment identification module is further configured to, in response to the second video style being the same as the first video style, determine pose information corresponding to each frame of video images in the video segment to be cropped; determine a change amplitude of the pose information between each frame of video images according to the pose information corresponding to each frame of video images; sort each frame of video images according to the change amplitude, and filter the sorted video images according to the time threshold to obtain at least one frame of video image satisfying the time threshold; and determine a target segment in the video segment to be cropped according to the at least one frame of video image satisfying the time threshold.
[0195] The above modules in the target video generation apparatus can be realized by software, hardware, or a combination thereof. The above modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the above modules.
[0196] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 12As shown in the figure. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store video segment data. The network interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement a target video generation method.
[0197] In one embodiment, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in the figure. Figure 13 As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. Wireless mode can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement a target video generation method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0198] Those skilled in the art can understand that, Figure 11 The structure shown in the figure is only a block diagram of part of the structure related to the present disclosure, and does not constitute a limitation on the computer device to which the present disclosure is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0199] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the above method embodiments.
[0200] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by the processor to implement the steps in the above method embodiments.
[0201] In one embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the above method embodiments.
[0202] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned method embodiments. Any reference to memory, database or other medium used in each embodiment provided by the present disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in each embodiment provided by the present disclosure can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in each embodiment provided by the present disclosure can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0203] Each technical feature of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of each technical feature in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present disclosure.
[0204] The above-described embodiments are merely illustrative of several embodiments of the present disclosure, which are described in a relatively specific and detailed manner, but should not be construed as limiting the scope of the patent of the present disclosure. It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present disclosure, and these all belong to the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the appended claims.
Claims
1. A target video generation method, characterized by, The method comprises: determining a video segment to be processed from a plurality of video segments according to positions of the plurality of video segments input into a clip template, and determining an adjusted time length matched to each of the video segment to be processed; in response to the video segment to be processed being a video segment to be expanded, determining a cluster number according to the adjusted time length matched to the video segment to be expanded, and clustering according to feature information of the video segment to be expanded, time information of each frame of video image in the video segment to be expanded, and the cluster number to obtain a plurality of video sub-segments; determining a video expansion mode between the plurality of video sub-segments according to image features of the plurality of video sub-segments, and generating a target segment; in response to the video segment to be processed being a video segment to be cropped, determining a first video style of the video segment to be cropped, and determining a target segment in the video segment to be cropped according to a relationship between the adjusted time length matched to the video segment to be cropped, the first video style, and a second video style of the clip template; generating a target video according to at least the generated target segment and the clip template.
2. The method of claim 1, wherein, The clustering according to the feature information of the video segment to be expanded, the time information of each frame of video image in the video segment to be expanded, and the cluster number to obtain a plurality of video sub-segments comprises: clustering the video image according to the feature information in each frame of video image, the time information corresponding to each frame of video image, and the cluster number to obtain a cluster result matched to the cluster number; determining a cluster boundary between the cluster results, and segmenting the video segment to be expanded according to the cluster boundary to obtain a plurality of video sub-segments.
3. The method of claim 2, wherein, The determining of the cluster boundary between the cluster results comprises: in response to a time overlap region existing in the cluster results, determining the cluster boundary between the cluster results based on a time point in the time overlap region; The determining of the cluster boundary between the cluster results based on the time point in the time overlap region comprises: determining a time median point in the time overlap region, and determining the cluster boundary between the cluster results according to the time median point; or, determining a total number of feature information contained in the time overlap region, and determining a first number of cluster results existing in the same time overlap region; determining a time point from the time overlap region, which divides the total number by the first number, and determining the cluster boundary between the cluster results based on the time point.
4. The method of claim 2, wherein, After the obtaining of the cluster result matched to the cluster number, the method further comprises: calculating a contour coefficient of the cluster result; in response to the contour coefficient being less than a preset coefficient threshold, decomposing the feature information by using a principal component analysis method to obtain feature information in a plurality of dimensions; combining the feature information in the plurality of dimensions and the time information corresponding to each frame of video image to obtain a modified cluster feature group; clustering according to the modified cluster feature group and the cluster number to obtain a modified cluster result.
5. The method of claim 1, wherein, The determining of the first video style of the video segment to be cropped comprises: extracting color information from each frame of the video image in the to-be-cut video segment; constructing a color atlas of each frame of the video image according to the color information, and determining a scene style of each frame of the video image according to a usage frequency of the color indicated by the color atlas and a combination of the color; determining a camera movement style between each frame of the video image in the to-be-cut video segment according to a change between each frame of the video image in the to-be-cut video segment; determining a first video style of the to-be-cut video segment according to the camera movement style between each frame of the video image and the scene style of each frame of the video image.
6. The method of claim 1, wherein, The determining of the target segment in the to-be-cut video segment according to the relationship between the adjustment time length matched with the to-be-cut video segment, the first video style and a second video style of the clip template comprises: in response to the second video style being the same as the first video style, determining a first material area in each frame of the video image in the to-be-cut video segment that has a matching relationship with the first video style; determining a style score of each frame of the video image according to the first material area, and determining the target segment in the to-be-cut video segment according to the style score of each frame of the video image and a time length threshold value, the time length threshold value being a video time length of each video segment in the clip template; in response to the second video style being different from the first video style, determining a second material area in each frame of the video image in the to-be-cut video segment that has a matching relationship with the second video style; determining a style score of each frame of the video image according to the second material area, and determining the target segment in the to-be-cut video segment according to the style score of each frame of the video image and the time length threshold value.
7. The method of claim 1, wherein, When the first video style is a motion style, the determining of the target segment in the to-be-cut video segment according to the relationship between the adjustment time length matched with the to-be-cut video segment, the first video style and the second video style of the clip template comprises: in response to the second video style being the same as the first video style, determining pose information corresponding to each frame of the video image in the to-be-cut video segment; determining a change amplitude of the pose information between each frame of the video image according to the pose information corresponding to each frame of the video image; sorting each frame of the video image according to the change amplitude, and screening the sorted video images according to a time length threshold value to obtain at least one frame of the video image satisfying the time length threshold value, the time length threshold value being the video time length of each video segment in the clip template; determining the target segment in the to-be-cut video segment according to the at least one frame of the video image satisfying the time length threshold value.
8. A target video generation apparatus characterized by comprising: The apparatus comprises: a video segment determination module configured to determine a to-be-processed video segment in a plurality of video segments and determine an adjustment time length matched with each to-be-processed video segment according to a position of the to-be-processed video segment input into a clip template; The clustering processing module is configured to, in response to the video segment to be processed being a video segment to be expanded, determine a number of clusters according to an adjusted time length matched by the video segment to be expanded, and perform clustering according to feature information of the video segment to be expanded, time information of each frame of video image in the video segment to be expanded, and the number of clusters, to obtain a plurality of video sub-segments. The expansion module is configured to determine a video expansion manner between the plurality of video sub-segments according to image features of the plurality of video sub-segments, and generate a target segment. The target segment determination module is configured to, in response to the video segment to be processed being a video segment to be cropped, determine a first video style of the video segment to be cropped, and determine a target segment in the video segment to be cropped according to a relationship between an adjusted time length matched by the video segment to be cropped, the first video style, and a second video style of the clip template. The target video generation module is configured to generate a target video according to at least the generated target segment and the clip template. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 7.
Citation Information
Patent Citations
Video synthesis method and device, electronic equipment and storage medium
CN112291484A
Video editing method and device, electronic equipment and storage medium
CN113301430A