Video editing method, device, equipment, storage medium and computer program

By sampling and analyzing the attribute information of video footage, and using the server to generate highlight and regular video segments, the problem of users having difficulty editing aesthetically pleasing videos is solved, and expressive video clips are generated efficiently and automatically.

CN119094858BActive Publication Date: 2025-10-28HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310666211.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-06
Publication Date
2025-10-28
Estimated Expiration
2043-06-06

AI Technical Summary

Technical Problem

Users often struggle to master video editing skills in a short period, resulting in an inability to create aesthetically pleasing and expressive finished videos.

Method used

By sampling multiple clips and obtaining the attribute information of the sampled frames, the server determines the target highlight template and highlight shot, generates highlight video segments and regular video segments, and finally automatically generates the final video.

Benefits of technology

It allows users to generate visually appealing and expressive videos without needing any editing knowledge, thus improving video editing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119094858B_ABST
    Figure CN119094858B_ABST
Patent Text Reader

Abstract

This application discloses a video editing method, apparatus, device, storage medium, and computer program, belonging to the field of media technology. The method includes: sampling multiple editing materials to obtain multiple sample frames; determining the attribute information of each sample frame; sending the attribute information of the multiple sample frames to a server; receiving a target highlight template and highlight sample frames corresponding to multiple highlight shots sent by the server; determining regular video segments and highlight video segments corresponding to the multiple editing materials based on the attribute information of the multiple editing materials, the multiple sample frames, the target highlight template, and the highlight sample frames corresponding to the multiple highlight shots; and generating a final video clip obtained after video editing of the multiple editing materials based on the regular video segments and highlight video segments. The video editing method provided by this application can automatically generate a final video clip, effectively improving the efficiency of video editing. Furthermore, the final video clip also possesses rich aesthetic appeal and expressiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of media technology, and in particular to a video editing method, apparatus, device, storage medium, and computer program. Background Technology

[0002] With the rapid development of multimedia and the Internet, users have increasingly higher requirements for the quality of the videos they shoot. They no longer just want to record their lives through videos, but want to showcase videos that are more aesthetically pleasing and expressive.

[0003] In related technologies, users can manually cut, merge, and recombine video footage through video editing to generate finished videos with different expressive qualities.

[0004] However, video editing is a specialized field, and learning and mastering how to use video editing software usually takes a lot of time. Users often cannot master video editing knowledge in a short period of time, making it difficult to edit high-quality segments from video footage, thus preventing users from showcasing more aesthetically pleasing and expressive videos. Summary of the Invention

[0005] This application provides a video editing method, apparatus, device, storage medium, and computer program, which can solve the problem of difficulty in generating expressive video clips in related technologies. The technical solution is as follows:

[0006] In a first aspect, a video editing method is provided, applied to a terminal. The method includes: sampling multiple editing materials to obtain multiple sample frames, the editing materials including videos or images; determining attribute information for each sample frame, the attribute information indicating the characteristics of the sample frame; sending the attribute information of the multiple sample frames to a server, so that the server determines a target highlight template and a highlight sample frame corresponding to each highlight lens among multiple highlight lenses included in the target highlight template, the highlight sample frame being the sample frame that matches the highlight lens among the multiple sample frames; receiving the target highlight template and the highlight sample frames corresponding to the multiple highlight lenses sent by the server; determining regular video segments and highlight video segments corresponding to the multiple editing materials based on the attribute information of the multiple sample frames, the target highlight template, and the highlight sample frames corresponding to the multiple highlight lenses; and generating a final video clip obtained after video editing of the multiple editing materials based on the regular video segments and the highlight video segments.

[0007] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the most exciting sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame. This allows for the generation of corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application can automatically generate final video segments without requiring users to possess editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video segment is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip generated from the final video segment based on the target highlight template and highlight sampling frames, the final video segment also possesses rich aesthetic appeal and expressiveness.

[0008] Optionally, the attribute information of the sampled frame includes a semantic feature vector, which indicates the image semantic information of the sampled frame. Optionally, the attribute information of the sampled frame includes feature labels, which include at least one type of feature label. Different types of feature labels can characterize the features of the sampled frame in different dimensions. Optionally, the feature labels can be divided into various types such as aesthetic labels, face labels, framing labels, object labels, camera movement labels, time labels, and location labels. Among them, aesthetic labels can characterize the features of the sampled frame in aesthetic dimensions such as composition and color, thereby reflecting the aesthetic value of the sampled frame; face labels can characterize the attributes of the face in the sampled frame; framing labels can characterize the features of the sampled frame in the framing dimension; object labels can characterize the objects contained in the sampled frame; camera movement labels can characterize the camera movement of the sampled frame; time labels can indicate the shooting time of the sampled frame; and location labels can indicate the shooting location of the sampled frame. Optionally, the feature labels of the sampled frame also include at least one label output by a multi-label model. In this case, at least one label output by the multi-label model can characterize the features of the sampled frame in at least one dimension. This at least one dimension includes the object dimension, the text dimension, and so on.

[0009] By determining the attribute information of the above sampling frames, we can fully understand the image features of each sampling frame. This makes the determination of highlight video segments and regular video segments more accurate, thus ensuring that the final video still has rich aesthetics and expressiveness.

[0010] Optionally, determining the regular video segment and highlight video segment corresponding to the multiple editing materials based on the attribute information of the multiple sampling frames, the target highlight template, and the highlight sampling frames corresponding to the multiple highlight shots includes: determining at least one set of regular sampling frames from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight shots based on the attribute information of the multiple sampling frames and the number of the multiple highlight shots; determining the regular shot corresponding to each regular sampling frame in the at least one set of regular sampling frames from the multiple editing materials; combining the regular shots corresponding to the regular sampling frames in the same group in the at least one set of regular sampling frames into a regular video segment to obtain at least one of the regular video segments; and determining at least one of the highlight video segments from the multiple editing materials based on the target highlight template and the highlight sampling frames corresponding to each of the highlight shots.

[0011] Optionally, determining at least one set of regular sampled frames from sampled frames other than those corresponding to the multiple highlight shots, based on the attribute information of the multiple sampled frames and the number of multiple highlight shots, includes: determining multiple candidate sampled frames from sampled frames other than those corresponding to the multiple highlight shots, based on the attribute information of the multiple sampled frames, wherein the image content of the multiple candidate sampled frames differs; clustering the multiple candidate sampled frames based on the attribute information of the multiple candidate sampled frames to obtain at least one cluster set, each cluster set including at least one of the candidate sampled frames; determining the number of regular shots included in the final video based on the number of the at least one cluster set and the number of multiple highlight shots; and selecting a set of regular sampled frames from each cluster set based on the number of regular shots and the number of candidate sampled frames in each cluster set to obtain the at least one set of regular sampled frames.

[0012] Because the image content of these multiple candidate sampling frames is different, it can effectively avoid the existence of similar images in the determined regular sampling frames, thereby ensuring that there are no duplicate video images in the final video, making the final video have rich aesthetics and expressiveness.

[0013] Optionally, after sending the attribute information of the plurality of sampled frames to the server, the method further includes: receiving target background music sent by the server, the target background music including at least one music segment, each music segment corresponding to a number of shots; determining the number of regular shots included in the final video based on the number of the at least one cluster set and the number of the plurality of highlight shots, including: determining a target music segment from the at least one music segment based on the number of the at least one cluster set, the number of the plurality of highlight shots, and the number of shots corresponding to each music segment; determining the number of shots corresponding to the target music segment as the number of shots in the final video; and determining the difference between the number of shots in the final video and the number of the plurality of highlight shots as the number of regular shots included in the final video.

[0014] Since the number of shots in the final video is determined based on the target background music, which is a segment of the target background music, this ensures that the final video matches the target background music better and makes the video more impactful with the support of the target music segment.

[0015] Optionally, after determining at least one set of regular sampling frames from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses based on the attribute information of the multiple sampling frames and the number of the multiple highlight lenses, the method further includes: sending the at least one set of regular sampling frames to the server, so that the server determines the video effect information corresponding to the final video based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight lens; receiving the video effect information sent by the server; and generating the final video obtained after video editing of the multiple editing materials based on the regular video segments and the highlight video segments includes: generating the final video based on the regular video segments, the highlight video segments, and the video effect information.

[0016] Since the video effect information is determined based on the regular sampling frames and the highlight sampling frames corresponding to each highlight shot, it is possible to determine the corresponding video effects for each shot in the final video clip in a more targeted manner. This ensures that the transitions between shots in the final video clip are more natural and that the final video clip is more impactful with the addition of video effects.

[0017] Optionally, generating the final video clip obtained after video editing of the multiple edited materials based on the regular video segments and the highlight video segments includes: displaying a video preview interface, the video preview interface including the regular video segments and the highlight video segments; responding to an editing operation on a target video segment, the target video segment being any one of the regular video segments and the highlight video segments; and generating the final video clip based on the edited regular video segments and highlight video segments.

[0018] In other words, after determining the regular video segments and the highlight video segments, users can replace unsatisfactory highlight segments and regular segments. That is, this application can provide users with more choices, so that the final video output can meet the user's preferences and requirements.

[0019] Secondly, a video editing method is provided, applied to a server. The method includes: receiving attribute information of each of a plurality of sampled frames sent by a terminal, wherein the plurality of sampled frames are obtained by sampling a plurality of editing materials, and the attribute information indicates the characteristics of the sampled frames; determining, based on the attribute information of the plurality of sampled frames, a target highlight template and a highlight sampled frame corresponding to each highlight shot among a plurality of highlight shots included in the target highlight template from a plurality of candidate highlight templates, wherein the highlight sampled frame is the sampled frame that matches the highlight shot among the plurality of sampled frames; sending the target highlight template and the highlight sampled frame corresponding to each highlight shot to the terminal, so that the terminal determines the regular video segment and the highlight video segment corresponding to the plurality of editing materials, and generates a final video clip obtained after video editing of the plurality of editing materials based on the regular video segment and the highlight video segment.

[0020] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the most exciting sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame. This allows for the generation of corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application can automatically generate final video segments without requiring users to possess editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video segment is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip generated from the final video segment based on the target highlight template and highlight sampling frames, the final video segment also possesses rich aesthetic appeal and expressiveness.

[0021] Optionally, each candidate highlight template has matching constraints; determining the target highlight template and the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template from the multiple candidate highlight templates based on the attribute information of the multiple sampling frames includes: determining the matching score between each sampling frame and each highlight lens in each candidate highlight template based on the attribute information of each sampling frame; determining multiple first highlight templates and the highlight sampling frame corresponding to each highlight lens in each first highlight template from the multiple candidate highlight templates based on the matching score between each sampling frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template; selecting at least one first highlight template as the target highlight template from the multiple first highlight templates, and using the highlight sampling frame corresponding to each highlight lens in the at least one first highlight template as the highlight sampling frame corresponding to each highlight lens in the target highlight template.

[0022] The highlight sampling frame corresponding to each highlight shot in the first highlight template is determined by the matching score between each sampling frame and each highlight shot in each candidate highlight template. This ensures that the final determined highlight sampling frame matches the highlight shot more accurately, making the final video have rich aesthetics and expressiveness.

[0023] Optionally, the attribute information of each sampled frame includes the feature label and semantic feature vector of the sampled frame, and each highlight shot in each candidate highlight template has label requirements and semantic feature vectors; determining the matching score between each sampled frame and each highlight shot in each candidate highlight template based on the attribute information of each sampled frame includes: for each candidate highlight template, determining the label score between each sampled frame and each highlight shot in the candidate highlight template based on the feature label of each sampled frame and the label requirements of each highlight shot in the candidate highlight template; determining the similarity score between each sampled frame and each highlight shot in the candidate highlight template based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight shot in the candidate highlight template; and determining the matching score between each sampled frame and each highlight shot in the candidate highlight template based on the label score and similarity score between each sampled frame and each highlight shot in the candidate highlight template.

[0024] The matching score between each sampled frame and each highlight shot in the candidate highlight template is determined by using a label score and a similarity score. The label score is related to the label requirements of each highlight shot, and the similarity score is related to the semantic feature vector. In this way, the final matching score can accurately guarantee the matching between the sampled frame and the highlight shot, thereby ensuring the accuracy of the highlight sampled frames determined based on the matching score.

[0025] Optionally, determining multiple first highlight templates and highlight sampling frames corresponding to each highlight lens in each first highlight template from the multiple candidate highlight templates based on the matching score between each sampling frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, includes: for each candidate highlight template, determining whether each highlight lens in the candidate highlight template has a matching sampling frame based on the matching score between each sampling frame and each highlight lens in the candidate highlight template and the matching constraints of the candidate highlight template; if each highlight lens in the candidate highlight template has a matching sampling frame, and the number of sampling frames matched by the target highlight lens in the candidate highlight template is multiple, then determining the highlight sampling frame corresponding to the target highlight lens based on the attribute information of the multiple sampling frames matched by the target highlight lens, and using the candidate highlight template as the first highlight template, and the target highlight lens as at least one highlight lens in the candidate highlight template.

[0026] Optionally, each candidate highlight template includes a category of beginning highlight template, middle highlight template, or ending highlight template; the step of selecting at least one first highlight template as the target highlight template from the plurality of first highlight templates, and using the highlight sampling frames corresponding to each highlight shot in the at least one first highlight template as the highlight sampling frames corresponding to each highlight shot in the target highlight template, includes: filtering the plurality of first highlight templates to obtain a plurality of second highlight templates, wherein the highlight sampling frames corresponding to highlight shots in any two second highlight templates do not overlap; if there are... If there are at least two beginning highlight templates and / or at least two ending highlight templates, then the beginning highlight template and the ending highlight template with the highest priority among the plurality of second highlight templates are selected; the beginning highlight template and the ending highlight template with the highest priority, as well as the middle highlight template among the plurality of second highlight templates, are determined as the target highlight template; and the highlight sampling frames corresponding to each highlight shot in the beginning highlight template and the ending highlight template with the highest priority, as well as the middle highlight template among the plurality of second highlight templates, are determined as the highlight sampling frames corresponding to each highlight shot in the target highlight template.

[0027] Since the highlight sampling frames corresponding to the highlight shots in each of the two second highlight templates do not overlap, this ensures that there are no similar images in the highlight segments, thereby ensuring that there are no repetitions in the video footage, making the finished video footage rich in aesthetics and expressiveness.

[0028] Optionally, the method further includes: determining a target background music from multiple candidate background music based on the attribute information of the multiple sampled frames, wherein the target background music includes at least one music segment, and each music segment corresponds to a number of shots; and sending the target background music to the terminal.

[0029] Since the final video also includes target background music, which is determined based on the attribute information of multiple sample frames, this ensures that the final target background music is more compatible with multiple materials, and makes the final video more impactful with the support of the target music segment.

[0030] Optionally, the method further includes: receiving at least one set of conventional sampling frames sent by the terminal;

[0031] Based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot in the target highlight template, determine the video effect information corresponding to the finished video; and send the video effect information to the terminal.

[0032] By determining the video effect information in the final video by using at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot in the target highlight template, the corresponding video effects can be determined more specifically for each shot in the final video. This ensures that the transitions between shots in the final video are more natural and that the final video is more impactful with the addition of video effects.

[0033] Thirdly, a video editing method is provided, the method comprising: sampling multiple editing materials to obtain multiple sample frames, the editing materials including videos or images; determining attribute information for each sample frame, the attribute information indicating the characteristics of the sample frame; based on the attribute information of the multiple sample frames, determining a target highlight template and a highlight sample frame corresponding to each highlight shot among multiple highlight shots included in the target highlight template from multiple candidate highlight templates, the highlight sample frame being the sample frame that matches the highlight shot among the multiple sample frames; determining regular video segments and highlight video segments corresponding to the multiple editing materials based on the attribute information of the multiple sample frames, the target highlight template, and the highlight sample frames corresponding to the multiple highlight shots; and generating a final video clip obtained after video editing of the multiple editing materials based on the regular video segments and the highlight video segments.

[0034] Optionally, each of the candidate specular templates has matching constraints;

[0035] The step of determining the target highlight template and the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template, based on the attribute information of the multiple sampling frames, includes:

[0036] Based on the attribute information of each sampled frame, a matching score is determined between each sampled frame and each highlight lens in each candidate highlight template;

[0037] Based on the matching score between each sampled frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, a plurality of first highlight templates and highlight sampled frames corresponding to each highlight lens in each first highlight template are determined from the plurality of candidate highlight templates.

[0038] At least one first highlight template is selected from the plurality of first highlight templates as the target highlight template, and the highlight sampling frame corresponding to each highlight lens in the at least one first highlight template is used as the highlight sampling frame corresponding to each highlight lens in the target highlight template.

[0039] Optionally, the attribute information of each sampled frame includes the feature label and semantic feature vector of the sampled frame, and each highlight shot in each candidate highlight template has label requirements and semantic feature vector;

[0040] The step of determining the matching score between each sampled frame and each highlight shot in each candidate highlight template based on the attribute information of each sampled frame includes:

[0041] For each candidate highlight template, based on the feature label of each sampled frame and the label requirement of each highlight shot in the candidate highlight template, the label score between each sampled frame and each highlight shot in the candidate highlight template is determined;

[0042] Based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight shot in the candidate highlight template, a similarity score is determined between each sampled frame and each highlight shot in the candidate highlight template.

[0043] Based on the label score and similarity score between each sampled frame and each highlight shot in the candidate highlight template, a matching score is determined between each sampled frame and each highlight shot in the candidate highlight template.

[0044] Optionally, determining a plurality of first highlight templates and the highlight sampling frames corresponding to each highlight lens in each of the candidate highlight templates based on the matching score between each sampled frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, includes:

[0045] For each candidate highlight template, based on the matching score between each sampled frame and each highlight shot in the candidate highlight template and the matching constraints of the candidate highlight template, it is determined whether each highlight shot in the candidate highlight template has a matching sampled frame.

[0046] If each highlight lens in the candidate highlight template has a matching sampling frame, and the number of sampling frames matching the target highlight lens in the candidate highlight template is multiple, then based on the attribute information of the multiple sampling frames matching the target highlight lens, the highlight sampling frame corresponding to the target highlight lens is determined, and the candidate highlight template is used as the first highlight template, and the target highlight lens is at least one highlight lens in the candidate highlight template.

[0047] Optionally, each of the candidate highlight templates may be categorized into a beginning highlight template, a middle highlight template, or a ending highlight template.

[0048] The step of selecting at least one first highlight template from the plurality of first highlight templates as the target highlight template, and using the highlight sampling frame corresponding to each highlight shot in the at least one first highlight template as the highlight sampling frame corresponding to each highlight shot in the target highlight template, includes:

[0049] The plurality of first highlight templates are filtered to obtain a plurality of second highlight templates, wherein the highlight sampling frames corresponding to the highlight lenses in each pair of second highlight templates do not overlap;

[0050] If there are at least two beginning highlight templates and / or at least two ending highlight templates among the plurality of second highlight templates, then the beginning highlight template and the ending highlight template with the highest priority among the plurality of second highlight templates are selected.

[0051] The highest priority opening highlight template and ending highlight template, as well as the middle highlight template among the plurality of second highlight templates, are determined as the target highlight template. The highlight sampling frame corresponding to each highlight shot in the highest priority opening highlight template and ending highlight template, as well as the middle highlight template among the plurality of second highlight templates, is determined as the highlight sampling frame corresponding to each highlight shot in the target highlight template.

[0052] Optionally, determining the regular video segment and highlight video segment corresponding to the multiple editing materials based on the attribute information of the multiple sampling frames, the target highlight template, and the highlight sampling frames corresponding to the multiple highlight shots includes:

[0053] Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses;

[0054] Determine the regular shot corresponding to each regular sampling frame in the at least one set of regular sampling frames from the plurality of clip materials;

[0055] By combining the regular shots corresponding to the regular sampling frames in the same group of the at least one group of regular sampling frames into a regular video segment, at least one of the regular video segments is obtained.

[0056] Based on the target highlight template and the highlight sampling frame corresponding to each highlight shot, at least one highlight video segment is determined from the plurality of edited materials.

[0057] Optionally, based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, determining at least one set of regular sampling frames from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses includes:

[0058] Based on the attribute information of the multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses, and the image content of the multiple candidate sampling frames is different;

[0059] Based on the attribute information of the multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one of the candidate sampling frames;

[0060] Based on the number of the at least one cluster set and the number of the plurality of highlight shots, the number of regular shots included in the final video is determined;

[0061] Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain the at least one set of regular sample frames.

[0062] Optionally, after determining the attribute information of each of the sampled frames, the method further includes:

[0063] Based on the attribute information of the multiple sampled frames, a target background music is determined from multiple candidate background music. The target background music includes at least one music segment, and each music segment corresponds to one shot.

[0064] Determining the number of regular shots included in the final video clip based on the number of the at least one cluster set and the number of the plurality of highlight shots includes:

[0065] Based on the number of the at least one cluster set, the number of the plurality of highlight shots, and the number of shots corresponding to each of the music segments, a target music segment is determined from the at least one music segment;

[0066] The number of shots corresponding to the target music segment is determined as the number of shots in the final video clip;

[0067] The difference between the number of shots in the final video and the number of the multiple highlight shots is determined as the number of regular shots included in the final video.

[0068] Optionally, after determining at least one set of regular sampling frames from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses based on the attribute information of the multiple sampling frames and the number of the multiple highlight lenses, the method further includes:

[0069] Based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot in the target highlight template, determine the video effect information corresponding to the finished video;

[0070] The process of generating a final video clip from the multiple edited materials, based on the regular video segments and the highlight video segments, includes:

[0071] The final video clip is generated based on the regular video clip, the highlight video clip, and the video effect information.

[0072] Optionally, after determining the attribute information of each of the multiple sampling frames, the method further includes:

[0073] Based on the attribute information of the multiple sampled frames, a target background music is determined from multiple candidate background music. The target background music includes at least one music segment, and each music segment corresponds to one shot.

[0074] Fourthly, a video editing apparatus is provided, which has the function of implementing the video editing method described in the first aspect. The video editing apparatus includes at least one module for implementing the video editing method provided in the first aspect.

[0075] Fifthly, a video editing apparatus is provided, which has the function of implementing the video editing method described in the second aspect above. The video editing apparatus includes at least one module for implementing the video editing method provided in the second aspect above.

[0076] Sixthly, a video editing apparatus is provided, which has the function of implementing the video editing method described in the third aspect above. The video editing apparatus includes at least one module for implementing the video editing method provided in the third aspect above.

[0077] A seventh aspect provides a terminal comprising a processor and a memory, the memory being used to store a computer program for executing the video editing method provided in the first or third aspect. The processor is configured to execute the computer program stored in the memory to implement the video editing method described in the first or third aspect.

[0078] Optionally, the terminal may further include a communication bus for establishing a connection between the processor and the memory.

[0079] Eighthly, a server is provided, the server including a processor and a memory, the memory being used to store a computer program for executing the video editing method provided in the second or third aspect above. The processor is configured to execute the computer program stored in the memory to implement the video editing method described in the second or third aspect above.

[0080] Optionally, the server may further include a communication bus for establishing a connection between the processor and the memory.

[0081] Ninthly, a computer-readable storage medium is provided, wherein the storage medium stores instructions that, when executed on a terminal, cause the terminal to perform the steps of the video editing method described in the first or third aspect.

[0082] In a tenth aspect, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a server, cause the server to perform the steps of the video editing method described in the second or third aspect above.

[0083] Eleventhly, a computer program product containing instructions is provided, which, when executed on a terminal, causes the terminal to perform the steps of the video editing method described in the first or third aspect. Alternatively, a computer program is provided that, when executed on a terminal, causes the terminal to perform the steps of the video editing method described in the first or third aspect.

[0084] In a twelfth aspect, a computer program product comprising instructions is provided, which, when executed on a server, cause the server to perform the steps of the video editing method described in the second or third aspect. Alternatively, a computer program is provided that, when executed on a server, causes the server to perform the steps of the video editing method described in the second or third aspect.

[0085] The technical effects achieved by the second to twelfth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here. Attached Figure Description

[0086] Figure 1 This is a schematic diagram of a video frame provided in an embodiment of this application;

[0087] Figure 2 This is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0088] Figure 3 This is a flowchart of a video editing method provided in an embodiment of this application;

[0089] Figure 4 This is a flowchart provided in an embodiment of the present application for determining a target specular template and the specular sampling frame corresponding to each specular lens among a plurality of specular lenses included in the target specular template;

[0090] Figure 5 This is a flowchart provided in an embodiment of the present application for determining the regular video segments and highlight video segments corresponding to the plurality of edited materials;

[0091] Figure 6 This is a flowchart of another video editing method provided in the embodiments of this application;

[0092] Figure 7 This is a schematic diagram of the structure of a video editing device provided in an embodiment of this application;

[0093] Figure 8 This is a schematic diagram of another video editing device provided in an embodiment of this application;

[0094] Figure 9 This is a schematic diagram of another video editing device provided in an embodiment of this application;

[0095] Figure 10 This is a structural block diagram of a terminal provided in an embodiment of this application;

[0096] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0097] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0098] To facilitate understanding, before providing a detailed explanation of the video editing method provided in the embodiments of this application, the terms, application scenarios, and implementation environments involved in the embodiments of this application will be introduced first.

[0099] Final video: refers to a complete video generated after cutting, merging, and recombining video, image, and other materials.

[0100] A shot is a continuous video segment, specifically the video clip between two cut points. It should be noted that images playing for a certain duration can also form a shot.

[0101] A segment, also known as a video sequence, is a video clip that can be composed of multiple segments. Each segment consists of several shots that are connected in the overall video clip and are conceptually closely related. For a given segment, the shots within that segment have clear differences and boundaries from the shots preceding and following it. Depending on different editing requirements and purposes, segments can also be classified into various types, such as regular segments, highlight segments, etc.

[0102] It should be noted that a single shot can constitute a segment, and multiple shots can also be combined into a single shot using split-screen or picture-in-picture techniques; this combined shot can also be called a segment. Please refer to [reference needed]. Figure 1 , Figure 1 It includes four shots, which can appear in the same video frame in a split-screen format and form a segment.

[0103] Highlight segments, also known as highlight video segments, are a typical segment type. Highlight segments are the most important and exciting parts of a final video, capable of attracting the viewer's attention, leaving a lasting impression, and even evoking emotional resonance. Highlight segments are typically generated by connecting at least one shot according to specific editing logic and techniques.

[0104] Based on their placement within the final video clip, highlight segments can be further divided into opening highlight segments, closing highlight segments, and mid-roll highlight segments. Opening highlight segments typically appear at the beginning of the video, closing highlight segments at the end, and mid-roll highlight segments in the middle. For example, using a series of rapidly changing shots at the beginning of a video to present an overview of the entire travel experience constitutes an opening highlight segment. Similarly, using a series of shots with a concluding message (such as a silhouette or closing credits) at the end of a video constitutes a closing highlight segment.

[0105] Regular segments: Also known as regular video segments, these are the most basic and important segment type in a finished video. Regular segments use camera cuts and transitions to present the development of the storyline and the actions of the characters. Regular segments typically follow a chronological order, allowing viewers to clearly understand the story's development and the characters' behavioral logic. In this embodiment, segments in the finished video, excluding highlight segments, are referred to as regular segments.

[0106] Shot type: This refers to the difference in the size of the subject within the frame caused by varying distances between the camera and the subject. Shot types are generally divided into five categories: close-up, medium shot, long shot, and extreme long shot.

[0107] With the rapid development of multimedia and the Internet, users have increasingly higher requirements for the quality of the videos they shoot. They no longer just want to record their lives through videos, but want to showcase videos that are more aesthetically pleasing and expressive.

[0108] Currently, users can manually cut, merge, and recombine video footage through video editing to generate finished videos with varying levels of expressiveness. Typically, experienced users can create videos with smooth transitions, clear storylines, rich layers, and artistic flair. In other words, the quality of a finished video highly depends on the user's editing skills. However, video editing is a specialized field, and learning and mastering video editing software usually takes considerable time. Users often cannot acquire video editing knowledge in a short period, making it difficult to edit high-quality segments from video footage and ultimately preventing them from creating more aesthetically pleasing and expressive videos.

[0109] Based on this, embodiments of this application provide a video editing method that can automatically generate a finished video based on the attribute information of multiple sampled frames. Since multiple sampled frames are obtained by sampling multiple editing materials, and the attribute information of multiple sampled frames can indicate the characteristics of the corresponding sampled frames, the method can determine the most compelling sampled frames from the editing materials that can form highlight segments, i.e., highlight sampled frames, based on the attribute information. It can also determine the highlight template matching the highlight sampled frame, thereby generating corresponding highlight video segments and regular video segments, ultimately generating the finished video. In other words, the video editing method provided by embodiments of this application can automatically generate a finished video without requiring the user to possess editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the finished video is generated based on regular video segments and highlight video segments, and the highlight segment is the most compelling video clip in the finished video generated based on the target highlight template and highlight sampled frames, the finished video also possesses rich aesthetic appeal and expressiveness.

[0110] The implementation environment involved in the embodiments of this application will be described next.

[0111] Please refer to Figure 2 , Figure 2 This is a schematic diagram of an implementation environment provided in an embodiment of this application. The implementation environment includes a terminal 10 and a server 20. The terminal 10 can communicate with the server 20. This communication connection can be a wired connection or a wireless connection, and this embodiment of the application does not limit it.

[0112] Terminal 10 can sample multiple clips to obtain multiple sample frames, and then determine the attribute information of each sample frame. Terminal 10 can send the attribute information of the multiple sample frames to server 20. Server 20 can receive the attribute information of each sample frame from the multiple sample frames sent by terminal 10, and based on the attribute information of the multiple sample frames, determine the target highlight template and the highlight sample frames corresponding to each highlight shot among the multiple highlight shots included in the target highlight template from multiple candidate highlight templates. Then, server 20 can send the target highlight template and the highlight sample frames corresponding to each highlight shot to terminal 10. Terminal 10 receives the target highlight template and the highlight sample frames corresponding to the multiple highlight shots sent by server 20, and based on the attribute information of the multiple clips, the multiple sample frames, the target highlight template, and the highlight sample frames corresponding to the multiple highlight shots, determines the regular video segments and highlight video segments corresponding to the multiple clips, and then generates the final video clip obtained after video editing of the multiple clips based on the regular video segments and highlight video segments.

[0113] In other words, terminal 10 is used to determine the attribute information of multiple sample frames, as well as to determine the regular video segments and highlight video segments corresponding to multiple editing materials and generate a finished video. Server 20 is used to determine the target highlight template and the highlight sample frames corresponding to each highlight shot among the multiple highlight shots included in the target highlight template. Terminal 10 and server 20 interact to jointly implement the above-described video editing method. In practical applications, either terminal 10 or server 20 can also execute the steps of determining the attribute information of multiple sample frames, determining the target highlight template and the highlight sample frames corresponding to each highlight shot among the multiple highlight shots included in the target highlight template, as well as determining the regular video segments and highlight video segments corresponding to multiple editing materials and generating a finished video. That is, either terminal 10 or server 20 can implement the above-described video editing method.

[0114] The following describes the process of implementing the video editing method described above on terminal 10, using the terminal as an example. Terminal 10 can sample multiple editing materials to obtain multiple sample frames, and determine the attribute information of each sample frame. Then, based on the attribute information of the multiple sample frames, it determines the target highlight template and the highlight sample frame corresponding to each highlight shot among the multiple highlight shots included in the target highlight template from multiple candidate highlight templates. Based on the attribute information of the multiple editing materials, the multiple sample frames, the target highlight template, and the highlight sample frames corresponding to the multiple highlight shots, terminal 10 determines the regular video segment and highlight video segment corresponding to the multiple editing materials. Then, based on the regular video segment and highlight video segment, it generates the final video clip obtained after video editing of the multiple editing materials.

[0115] The terminal 10 can be any electronic product that can interact with the user through one or more means such as a keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, such as PC (Personal Computer), mobile phone, smartphone, PDA (Personal Digital Assistant), wearable device, PPC (Pocket PC), tablet computer, smart car system, smart TV, smart speaker, etc.

[0116] Server 20 can be a standalone server, a server cluster or distributed system composed of multiple physical servers, a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, or a cloud computing service center.

[0117] Those skilled in the art should understand that the above-described terminal 10 and server 20 are merely examples. Other existing or future terminals or servers that are applicable to the embodiments of this application should also be included within the scope of protection of the embodiments of this application, and are hereby incorporated by reference.

[0118] It should be noted that the application scenarios and implementation environments described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, as the implementation environment evolves, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0119] The video editing method provided in the embodiments of this application will now be explained in detail. Figure 3 This is a flowchart illustrating a video editing method provided in an embodiment of this application. Please refer to it. Figure 3 The method includes the following steps.

[0120] Step 301: The terminal samples multiple clips to obtain multiple sample frames, where any one of the multiple clips can be a video or an image.

[0121] In some embodiments, for any one of the multiple clips, the terminal can sample that clip to obtain at least one sampled frame corresponding to that clip. Each of the multiple clips is sampled in the same manner to obtain at least one sampled frame corresponding to each of the multiple clips, ultimately resulting in multiple sampled frames.

[0122] Since any one of these multiple clips includes either video or image, meaning each clip can be either a video or an image, the terminal's sampling method differs depending on the type of clip. These methods will be described separately below.

[0123] When the clip material is an image, the terminal can directly use that clip material as the corresponding sampling frame.

[0124] When the edited material is video, the terminal can sample the edited material according to the reference sampling frequency to obtain at least one sampled frame corresponding to the edited material.

[0125] The reference sampling frequency is preset and can be adjusted according to different needs under different circumstances. For example, the reference sampling frequency can be 1 hertz (Hz). In this case, the clip can be sampled at a frequency of 1 Hz, that is, one video frame of the clip is captured every second as a sample frame.

[0126] In some embodiments, the terminal may also acquire the clip material before sampling it.

[0127] Optionally, the terminal has a camera that can capture videos or images, which can then be used as editing material. Alternatively, the terminal can also receive videos or images captured by other devices, which are devices with shooting capabilities other than the terminal mentioned above, and use those videos or images as editing material. In other words, the other device can send its own captured videos or images to the terminal. In practical applications, the terminal can also download videos or images from the internet and use those downloaded videos or images as editing material.

[0128] In other words, the edited material can be captured by the terminal itself, by other devices with shooting capabilities, or by videos or images downloaded by the terminal from the Internet. Of course, the terminal can also obtain the edited material through other means, and this application embodiment does not limit this.

[0129] Step 302: The terminal determines the attribute information of each of the multiple sampled frames, which indicates the characteristics of the sampled frame.

[0130] For any one of the multiple sampled frames, the terminal can determine the attribute information of that sampled frame. By processing each of the multiple sampled frames in the same way, the attribute information of each sampled frame can be determined.

[0131] In some embodiments, the attribute information of the sampled frame includes a semantic feature vector, which indicates the image semantic information of the sampled frame. In this case, the terminal can input the sampled frame into a multimodal semantic understanding model to obtain the semantic feature vector of the sampled frame output by the multimodal semantic understanding model.

[0132] As an example, the multimodal semantic understanding model can be a contrastive language and image pre-training (CLIP) model, a Transformer-based vision-and-language BERT (ViLBERT) model, etc., and this application does not limit it to this.

[0133] Optionally, the multimodal semantic understanding model can output a semantic feature vector with the data type of the target data type and the number of dimensions of the target dimension.

[0134] The target data type and the number of target dimensions are preset and can be adjusted according to different needs under different circumstances. For example, the target data type can be floating-point (float), and the number of target dimensions can be 512. In this case, the multimodal semantic understanding model can output a 512-dimensional floating-point semantic feature vector.

[0135] In some embodiments, the attribute information of the sampled frame includes feature labels for the sampled frame. These feature labels include at least one type of feature label, and different types of feature labels can characterize the features of the sampled frame in different dimensions. For example, the feature labels can be categorized into various types such as aesthetic labels, face labels, framing labels, object labels, camera movement labels, time labels, and location labels. Specifically, aesthetic labels can characterize the features of the sampled frame in aesthetic dimensions such as composition and color, thereby reflecting the aesthetic value of the sampled frame; face labels can characterize the attributes of the face in the sampled frame; framing labels can characterize the features of the sampled frame in the framing dimension; object labels can characterize the objects contained in the sampled frame; camera movement labels can characterize the camera movement of the sampled frame; time labels can indicate the shooting time of the sampled frame; and location labels can indicate the shooting location of the sampled frame.

[0136] Since there are many types of feature labels, the process of determining each type of feature label listed above will be described below.

[0137] When the feature label of a sampled frame includes an aesthetic label, the terminal can determine the aesthetic score of the sampled frame and then use the aesthetic score as the aesthetic label of the sampled frame.

[0138] In some embodiments, the terminal can input the sampled frame into an aesthetic model to obtain an aesthetic score for the sampled frame output by the aesthetic model.

[0139] When the feature label of a sampling frame includes a face label, the terminal can determine the attributes of the face in the sampling frame and then use the attributes of the face as the face label of the sampling frame.

[0140] In some embodiments, the terminal can input the sampled frame into a face model to obtain the attributes of the face in the sampled frame output by the face model. Optionally, the face attributes include the position of the face in the sampled frame, the angle of the face, and the expression of the face, etc. For example, the angle of the face includes frontal view, side view, etc., and the expression of the face includes crying face, smiling face, funny face, etc.

[0141] It should be noted that if a sampled frame does not contain a face, the face label of that sampled frame can indicate that the sampled frame does not contain a face.

[0142] When the feature label of a sampled frame includes a scene type label, the terminal can determine the scene type of the sampled frame and then use the scene type of the sampled frame as the scene type label of the sampled frame.

[0143] In some embodiments, the terminal can input the sampled frame into a scene model to obtain the scene type of the sampled frame output by the scene model, and use the scene type of the sampled frame as the scene type label of the sampled frame.

[0144] When the feature label of a sampling frame includes an object label, the terminal can determine the name of the object in the sampling frame and then use the name of the object as the object label of the sampling frame.

[0145] In some embodiments, the terminal can input the sampled frame into an object detection model to obtain the name of the object in the sampled frame output by the object detection model, and use the name of the object as the object label of the sampled frame.

[0146] Optionally, the terminal can also determine the location of the object corresponding to its name within the sampling frame, and then store the object's name and its location within the sampling frame accordingly. In this case, the object detection model can also output the location of the object in the sampling frame.

[0147] It should be noted that the position of an object in a sampling frame can be represented by its coordinates or its orientation within the frame. For example, the position of an object in a sampling frame can be represented as the object being located within the region enclosed by the four coordinates (x1, y1), (x2, y2), (x3, y3), and (x4, y4). Alternatively, the position of an object in a sampling frame can be represented as the object being located at the center of the sampling frame; however, this embodiment does not impose such limitations.

[0148] When the feature tags of a sampled frame include camera movement tags, the terminal can determine the camera movement direction and camera movement speed of the sampled frame based on the sampled frame and adjacent sampled frames, and use the camera movement direction and camera movement speed of the sampled frame as the camera movement tags of the sampled frame. The adjacent sampled frame corresponds to the same clip material as the sampled frame, and is either the previous or next sampled frame adjacent to the sampled frame.

[0149] In some embodiments, the terminal can input the sampled frame and adjacent sampled frames into the camera movement model to obtain the camera movement direction and camera movement speed of the sampled frame output by the camera movement model.

[0150] As an example, the camera movement direction of the sampled frame may include upward camera movement, downward camera movement, leftward camera movement, rightward camera movement, forward camera movement, and backward camera movement, etc., and this application embodiment does not limit this.

[0151] When the feature label of a sampled frame includes a time label, the terminal can use the capture time of the sampled frame as the corresponding time label.

[0152] When the feature label of a sampled frame includes a location label, the terminal can use the location where the sampled frame was captured as the location label corresponding to that sampled frame.

[0153] Each of the aforementioned feature labels can only characterize the features of the sampled frame in one dimension. However, in practical applications, the feature labels of the sampled frame also include at least one label output by a multi-label model. In this case, the at least one label output by the multi-label model can characterize the features of the sampled frame in at least one dimension. This at least one dimension includes the object dimension, the text dimension, and so on. For example, if the image content of sampled frame 1 includes a person walking a dog standing in front of a shop called "Time Cafe," then the labels that sampled frame 1 can output through the multi-label model include "person," "dog," and "cafe." Among these, the labels "person" and "dog" are the features of the sampled frame in the object dimension, and "cafe" is the feature of the sampled frame in the text dimension. For ease of description, the at least one label output by the multi-label model will be referred to as the first label.

[0154] In some embodiments, the terminal can directly input the sampled frame into a multi-label model to obtain at least one label of the sampled frame output by the multi-label model, and use the at least one label of the sampled frame as the first label of the sampled frame.

[0155] In practical applications, the feature labels of the sampling frames may also include other types of labels, which will not be listed one by one in this application embodiment.

[0156] Step 303: The terminal sends the attribute information of multiple sampled frames to the server.

[0157] Step 304: The server receives the attribute information of each of the multiple sampling frames sent by the terminal, and based on the attribute information of the multiple sampling frames, determines the target highlight template and the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template from the multiple candidate highlight templates. The highlight sampling frame is the sampling frame that matches the highlight lens among the multiple sampling frames.

[0158] In some embodiments, each candidate specular template has matching constraints, in which case please refer to Figure 4 The server can follow Figure 4 Steps 3041-3043 shown determine the target specular template and the specular sampling frame corresponding to each specular lens among the multiple specular lenses included in the target specular template.

[0159] 3041: Based on the attribute information of each sampled frame, determine the matching score between each sampled frame and each highlight shot in each candidate highlight template.

[0160] In some embodiments, the attribute information of each sampled frame includes the feature label and semantic feature vector of the sampled frame, and each highlight shot in each candidate highlight template has label requirements and semantic feature vector. In this case, the matching score between each sampled frame and each highlight shot in each candidate highlight template can be determined according to the following steps (1) to (3).

[0161] (1) For each candidate specular template, based on the feature label of each sampled frame and the label requirement of each specular shot in the candidate specular template, determine the label score between each sampled frame and each specular shot in the candidate specular template.

[0162] For any highlight shot in any candidate highlight template and any sample frame in any sample frame, the label requirement for the highlight shot indicates the feature label that the sample frame matching the highlight shot must satisfy. This label requirement includes at least one sub-label requirement. In this case, the number of sub-label requirements that the feature label of the sample frame can satisfy is taken as the label score between the sample frame and the highlight shot. Processing each sample frame and each highlight shot in the candidate highlight template in the same way yields the label score between each sample frame and each highlight shot in the candidate highlight template.

[0163] As an example, if the labeling requirements for highlight shot A in candidate highlight template 1 include sub-labeling requirement 1 and sub-labeling requirement 2, where sub-labeling requirement 1 is that the object label of the sampled frame includes "person" and "door", and sub-labeling requirement 2 is that the scene type label of the sampled frame includes "medium shot" or "full shot". Sample frame a has object labels including "person" and "door" and a scene type label including "full shot", while sample frame b has object labels including "person" and a scene type label including "medium shot". In this case, the feature labels of sample frame a satisfy both sub-labeling requirements of highlight shot A, namely sub-labeling requirement 1 and sub-labeling requirement 2, then the labeling score between sample frame a and highlight shot A is 2. The feature labels of sample frame b satisfy one sub-labeling requirement of highlight shot A, namely sub-labeling requirement 2, then the labeling score between sample frame b and highlight shot A is 1.

[0164] (2) Based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight shot in the candidate highlight template, determine the similarity score between each sampled frame and each highlight shot in the candidate highlight template.

[0165] In some embodiments, for any highlight shot in any candidate highlight template and any sample frame in any sample frame, the semantic feature vector of the highlight shot indicates the image semantic information corresponding to the sample frame that best matches the highlight shot. In this case, the similarity score between the semantic feature vector of the sample frame and the semantic feature vector of the highlight shot can be determined based on a relevant similarity measurement algorithm, or by inputting the semantic feature vectors of the sample frame and the highlight shot into a similarity score calculation model. Processing each sample frame and each highlight shot in the candidate highlight template in the same way yields a similarity score between each sample frame and each highlight shot in the candidate highlight template.

[0166] (3) Based on the label score and similarity score between each sampled frame and each highlight shot in the candidate highlight template, determine the matching score between each sampled frame and each highlight shot in the candidate highlight template.

[0167] In some embodiments, for any highlight shot in any candidate highlight template and any sample frame in any sample frame, the sum of the label score and similarity score between the sample frame and the highlight shot can be directly used as the matching score between the sample frame and the highlight shot. Processing each sample frame and each highlight shot in the candidate highlight template in the same way yields the matching score between each sample frame and each highlight shot in the candidate highlight template.

[0168] In other embodiments, the server stores weights corresponding to label scores and similarity scores, respectively. In this case, for any highlight shot in any candidate highlight template and any sample frame in any sample frame, the server multiplies the label score between the sample frame and the highlight shot by the weight corresponding to the label score as the target label score, multiplies the similarity score between the sample frame and the highlight shot by the weight corresponding to the similarity score as the target similarity score, and uses the sum of the target label score and the target similarity score as the matching score between the sample frame and the highlight shot. By processing each sample frame and each highlight shot in the candidate highlight template in the same way, a matching score can be obtained between each sample frame and each highlight shot in the candidate highlight template.

[0169] In summary, for any highlight shot within any candidate highlight template and any sample frame within any sample frame, the server can determine the label score and similarity score between the highlight shot and the sample frame. Based on these scores, a matching score can be determined between the highlight shot and the sample frame. In practical applications, the server can also determine only one of the label score and similarity score, directly using either score as the matching score between the highlight shot and the sample frame. If only the label score is determined, the attribute information of each sample frame includes its feature label, and each highlight shot in each candidate highlight template has label requirements. If only the similarity score is determined, the attribute information of each sample frame includes a semantic feature vector, and each highlight shot in each candidate highlight template has a semantic feature vector.

[0170] 3042: Based on the matching score between each sampled frame and each highlight shot in each candidate highlight template, and the matching constraints of each candidate highlight template, determine multiple first highlight templates and the highlight sampled frames corresponding to each highlight shot in each first highlight template from multiple candidate highlight templates.

[0171] Next, the target specular template and the specular sampling frame corresponding to each specular lens among the multiple specular lenses included in the target specular template can be determined by following the steps (1)-(2).

[0172] (1) For each candidate highlight template, based on the matching score between each sampled frame and each highlight shot in the candidate highlight template and the matching constraint conditions of the candidate highlight template, determine whether each highlight shot in the candidate highlight template has a matching sampled frame.

[0173] In some embodiments, based on the matching score between each sampled frame and each highlight shot in the candidate highlight template, as well as the matching constraints of the candidate highlight template, a brute-force solution or a bipartite graph matching algorithm can be used to determine whether each highlight shot in the candidate highlight template has a matching sampled frame, i.e., whether the candidate highlight template has a solution. Of course, in practical applications, other algorithms can also be used to determine whether the candidate highlight template has a solution, and this application embodiment does not limit this.

[0174] It should be noted that if at least one highlight lens in the candidate highlight template does not have a matching sampling frame, then the candidate highlight template is determined to be unsolvable.

[0175] In some embodiments, the matching constraints of candidate specular templates include conventional constraints and dynamic constraints. Conventional constraints refer to the constraints that each candidate specular template must follow in determining whether a solution exists; the conventional constraints are the same for each candidate specular template. Dynamic constraints are determined based on the matching score threshold corresponding to each candidate specular template.

[0176] As an example, common constraints include: for any one of the multiple candidate highlight templates, the sampled frames matched by each highlight shot in that candidate highlight template do not appear repeatedly.

[0177] In some embodiments, the server further stores a matching score threshold corresponding to each candidate highlight template. This matching score threshold includes a total score threshold and a score threshold corresponding to each shot in the respective candidate highlight template. In this case, for any one of the multiple candidate highlight templates, the dynamic constraints of the candidate highlight template include: the matching score between any highlight shot in the candidate highlight template and the matched sampling frame is not less than the score threshold corresponding to that shot, and the sum of the matching scores between each highlight shot in the candidate highlight template and the matched sampling frame is not less than the total score threshold corresponding to the candidate highlight template.

[0178] In other words, under the constraints of the matching conditions of the above candidate specular templates, the sampled frames matched by each specular lens in the candidate specular templates with a solution satisfy the above dynamic constraints and conventional constraints. That is, the image semantic information corresponding to the sampled frames matched by each specular lens is similar to that of the most matched sampled frame of the corresponding specular lens, and can satisfy the label requirements of the corresponding specular lens, and there is no duplication of the sampled frames matched by each specular lens.

[0179] For example, candidate highlight template a includes two highlight lenses, highlight lens a and highlight lens b. If candidate highlight template a has a solution, and the sampled frame matched by highlight lens a is sampled frame A, and the sampled frame matched by highlight lens b is sampled frame B, then the matching score between sampled frame A and highlight lens a is not less than the score threshold corresponding to highlight lens a, the matching score between sampled frame B and highlight lens b is not less than the score threshold corresponding to highlight lens b, and the sum of the matching scores between sampled frame A and highlight lens a and between sampled frame B and highlight lens b is not less than the total score threshold corresponding to candidate highlight template a.

[0180] In some embodiments, the category of each candidate highlight template includes a beginning highlight template, a middle highlight template, or an end highlight template, and the matching constraints of the candidate highlight templates also include time constraints, which are determined based on the category of the corresponding candidate highlight template.

[0181] Optionally, the server can determine the shooting time range of the multiple clips, and based on the shooting time range, determine the time constraints corresponding to the candidate highlight templates for each category.

[0182] In some embodiments, the attribute information of the plurality of sampled frames includes time tags, and the server can directly use the time range corresponding to the largest and smallest time tags among the plurality of time tags as the shooting time range of the plurality of clips.

[0183] Because the shooting time ranges of multiple clips correspond to different numbers of days, the implementation methods for determining the time constraints for each category of candidate highlight templates based on these shooting time ranges differ. These will be described separately below. It should be noted that the number of days corresponding to the shooting time range refers to the span of shooting dates within that shooting time range. If there are two shooting dates within the shooting time range, then the number of days corresponding to that shooting time range is determined to be 2 days.

[0184] If the shooting time range of multiple clips corresponds to one day, meaning the clips were shot within a single day, then because the shooting time is relatively concentrated, the time constraints for each category of candidate highlight templates can be disregarded; that is, the time constraints for each category of candidate highlight templates are empty. However, in practical applications, the daytime period (sunrise to sunset) can be used as the time constraints for the opening and middle highlight templates, while the period after sunset can be used as the time constraints for the ending highlight template.

[0185] If the shooting time range of multiple clips corresponds to 2 days, that is, the multiple clips were shot within 2 days, then the time range corresponding to the first date in the shooting time range can be used as the time constraint condition corresponding to the opening highlight template and the middle highlight template, and the time range corresponding to the second date in the shooting time range can be used as the time constraint condition corresponding to the ending highlight template, where the first date is earlier than the second date.

[0186] If the shooting time range of multiple clips corresponds to 3 days, that is, the multiple clips were shot within 3 days, then the time range corresponding to the first date in the shooting time range can be used as the time constraint condition for the opening highlight template, the time range corresponding to the second date in the shooting time range can be used as the time constraint condition for the middle highlight template, and the time range corresponding to the third date in the shooting time range can be used as the time constraint condition for the ending highlight template. The first date is earlier than the second and third dates, and the second date is earlier than the third date.

[0187] If the shooting time range of multiple clips corresponds to more than 3 days, that is, the multiple clips were shot within at least four days, then the time range corresponding to x1% of the shooting time range can be used as the time constraint condition for the opening highlight template, the time range corresponding to x2% of the shooting time range can be used as the time constraint condition for the middle highlight template, and the time range corresponding to x3% of the shooting time range can be used as the time constraint condition for the ending highlight template. The sum of x1, x2, and x3 is not greater than 100, the time ranges corresponding to x1%, x2, and x3 do not overlap, and the time range corresponding to x1% is earlier than the time ranges corresponding to x2% and x3%, and the time range corresponding to x2% is earlier than the time range corresponding to x3%.

[0188] Among them, x1, x2 and x3 are preset and can be adjusted according to different needs under different circumstances.

[0189] (2) If each highlight lens in the candidate highlight template has a matching sampling frame, and the number of sampling frames matched by the target highlight lens in the candidate highlight template is multiple, then based on the attribute information of the multiple sampling frames matched by the target highlight lens, the highlight sampling frame corresponding to the target highlight lens is determined, and the candidate highlight template is used as the first highlight template, and the target highlight lens is at least one highlight lens in the candidate highlight template.

[0190] If each highlight shot in the candidate highlight template has a matching sample frame, and the target highlight shot in the candidate highlight template matches multiple sample frames (i.e., the candidate highlight template has multiple solutions, and at least one highlight shot in the candidate highlight template can match multiple sample frames), then it is necessary to filter the multiple solutions corresponding to the candidate highlight template.

[0191] In some embodiments, the attribute information of the multiple sample frames matched by the target highlight lens includes aesthetic tags. In this case, for one highlight lens in the target highlight lens, the sample frame with the largest aesthetic tag among the multiple sample frames matched by the highlight lens is taken as the highlight sample frame corresponding to that highlight lens. Each highlight lens in the target highlight lens is processed in the same way to obtain the highlight sample frame corresponding to each highlight lens in the target highlight lens.

[0192] In other embodiments, for one of the highlight lenses in the target highlight lens, a matching score between each sampling frame and the highlight lens can be determined based on the attribute information of multiple sampling frames matched by the highlight lens. The sampling frame with the highest matching score is then taken as the highlight sampling frame corresponding to the highlight lens. This process is repeated for each highlight lens in the target highlight lens to obtain the highlight sampling frame corresponding to each highlight lens. Of course, in practical applications, other methods can be used to determine the highlight sampling frame corresponding to the highlight lens. For example, any one of the multiple sampling frames matched by the highlight lens can be taken as the highlight sampling frame corresponding to the highlight lens. This application does not limit this approach.

[0193] The process of determining the matching score between each sample frame and the highlight lens based on the attribute information of multiple sample frames matched by the highlight lens is similar to the process of determining the matching score between each sample frame and each highlight lens in each candidate highlight template based on the attribute information of each sample frame. For details, please refer to the relevant content above, which will not be repeated here.

[0194] 3043: Select at least one first highlight template from multiple first highlight templates as the target highlight template, and take the highlight sampling frame corresponding to each highlight shot in the at least one first highlight template as the highlight sampling frame corresponding to each highlight shot in the target highlight template.

[0195] In some embodiments, each candidate highlight template is categorized into an initial highlight template, a middle highlight template, or an initial highlight template. In this case, multiple first highlight templates can be filtered to obtain multiple second highlight templates. The highlight sampling frames corresponding to highlight shots in any two second highlight templates do not overlap. If at least two initial highlight templates and / or at least two initial highlight templates exist among the multiple second highlight templates, then the initial highlight template and the initial highlight template with the highest priority and the initial highlight template with the highest priority and the initial highlight template with the highest priority and the initial highlight template with the highest priority and the middle highlight template among the multiple second highlight templates are determined as the target highlight template. The highlight sampling frames corresponding to each highlight shot in the initial highlight template with the highest priority and the initial highlight template with the highest priority and the middle highlight template among the multiple second highlight templates are determined as the highlight sampling frames corresponding to each highlight shot in the target highlight template.

[0196] Although the specular sampling frames corresponding to different specular shots within the same first specular template do not overlap, there may be overlapping specular sampling frames between different first specular templates. To avoid overlapping sampling frames in the final generated specular video segment, it is necessary to filter the multiple first specular templates to obtain multiple second specular templates.

[0197] In some embodiments, the server stores the priorities corresponding to multiple candidate highlight templates, that is, the server stores the priorities corresponding to multiple first highlight templates. In this case, the specific implementation process of filtering multiple first highlight templates to obtain multiple second highlight templates includes: counting the occurrence count of highlight sampling frames corresponding to each highlight lens in the multiple first highlight templates; if there are no highlight sampling frames with an occurrence count greater than 2 in the statistical results, the multiple first highlight templates are used as multiple second highlight templates. If there are highlight sampling frames with an occurrence count greater than 2 in the statistical results, the first highlight template with the lowest priority among the first highlight templates corresponding to the highlight sampling frames with an occurrence count greater than 2 is deleted from the multiple first highlight templates, and the above step of counting the occurrence count of highlight sampling frames corresponding to each highlight lens in the multiple first highlight templates is returned. At this time, the multiple second highlight templates are the multiple first highlight templates after deletion.

[0198] If the statistical results do not contain any highlight sampling frames that appear more than twice, it means that there are no duplicate highlight sampling frames among the multiple first highlight templates. Therefore, these multiple first highlight templates can be used as multiple second highlight templates. If the statistical results contain any highlight sampling frames that appear more than twice, it means that there are duplicate highlight sampling frames among the multiple first highlight templates. Therefore, the first highlight template with the lowest priority among the first highlight templates corresponding to the highlight sampling frames that appear more than twice can be deleted from the multiple first highlight templates, and the steps described above for counting the occurrences of highlight sampling frames corresponding to each highlight lens in the multiple first highlight templates are returned.

[0199] Since a video clip typically contains only one beginning highlight template and / or one ending highlight template, and the aforementioned multiple second highlight templates may contain multiple beginning highlight templates and / or at least two ending highlight templates, in order to avoid duplicate highlight segments in the final generated highlight video segments, it is necessary to filter the multiple second highlight templates to obtain the target highlight template.

[0200] In some embodiments, the server stores the priorities corresponding to multiple candidate highlight templates. In this case, if at least two beginning highlight templates and / or at least two ending highlight templates exist among the multiple second highlight templates, then the beginning highlight template and the ending highlight template with the highest priority among the multiple second highlight templates are selected. The beginning highlight template and the ending highlight template with the highest priority, as well as the middle highlight template among the multiple second highlight templates, are determined as the target highlight template. The highlight sampling frames corresponding to each highlight shot in the beginning highlight template and the ending highlight template, as well as the middle highlight template among the multiple second highlight templates, are determined as the highlight sampling frames corresponding to each highlight shot in the target highlight template. If at least two beginning highlight templates and / or at least two ending highlight templates do not exist among the multiple second highlight templates, then the multiple second highlight templates are directly determined as the target highlight template, and the highlight sampling frames corresponding to each highlight shot in the multiple second highlight templates are determined as the highlight sampling frames corresponding to each highlight shot in the target highlight template.

[0201] In summary, when at least two beginning highlight templates and / or at least two ending highlight templates exist among multiple second highlight templates, the beginning highlight template with higher priority and the at least two ending highlight templates can be used as the target highlight template. In practical applications, when at least two beginning highlight templates and / or at least two ending highlight templates exist among multiple second highlight templates, any one of the at least two beginning highlight templates and any one of the at least two ending highlight templates can be selected, and the arbitrarily selected beginning highlight template and ending highlight template can be used as the target highlight template. This application does not limit this approach.

[0202] Step 305: The server sends the target highlight template and the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template to the terminal.

[0203] Step 306: The terminal receives the target highlight template and the highlight sampling frames corresponding to the multiple highlight shots included in the target highlight template sent by the server, and determines the regular video segment and highlight video segment corresponding to the multiple clips based on the attribute information of the multiple clips and the multiple sampling frames, the target highlight template and the highlight sampling frames corresponding to the multiple highlight shots included in the target highlight template.

[0204] In some embodiments, please refer to Figure 5 The terminal can Figure 5 Steps 3061-3064 shown determine the regular video segments and highlight video segments corresponding to the multiple clip materials.

[0205] 3061: Based on the attribute information of the multiple sampling frames and the number of the multiple highlight lenses, determine at least one set of regular sampling frames from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses.

[0206] In some embodiments, the terminal may determine at least one set of regular sampling frames according to the following steps (1) to (4).

[0207] (1) Based on the attribute information of the multiple sampled frames, multiple candidate sampled frames are determined from the sampled frames other than the highlight sampled frames corresponding to the multiple highlight lenses. The image content of the multiple candidate sampled frames is different.

[0208] In some embodiments, the attribute information of each sampled frame includes a semantic feature vector. Based on the semantic feature vector of the sampled frames other than the highlight sampled frames corresponding to the multiple highlight lenses, a similarity score between every two sampled frames in the sampled frames other than the highlight sampled frames corresponding to the multiple highlight lenses is determined to obtain multiple first scores. Based on the multiple first scores, multiple candidate sampled frames are determined from the sampled frames other than the highlight sampled frames corresponding to the multiple highlight lenses.

[0209] For ease of description, the sampled frames other than the highlight sampled frames corresponding to the multiple highlight shots will be referred to as the remaining sampled frames.

[0210] The specific implementation process for determining the similarity score between any two sampled frames in the remaining sampled frames includes: for any two sampled frames in the remaining sampled frames, determining the similarity score between the two sampled frames based on the semantic feature vectors of the two sampled frames. This process is repeated for every two sampled frames in the remaining sampled frames to obtain the similarity score between every two sampled frames in the remaining sampled frames, i.e., multiple first scores.

[0211] The specific implementation method for determining the similarity score between the two sampled frames based on the semantic feature vectors of the two sampled frames is similar to the implementation method for determining the similarity score between each sampled frame and each highlight lens in the candidate highlight template based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight lens in the candidate highlight template. For details, please refer to the relevant content above, which will not be repeated here.

[0212] The specific implementation process for determining multiple candidate sample frames from the remaining sample frames based on the multiple first scores includes: if there is at least one first score among the multiple first scores that is greater than a first similarity threshold, then the first target sample frame in the two sample frames corresponding to the first target score of the at least one first score in the remaining sample frames is deleted, and the step of determining the similarity score between every two sample frames in the remaining sample frames is returned. At this time, the remaining sample frames are the remaining sample frames after deletion, the first target score is any one of the at least one first scores, or the first score with the largest value among the at least one first scores, and the first target sample frame is any one of the two sample frames corresponding to the first target score, or the sample frame with the smallest aesthetic label among the two sample frames corresponding to the first target score. If there is no first score among the multiple first scores that is greater than the first similarity threshold, then the remaining sample frames are used as the multiple candidate sample frames.

[0213] If at least one of the multiple first scores is greater than the first similarity threshold, it indicates that there are sampled frames with similar image content in the remaining sampled frames. Therefore, the first target sampled frame in the two sampled frames corresponding to the first target score of the at least one first score in the remaining sampled frames can be deleted. If there is no first score greater than the first similarity threshold among the multiple first scores, it indicates that there are no sampled frames with similar image content in the remaining sampled frames. Therefore, the remaining sampled frames can be used as the multiple candidate sampled frames.

[0214] The first similarity threshold is set in advance and can be adjusted according to different needs under different circumstances.

[0215] In summary, the terminal can eliminate samples with small differences in image content from the remaining sampled frames based on the similarity scores of two sampled frames, and then use the eliminated remaining sampled frames as the multiple candidate sampled frames. In practical applications, the terminal can also eliminate remaining sampled frames based on the feature labels of the sampled frames.

[0216] In some embodiments, the feature label of the sampled frame includes a location label. In this case, if there are at least two sampled frames with the same location label among the remaining sampled frames, then all sampled frames other than the second sampled frame among the at least two sampled frames with the same location label are deleted to obtain multiple candidate sampled frames. The second sampled frame is any one of the at least two sampled frames with the same location label, or it is the sampled frame with the largest aesthetic label among the at least two sampled frames with the same location label. If there are no sampled frames with the same location label among the remaining sampled frames, then the remaining sampled frames are used as the multiple candidate sampled frames.

[0217] In some embodiments, the feature label of the sampled frame includes a time label. In this case, the difference in time labels between every two sampled frames in the remaining sampled frames is determined to obtain multiple time differences. If at least one time difference among the multiple time differences is less than a time difference threshold, the second target sampled frame of the two sampled frames corresponding to the target time difference of the at least one time difference in the remaining sampled frames is deleted, and the step of determining the difference in time labels between every two sampled frames in the remaining sampled frames is returned. At this time, the remaining sampled frames are the remaining sampled frames after deletion, the target time difference is any one of the at least one time difference, or the smallest time difference among the at least one time difference, and the second target sampled frame is any one of the two sampled frames corresponding to the target time difference, or the sampled frame with the smallest aesthetic label among the two sampled frames corresponding to the target time difference. If there is no time difference less than the time difference threshold among the multiple time differences, the remaining sampled frame is used as the multiple candidate sampled frames.

[0218] In some embodiments, the feature label of the sampled frame includes at least one label output by a multi-label model. In this case, the number of identical labels in the first labels of every two sampled frames in the remaining sampled frames is determined to obtain multiple label counts. If at least one label count among the multiple label counts is greater than a threshold, the target third sampled frame in the two sampled frames corresponding to the difference in the target label count among the at least one label count in the remaining sampled frames is deleted, and the step of determining the number of identical labels in the first labels of every two sampled frames in the remaining sampled frames is returned. At this time, the remaining sampled frame is the remaining sampled frame after deletion, the target label count is any one of the at least one label counts, or the largest label count among the at least one label counts, and the target third sampled frame is any one of the two sampled frames corresponding to the target label count, or the sampled frame with the smallest aesthetic label among the two sampled frames corresponding to the target label count. If no label count among the multiple label counts is greater than the threshold, the remaining sampled frame is used as the multiple candidate sampled frames.

[0219] Since some of the candidate sample frames may be located in shots with unstable camera movements, after determining these candidate sample frames, they can be further filtered to obtain filtered candidate sample frames in shots with stable camera movements.

[0220] The specific implementation process for further filtering the multiple candidate sample frames to obtain the filtered multiple candidate sample frames includes: determining the target shot duration; for any candidate sample frame among the multiple candidate sample frames, cutting a video clip with a duration equal to the target shot duration from the multiple clip materials, the video clip containing the candidate sample frame; taking the video clip containing the candidate sample frame as the first shot corresponding to the candidate sample frame; if the camera movement speed of the first shot is greater than the camera movement speed threshold and / or the camera movement variance is greater than the variance threshold, then the candidate sample frame is deleted from the multiple candidate sample frames; otherwise, the candidate sample frame is not deleted from the multiple candidate sample frames. The target shot duration refers to the duration of each shot in the final video. Each candidate sample frame among the multiple candidate sample frames is processed in the same way to obtain the filtered multiple candidate sample frames.

[0221] In some embodiments, the terminal stores the target shot duration, in which case the terminal can directly obtain the target shot duration. In other embodiments, the terminal can obtain target background music, which includes the shot duration corresponding to the target background music, and the terminal uses the shot duration corresponding to the target background music as the target shot duration.

[0222] Optionally, after the terminal sends the attribute information of the multiple sampled frames to the server, the server can also determine the target background music from multiple candidate background music based on the attribute information of the multiple sampled frames. The target background music includes the shot duration corresponding to the target background music, and then sends the target background music to the terminal.

[0223] The process of determining the target background music from multiple candidate background music based on the attribute information of multiple sampled frames includes: fusing the attribute information of the multiple sampled frames to obtain fused first information; determining the candidate background music corresponding to the fused first information from multiple candidate background music based on the fused first information; and using the candidate background music corresponding to the fused first information as the target background music. The fused first information can characterize the common features of the multiple sampled frames.

[0224] In some embodiments, the attribute information of the plurality of sampled frames includes a first label and / or a semantic feature vector, and the fused information includes a fused first label and / or a fused semantic feature vector. In this case, the union of the first labels of the plurality of sampled frames is used as the fused first label, and the average of the semantic feature vectors of the plurality of sampled frames is used as the fused semantic feature vector.

[0225] As an example, the candidate background music can be obtained by following a matching recommendation algorithm or by inputting the fused information into a matching recommendation model. Of course, in practical applications, other methods can also be used to determine the candidate background music corresponding to the fused information, and this application embodiment does not limit this approach.

[0226] In some embodiments, the average of the camera movement speeds of at least one sampled frame included in the first shot is taken as the camera movement speed of the first shot, and the variance of the camera movement values ​​corresponding to at least one sampled frame included in the first shot is taken as the camera movement variance of the shot. The camera movement value corresponding to each sampled frame can indicate the camera movement direction of the sampled frame.

[0227] Optionally, the camera movement direction of the sampled frame may include upward, downward, leftward, rightward, forward, and backward camera movement. In this case, the camera movement values ​​corresponding to the upward camera sampled frame and the downward camera sampled frame are opposites of each other, the camera movement values ​​corresponding to the leftward camera sampled frame and the rightward camera sampled frame are opposites of each other, and the camera movement values ​​corresponding to the forward camera sampled frame and the backward camera sampled frame are opposites of each other.

[0228] Among them, the time difference threshold, quantity threshold, camera movement speed threshold, and variance threshold are all preset, and can be adjusted according to different needs under different circumstances.

[0229] It should be noted that if the multiple candidate sampling frames are further filtered to obtain multiple filtered candidate sampling frames, then the multiple candidate sampling frames in subsequent steps are the multiple filtered candidate sampling frames.

[0230] (2) Based on the attribute information of the multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one candidate sampling frame.

[0231] In some embodiments, the multiple candidate sampled frames can be clustered based on their semantic feature vectors according to a relevant clustering algorithm to obtain at least one cluster set.

[0232] In other embodiments, the multiple candidate sampled frames can be clustered according to relevant clustering algorithms based on the semantic feature vectors and clustering constraints of the multiple candidate sampled frames to obtain at least one cluster set.

[0233] For example, the clustering constraints include: for each cluster set, the time labels and / or location labels of the candidate sampling frames in the cluster set are the same, and / or, the first labels of the candidate sampling frames in the cluster set have intersection, and / or, the number of identical labels among the first labels of the candidate sampling frames in the cluster set is greater than the cluster label threshold. Of course, in practical applications, the clustering constraints may also include other contents, and this application embodiment does not limit them.

[0234] The clustering label threshold is preset and can be adjusted according to different needs under different circumstances.

[0235] (3) Based on the number of the at least one cluster set and the number of the multiple highlight shots, determine the number of regular shots included in the final video.

[0236] There are several ways to determine the number of standard shots included in a final video clip. Two of these methods will be introduced below.

[0237] The first implementation method determines the number of shots in the final video based on the number of the at least one cluster set and the number of the multiple highlight shots, and determines the number of regular shots included in the final video by the difference between the number of shots in the final video and the number of the multiple highlight shots.

[0238] Optionally, the sum of the number of the at least one cluster set multiplied by the number of reference shots and the number of the plurality of highlight shots is used as the number of shots in the final video.

[0239] In the second implementation, the terminal can acquire target background music, which includes at least one music segment, with each music segment corresponding to a certain number of shots. In this case, based on the number of the at least one cluster set, the number of multiple highlight shots, and the number of shots corresponding to each music segment, the target music segment is determined from the at least one music segment. The number of shots corresponding to the target music segment is determined as the number of shots in the final video clip. The difference between the number of shots in the final video clip and the number of multiple highlight shots is determined as the number of regular shots included in the final video clip.

[0240] Optionally, after the terminal sends the attribute information of the multiple sampled frames to the server, the server can also determine the target background music from multiple candidate background music based on the attribute information of the multiple sampled frames. The target background music includes at least one music segment, and each music segment corresponds to a number of shots, and then sends the target background music to the terminal.

[0241] The process of determining the target background music from multiple candidate background music based on the attribute information of these multiple sampling frames has been described above. For details, please refer to the relevant content above, and it will not be repeated here.

[0242] The specific implementation process of determining the target music segment from the at least one music segment based on the number of the at least one cluster set, the number of multiple highlight shots, and the number of shots corresponding to each music segment includes: multiplying the number of the at least one cluster set by the number of reference shots and the sum of the number of multiple highlight shots as the first number of shots, and taking the music segment corresponding to the number of shots in the at least one that is greater than the first number of shots and has the smallest difference from the first number of shots as the target music segment.

[0243] (4) Based on the number of regular shots and the number of candidate sampling frames in each cluster set, select a set of regular sampling frames from each cluster set to obtain at least one set of regular sampling frames.

[0244] Based on the number of regular shots and the number of candidate sample frames in each cluster set, the number of regular sample frames in each cluster set is determined. Based on the number of regular sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain at least one set of regular sample frames.

[0245] For any cluster in at least one cluster set, the number of regular shots divided by the number of candidate sampled frames is multiplied by the number of candidate sampled frames in that cluster set to obtain a first value. Based on this first value, the number of regular sampled frames in that cluster set is determined. The same method is applied to each cluster in at least one cluster set to obtain the number of regular sampled frames in each cluster set.

[0246] In some embodiments, the first value is rounded down to the nearest integer and used as the number of regular sampled frames in the cluster set. In other embodiments, the multiple reference value with the smallest difference from the first value among a plurality of multiple reference values ​​is used as the number of regular sampled frames in the cluster set, where the multiple reference value is a positive integer multiple of the reference value.

[0247] There are several ways to select a set of regular sampling frames from each cluster set based on the number of regular sampling frames in each cluster set. Two of these methods will be introduced below.

[0248] The first implementation method is as follows: For any cluster in at least one cluster set, randomly select candidate sampling frames from that cluster set, with the number of candidate frames equal to the number of regular sampling frames in that cluster set, to obtain a set of regular sampling frames corresponding to that cluster set. Process each cluster set in the same way to obtain at least one set of regular sampling frames.

[0249] The second implementation method: For any cluster within at least one cluster set, the candidate sampled frames in that cluster set are sorted according to aesthetic labels from high to low, and the top H candidate sampled frames are taken as a set of regular sampled frames corresponding to that cluster set. The same method is applied to each cluster within at least one cluster set to obtain at least one set of regular sampled frames. Here, H is the number of regular sampled frames in that cluster set.

[0250] As an example, if the number of regular sampled frames in cluster set 1 is 4, then the candidate sampled frames in cluster set 1 are sorted in descending order of aesthetic labels, and the first 4 candidate sampled frames are taken as a set of regular sampled frames corresponding to cluster set 1.

[0251] 3062: Determine the regular shot corresponding to each regular sample frame in at least one set of regular sample frames from the multiple clips.

[0252] For any regular sample frame in any group of at least one set of regular sample frames, a video segment with a duration equal to the target shot duration is clipped from the multiple clips. This video segment contains the regular sample frame, and the video segment containing the regular sample frame is taken as the regular shot corresponding to that regular sample frame. Each regular sample frame in each group of at least one set of regular sample frames is processed in the same way, and finally, the regular shot corresponding to each regular sample frame in at least one set of regular sample frames can be obtained.

[0253] 3063: Combine the regular shots corresponding to the regular sampled frames in the same group of at least one group of regular sampled frames into a regular video segment to obtain at least one regular video segment.

[0254] 3064: Based on the target highlight template and the highlight sampling frames corresponding to each highlight shot, determine at least one highlight video segment from the multiple clips.

[0255] In some embodiments, for any one of the target highlight templates, the target highlight template includes a shot duration coefficient corresponding to each highlight shot in the target highlight template. Based on the shot duration ratio corresponding to each highlight shot and the highlight sampling frame corresponding to each highlight shot, the highlight video segment corresponding to the target highlight template is determined from the multiple clips. Each target highlight template is processed in the same way to ultimately obtain at least one highlight video segment.

[0256] For any highlight shot in the target highlight template, the duration of the highlight shot is obtained by multiplying the duration coefficient corresponding to that highlight shot by the duration of the target shot. A video segment containing the highlight sampling frame corresponding to that highlight shot is then edited from the multiple clips to obtain the corresponding video segment. This process is repeated for each highlight shot in the target highlight template to ultimately obtain the video segment corresponding to each highlight shot.

[0257] In some embodiments, the target highlight template includes the order of video segments corresponding to each highlight shot in the target highlight template. The video segments corresponding to each highlight shot are merged according to the order of the video segments corresponding to each highlight shot to obtain the highlight video segment corresponding to the target highlight template.

[0258] Step 307: The terminal generates the final video clip obtained after video editing based on the regular video segments and highlight video segments corresponding to multiple editing materials.

[0259] There are several ways to generate a final video clip from multiple edited clips, including regular and highlight video segments. Two of these methods will be described below.

[0260] The first implementation method is as follows: The terminal directly determines the priority of the regular video segments and highlight video segments corresponding to multiple editing materials, and merges the regular video segments and highlight video segments in order of priority from high to low to obtain the final video. The priority of the regular video segments and highlight video segments indicates their sequential position in the final video.

[0261] For any highlight video segment within the highlight video segments, if the target highlight template category corresponding to that highlight video segment is the beginning highlight segment, then that highlight video segment is determined to have the highest priority. If the target highlight template category corresponding to that highlight video segment is the end highlight segment, then that highlight video segment is determined to have the lowest priority. If the target highlight template category corresponding to that highlight video segment is the middle highlight segment, then the average shooting time of that highlight video segment is determined. Based on the average shooting time of that highlight video segment and the average shooting time of each regular video segment within the regular video segments, the priority of that highlight video segment and each regular video segment is determined.

[0262] The average time of shooting for the highlight video segment is taken as the average time of shooting for that segment. The method for determining the average time of shooting for each regular video segment is the same as that for determining the average time of shooting for the highlight video segment, and will not be repeated here.

[0263] The average shooting time of the highlight video segment and the average shooting time of each regular video segment are sorted from early to late. This sorting result is used as the priority result for the highlight video segment and each regular video segment, with video segments with earlier average shooting times having higher priority than those with later average shooting times.

[0264] In some embodiments, the terminal can generate the final video based on the regular video segments corresponding to the multiple clips, the highlight video segments corresponding to the multiple clips, and the target music segment.

[0265] In this scenario, the target music includes at least one musical segment, each corresponding to a melodic transition point. Based on the regular and highlight video segments corresponding to multiple clips, their priorities are determined. These segments are then merged in descending order of priority, and the target music segment is used as background music to obtain the final video clip. The priority of the regular and highlight video segments indicates their sequential position within the final video clip.

[0266] For any highlight video segment, if the target highlight template category corresponding to that highlight video segment is the beginning highlight segment, then that highlight video segment has the highest priority. If the target highlight template category corresponding to that highlight video segment is the end highlight segment, then that highlight video segment has the lowest priority. If the target highlight template category corresponding to that highlight video segment is the middle highlight segment, then the melody transition time point corresponding to the target music segment is determined as the average shooting time of that highlight video segment. Based on the average shooting time of that highlight video segment and the average shooting time of each regular video segment among the regular video segments, the priority of that highlight video segment and each regular video segment is determined.

[0267] In some embodiments, the terminal can also generate the final video clip based on the video effect information corresponding to the regular video segments, the highlight video segments, and the final video clip.

[0268] In this scenario, after the terminal determines at least one set of regular sampling frames, it can send the at least one set of regular sampling frames to the server. The server then receives the at least one set of regular sampling frames sent by the terminal, and based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight lens in the target highlight template, it determines the video effect information corresponding to the final video from multiple candidate effect information, and sends the video effect information to the terminal.

[0269] In some embodiments, candidate effect information includes candidate titles, candidate filters, candidate transitions, and candidate sound effects. The video effect information corresponding to the final video clip includes target titles, target filters, target transitions, and target sound effects. Target transitions include transitions for highlight shots and transitions for regular shots; target shots refer to shots corresponding to regular sampling frames. Target sound effects include sound effects corresponding to highlight segments and sound effects corresponding to regular segments.

[0270] Based on the attribute information of at least one set of regular sampling frames and the attribute information of the highlight sampling frames corresponding to each highlight lens in the target highlight template, the server determines the fused second information. Based on the fused second information, the server determines the candidate title and candidate filter corresponding to the fused second information from multiple candidate titles and candidate filters, and uses the candidate title and candidate filter corresponding to the fused second information as the target title and target filter.

[0271] The process of determining the fused second information and, based on the fused second information, determining the candidate title and candidate filter corresponding to the fused second information from multiple candidate titles and candidate filters is similar to the process of determining the fused first information and, based on the fused first information, determining the candidate background music corresponding to the fused first information from multiple candidate background music. For details, please refer to the relevant content above, which will not be repeated here.

[0272] In some embodiments, each candidate transition has a corresponding priority. In this case, for any one regular sampled frame in any set of at least one set of regular sampled frames, based on the attribute information of the regular sampled frame, a matching score is determined between the regular sampled frame and each of the multiple candidate transitions to obtain multiple transition matching scores. The candidate transition with the highest priority among the multiple transition matching scores that is greater than the transition matching score threshold is taken as the transition for the regular shot corresponding to the regular sampled frame. This process can be applied to each regular sampled frame in the same way to determine the transition for the regular shot corresponding to each regular sampled frame.

[0273] The transition matching score threshold is preset and can be adjusted according to different needs under different circumstances.

[0274] Based on the attribute information of the regular sampled frame, the matching score between the regular sampled frame and each candidate transition in multiple candidate transitions is determined. This process is similar to the above process of determining the matching score between each sampled frame and each highlight shot in each candidate highlight template based on the attribute information of each sampled frame. For details, please refer to the relevant content above, which will not be repeated here.

[0275] In some embodiments, each target highlight template includes a transition corresponding to each highlight shot in that target highlight template. In this case, the server can directly determine the transition corresponding to each highlight shot. Of course, the transition corresponding to the highlight shot can also be determined in the same way as the transition of a regular shot described above, and this application embodiment does not limit this.

[0276] In some embodiments, each tag in the first tag has a corresponding priority, and each tag corresponds to a sound effect. In this case, for any regular video segment within the regular video segment, the server can use the sound effect corresponding to the highest priority tag in the intersection of the first tags of the sampled frames in the regular video segment as the sound effect corresponding to that regular video segment. By processing each regular video segment in the same way, the sound effect corresponding to each regular video segment can be determined.

[0277] In some embodiments, each target highlight template includes the sound effect of that target highlight template. In this case, for any highlight video segment within the highlight video segments, the server can determine the sound effect of the target highlight template corresponding to that highlight video segment as the sound effect corresponding to that highlight video segment. Of course, the sound effect corresponding to the highlight video segment can also be determined in the same way as the sound effect corresponding to a regular video segment described above, and this application embodiment does not limit this.

[0278] The second implementation method is as follows: The terminal displays a video preview interface, which includes the regular video segment and the highlight video segment. In response to the editing operation of the target video segment, the editing operation is performed on the target video segment, which can be either the regular video segment or the highlight video segment. Based on the edited regular video segment and highlight video segment, the final video is generated.

[0279] Since users may not be satisfied with the automatically generated regular video segments and highlight video segments, they can trigger the editing operation of the target video segment in the video preview interface. In response to the editing operation of the target video segment, the terminal displays at least one replacement segment corresponding to the target video segment. The user selects a replacement segment from the at least one replacement segment corresponding to the target video segment and triggers the replacement confirmation operation to instruct the terminal to replace the target video segment with the replacement segment selected by the user.

[0280] If the target video segment is a regular video segment, then the cluster set corresponding to the target video segment is determined. From the candidate sample frames in the cluster set, excluding those corresponding to the target video segment, at least one set of sample frames with the same number of shots as the target video segment is selected. Based on this at least one set of sample frames, shots corresponding to each sample frame in the at least one set of sample frames are determined from multiple clips. Shots corresponding to sample frames in the same group within the at least one set of sample frames are combined into a replacement segment to obtain at least one replacement segment corresponding to the target video segment. If the target video segment is a highlight video segment, then the category of the target highlight template corresponding to the target video segment is determined. From multiple first highlight templates, at least one other first highlight template belonging to the same category as the target highlight template is determined. From multiple second highlight templates, at least one other second highlight template belonging to the same category as the target highlight template is determined. Based on this at least one first highlight template and / or at least one second highlight template, at least one replacement segment corresponding to the target video segment is determined from multiple clips.

[0281] The method of determining at least one replacement segment corresponding to the target video segment from multiple clips based on at least one first highlight template and / or at least one second highlight template is similar to the method described above of determining at least one highlight video segment from multiple clips based on the target highlight template and the highlight sampling frame corresponding to each highlight shot. For details, please refer to the relevant content above, which will not be repeated here.

[0282] The process of generating the final video clip based on the edited regular and highlight video segments is similar to the first method described above. For detailed implementation, please refer to the relevant content above, which will not be repeated here.

[0283] In practical applications, users can also modify the duration of the target shot, that is, modify it to a positive integer multiple of the target shot duration. In this case, the duration of each shot in the regular video segment and highlight video segment corresponding to the multiple clips is the modified target shot duration. Since the duration of each shot is usually longer than the original target shot duration after modification, the playback duration of the final video is a positive integer multiple of the target music segment. In this case, the audio in the target background music that is a positive integer multiple of the duration of the target music segment after the start time of the target music segment is used as the background music of the final video.

[0284] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the best sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame, thereby generating corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application embodiment can automatically generate a final video without requiring the user to have editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip in the final video generated based on the target highlight template and highlight sampling frames, the final video also possesses rich aesthetics and expressiveness. Moreover, this application embodiment can also replace highlight segments and regular segments, that is, provide users with more choices, so that the final generated video can meet the user's preferences and requirements.

[0285] The above embodiment achieves the video editing method through interaction between a terminal and a server. In practical applications, the video editing method can also be implemented by either the terminal or the server. Please refer to... Figure 6 , Figure 6 This is a flowchart of another video editing method provided in an embodiment of this application. The method includes the following steps.

[0286] Step 601: Sample multiple clips to obtain multiple sample frames. Any one of these clips can be a video or an image.

[0287] The implementation of this step can refer to the implementation method in step 301 of the aforementioned embodiment, where the terminal samples multiple clip materials to obtain multiple sample frames. However, unlike the aforementioned embodiment, the execution subject in this embodiment can also be a server. When the execution subject is a server, the multiple clip materials can be sent from the terminal to the server, or they can be pre-stored in the server, or they can be obtained by the server through other means; this application embodiment does not limit this.

[0288] Step 602: Determine the attribute information of each sample frame in the multiple sample frames, which indicates the characteristics of the sample frame.

[0289] The implementation of this step can refer to the relevant method described in step 302 of the foregoing embodiment, and will not be repeated here in this application embodiment.

[0290] Step 603: Based on the attribute information of each sample frame in multiple sample frames, determine the target highlight template and the highlight sample frame corresponding to each highlight lens in the multiple highlight lenses included in the target highlight template from multiple candidate highlight templates. The highlight sample frame is the sample frame that matches the highlight lens in the multiple sample frames.

[0291] The implementation of this step can refer to the relevant methods described in steps 3041-3043 of the foregoing embodiments, and will not be repeated here in the embodiments of this application.

[0292] Step 604: Based on the attribute information of each sample frame in multiple clips, the target highlight template, and the highlight sample frames corresponding to multiple highlight shots, determine the regular video segments and highlight video segments corresponding to the multiple clips.

[0293] The implementation of this step can refer to the relevant methods described in steps 3061-3064 of the foregoing embodiments, and will not be repeated here in the embodiments of this application.

[0294] Step 605: Based on the regular video segments and highlight video segments corresponding to multiple editing materials, generate the final video clip obtained after video editing of these multiple editing materials.

[0295] The implementation of this step can refer to step 307 in the aforementioned embodiment, where the terminal generates a final video clip based on regular video segments and highlight video segments corresponding to multiple editing materials after video editing. However, unlike the aforementioned embodiment, the execution entity in this embodiment can also be a server. When the execution entity is a server, in the second implementation of step 307, the server can send the regular video segment and the highlight video segment to the terminal so that the terminal can display a video preview interface. The user can trigger the editing operation of the target video segment in the video preview interface. In response to the editing operation of the target video segment, the terminal sends the target video segment to the server. The server receives the target video segment and determines at least one replacement segment corresponding to the target video segment. The server sends the at least one replacement segment corresponding to the target video segment to the terminal. The terminal displays the at least one replacement segment corresponding to the target video segment. The user selects a replacement segment from the at least one replacement segment corresponding to the target video segment and triggers a replacement confirmation operation to instruct the terminal to replace the target video segment with the replacement segment selected by the user. The terminal sends the replacement segment selected by the user to the server. The server receives the replacement segment selected by the user to obtain the edited regular video segment and the highlight video segment. Based on the edited regular video segment and the highlight video segment, the final video is generated.

[0296] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the best sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame, thereby generating corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application embodiment can automatically generate a final video without requiring the user to have editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip in the final video generated based on the target highlight template and highlight sampling frames, the final video also possesses rich aesthetics and expressiveness. Moreover, this application embodiment can also replace highlight segments and regular segments, that is, provide users with more choices, so that the final generated video can meet the user's preferences and requirements. Furthermore, in this embodiment, the video editing method described above can be implemented using only one device, without the need for interaction between the terminal and the server, thereby saving communication time and greatly improving the efficiency of video editing.

[0297] Figure 7 This is a schematic diagram of the structure of a video editing device provided in an embodiment of this application. This video editing device can be implemented as part or all of a terminal by software, hardware, or a combination of both. See also... Figure 7 The device includes: a sampling module 701, a first determining module 702, a first sending module 703, a first receiving module 704, a second determining module 705, and a generating module 706.

[0298] The sampling module 701 is used to sample multiple clip materials to obtain multiple sample frames. These clip materials include videos or images. For detailed implementation processes, please refer to the corresponding content in the above embodiments; they will not be repeated here.

[0299] The first determining module 702 is used to determine the attribute information of each sampling frame, which indicates the characteristics of the sampling frame. For detailed implementation details, please refer to the corresponding content in the above embodiments; they will not be repeated here.

[0300] The first sending module 703 is used to send attribute information of multiple sampling frames to the server, so that the server can determine the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template. The highlight sampling frame is the sampling frame that matches the highlight lens among the multiple sampling frames. For detailed implementation process, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0301] The first receiving module 704 is used to receive the target highlight template and highlight sampling frames corresponding to multiple highlight lenses sent by the server. For detailed implementation details, please refer to the corresponding contents in the above embodiments, which will not be repeated here.

[0302] The second determining module 705 is used to determine the regular video segments and highlight video segments corresponding to the multiple clip materials based on the attribute information of the multiple clipping materials, the multiple sampling frames, the target highlight template, and the highlight sampling frames corresponding to the multiple highlight shots. For detailed implementation processes, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0303] The generation module 706 is used to generate a final video clip based on regular video segments and highlight video segments corresponding to multiple editing materials after video editing. For detailed implementation processes, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0304] Optionally, the second determining module 705 is specifically used for:

[0305] Based on the attribute information of multiple sampling frames and the number of multiple highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses.

[0306] Identify the regular shot corresponding to each regular sample frame in at least one set of regular sample frames from multiple clips;

[0307] By combining regular shots corresponding to regular sample frames in the same group from at least one group of regular sample frames into a regular video segment, at least one regular video segment is obtained.

[0308] Based on the target highlight template and the highlight sampling frames corresponding to each highlight shot, at least one highlight video segment is determined from multiple clips.

[0309] Optionally, the second determining module 705 is specifically used for:

[0310] Based on the attribute information of multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to multiple highlight lenses. The image content of these multiple candidate sampling frames is different.

[0311] Based on the attribute information of multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one candidate sampling frame.

[0312] The number of regular shots included in the final video is determined based on the number of at least one cluster set and the number of multiple highlight shots;

[0313] Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain at least one set of regular sample frames.

[0314] Optionally, the video editing device also includes:

[0315] The second receiving module is used to receive target background music sent by the server. The target background music includes at least one music segment, and each music segment corresponds to one shot.

[0316] The second determining module 705 is specifically used for:

[0317] The target music segment is determined from at least one music segment based on the number of at least one cluster set, the number of multiple highlight shots, and the number of shots corresponding to each music segment.

[0318] The number of shots corresponding to the target music segment is determined as the number of shots in the final video.

[0319] The difference between the number of shots in the final video and the number of multiple highlight shots is determined as the number of regular shots included in the final video.

[0320] Optionally, the video editing device also includes:

[0321] The second sending module is used to send at least one set of regular sampling frames to the server so that the server can determine the video effect information corresponding to the final video based on at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot.

[0322] The third receiving module is used to receive video effect information sent by the server.

[0323] Module 706 is specifically used for:

[0324] The final video clip is generated based on regular video segments, highlight video segments, and video effect information.

[0325] Optionally, the generation module 706 is specifically used for:

[0326] Displays a video preview interface, which includes regular video segments and highlight video segments;

[0327] In response to an editing operation on the target video segment, perform the editing operation on the target video segment, which can be either a regular video segment or a highlight video segment.

[0328] Based on the edited regular video segments and highlight video segments, a final video clip is generated.

[0329] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the best sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame, thereby generating corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application embodiment can automatically generate a final video without requiring the user to have editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip in the final video generated based on the target highlight template and highlight sampling frames, the final video also possesses rich aesthetics and expressiveness. Moreover, this application embodiment can also replace highlight segments and regular segments, that is, provide users with more choices, so that the final generated video can meet the user's preferences and requirements.

[0330] It should be noted that the video editing device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video editing device and the video editing method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0331] Figure 8 This is a schematic diagram of the structure of a video editing device provided in an embodiment of this application. This video editing device can be implemented as part or all of a server by software, hardware, or a combination of both. See also... Figure 8 The device includes: a first receiving module 801, a first determining module 802, and a first sending module 803.

[0332] The first receiving module 801 is used to receive attribute information of each of the multiple sampling frames sent by the terminal. The multiple sampling frames are obtained by sampling multiple clip materials, and the attribute information indicates the characteristics of the sampling frames. For detailed implementation process, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0333] The first determining module 802 is used to determine, based on the attribute information of multiple sampling frames, a target highlight template and a highlight sampling frame corresponding to each highlight lens among multiple candidate highlight templates, wherein the highlight sampling frame is the sampling frame that matches the highlight lens among the multiple sampling frames. Detailed implementation processes are described in the corresponding contents of the above embodiments and will not be repeated here.

[0334] The first sending module 803 is used to send the target highlight template and the highlight sampling frames corresponding to each highlight shot to the terminal, so that the terminal can determine the regular video segments and highlight video segments corresponding to multiple editing materials, and generate a final video clip obtained after video editing of multiple editing materials based on the regular video segments and highlight video segments. For detailed implementation processes, please refer to the corresponding contents in the above embodiments, which will not be repeated here.

[0335] Optionally, each candidate specular template has matching constraints;

[0336] The first determining module 802 is specifically used for:

[0337] Based on the attribute information of each sampled frame, determine the matching score between each sampled frame and each highlight shot in each candidate highlight template;

[0338] Based on the matching score between each sampled frame and each highlight shot in each candidate highlight template, and the matching constraints of each candidate highlight template, multiple first highlight templates and the highlight sampled frames corresponding to each highlight shot in each first highlight template are determined from multiple candidate highlight templates.

[0339] Select at least one first highlight template from multiple first highlight templates as the target highlight template, and use the highlight sampling frame corresponding to each highlight shot in the at least one first highlight template as the highlight sampling frame corresponding to each highlight shot in the target highlight template.

[0340] Optionally, the attribute information of each sampled frame includes the feature label and semantic feature vector of the sampled frame, and each highlight shot in each candidate highlight template has label requirements and semantic feature vector;

[0341] The first determining module 802 is specifically used for:

[0342] For each candidate specular template, based on the feature label of each sampled frame and the label requirements of each specular shot in the candidate specular template, the label score between each sampled frame and each specular shot in the candidate specular template is determined;

[0343] Based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight shot in the candidate highlight template, the similarity score between each sampled frame and each highlight shot in the candidate highlight template is determined.

[0344] Based on the label score and similarity score between each sampled frame and each highlight shot in the candidate highlight template, a matching score is determined between each sampled frame and each highlight shot in the candidate highlight template.

[0345] Optionally, the first determining module 802 is specifically used for:

[0346] For each candidate highlight template, based on the matching score between each sampled frame and each highlight shot in the candidate highlight template, as well as the matching constraints of the candidate highlight template, it is determined whether each highlight shot in the candidate highlight template has a matching sampled frame.

[0347] If each highlight lens in the candidate highlight template has a matching sampling frame, and the number of sampling frames matching the target highlight lens in the candidate highlight template is multiple, then based on the attribute information of the multiple sampling frames matching the target highlight lens, the highlight sampling frame corresponding to the target highlight lens is determined, and the candidate highlight template is used as the first highlight template, and the target highlight lens is at least one highlight lens in the candidate highlight template.

[0348] Optionally, each candidate highlight template can be categorized into a beginning highlight template, a middle highlight template, or a ending highlight template.

[0349] The first determining module 802 is specifically used for:

[0350] Multiple first highlight templates are filtered to obtain multiple second highlight templates, and the highlight sampling frames corresponding to the highlight lenses in each pair of second highlight templates do not overlap;

[0351] If there are at least two beginning highlight templates and / or at least two ending highlight templates among the multiple second highlight templates, then select the beginning highlight template and the ending highlight template with the highest priority among the multiple second highlight templates;

[0352] The highest priority opening highlight template and closing highlight template, as well as the middle highlight template among multiple second highlight templates, are determined as the target highlight template. The highlight sampling frames corresponding to each highlight shot in the highest priority opening highlight template and closing highlight template, as well as the middle highlight template among multiple second highlight templates, are determined as the highlight sampling frames corresponding to each highlight shot in the target highlight template.

[0353] Optionally, the video editing device further includes:

[0354] The second determining module is used to determine the target background music from multiple candidate background music based on the attribute information of multiple sampled frames. The target background music includes at least one music segment, and each music segment corresponds to one shot.

[0355] The second sending module is used to send the target background music to the terminal.

[0356] Optionally, the video editing device further includes:

[0357] The second receiving module is used to receive at least one set of regular sampling frames sent by the terminal;

[0358] The third determining module is used to determine the video effect information corresponding to the final video based on at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot in the target highlight template.

[0359] The third sending module is used to send video effect information to the terminal.

[0360] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the best sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame, thereby generating corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application embodiment can automatically generate a final video without requiring the user to have editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip in the final video generated based on the target highlight template and highlight sampling frames, the final video also possesses rich aesthetics and expressiveness. Moreover, this application embodiment can also replace highlight segments and regular segments, that is, provide users with more choices, so that the final generated video can meet the user's preferences and requirements.

[0361] It should be noted that the video editing device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video editing device and the video editing method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0362] Figure 9 This is a schematic diagram of the structure of a video editing device provided in an embodiment of this application. The video editing device can be implemented as part or all of a terminal by software, hardware, or a combination of both, or it can be implemented as part or all of a server by software, hardware, or a combination of both. See also Figure 9 The device includes: a sampling module 901, a first determining module 902, a second determining module 903, a third determining module 904, and a generating module 905.

[0363] The sampling module 901 is used to sample multiple clip materials to obtain multiple sample frames. These clip materials include videos or images. For detailed implementation processes, please refer to the corresponding content in the above embodiments; they will not be repeated here.

[0364] The first determining module 902 is used to determine the attribute information of each sampling frame, which indicates the characteristics of the sampling frame. For detailed implementation details, please refer to the corresponding content in the above embodiments; these details will not be repeated here.

[0365] The second determining module 903 is used to determine, based on the attribute information of multiple sampling frames, a target highlight template and a highlight sampling frame corresponding to each highlight lens among multiple candidate highlight templates, wherein the highlight sampling frame is the sampling frame that matches the highlight lens among the multiple sampling frames. For detailed implementation processes, please refer to the corresponding contents in the above embodiments, which will not be repeated here.

[0366] The third determining module 904 is used to determine the regular video segments and highlight video segments corresponding to multiple clip materials based on the attribute information of multiple clipping materials, multiple sampling frames, the target highlight template, and the highlight sampling frames corresponding to multiple highlight shots. For detailed implementation processes, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0367] The generation module 905 is used to generate a final video clip based on regular video segments and highlight video segments corresponding to multiple editing materials after video editing. For detailed implementation processes, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0368] Optionally, each candidate specular template has matching constraints;

[0369] The second determining module 903 is specifically used for:

[0370] Based on the attribute information of each sampled frame, determine the matching score between each sampled frame and each highlight shot in each candidate highlight template;

[0371] Based on the matching score between each sampled frame and each highlight shot in each candidate highlight template, and the matching constraints of each candidate highlight template, multiple first highlight templates and the highlight sampled frames corresponding to each highlight shot in each first highlight template are determined from multiple candidate highlight templates.

[0372] Select at least one first highlight template from multiple first highlight templates as the target highlight template, and use the highlight sampling frame corresponding to each highlight shot in the at least one first highlight template as the highlight sampling frame corresponding to each highlight shot in the target highlight template.

[0373] Optionally, the third determining module 904 is specifically used for:

[0374] Based on the attribute information of multiple sampling frames and the number of multiple highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses.

[0375] Identify the regular shot corresponding to each regular sample frame in at least one set of regular sample frames from multiple clips;

[0376] By combining regular shots corresponding to regular sample frames in the same group from at least one group of regular sample frames into a regular video segment, at least one regular video segment is obtained.

[0377] Based on the target highlight template and the highlight sampling frames corresponding to each highlight shot, at least one highlight video segment is determined from multiple clips.

[0378] Optionally, the third determining module 904 is specifically used for:

[0379] Based on the attribute information of multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to multiple highlight lenses. The image content of the multiple candidate sampling frames is different.

[0380] Based on the attribute information of multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one candidate sampling frame.

[0381] The number of regular shots included in the final video is determined based on the number of at least one cluster set and the number of multiple highlight shots;

[0382] Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain at least one set of regular sample frames.

[0383] Optionally, the device further includes:

[0384] The fourth determining module is used to determine the target background music from multiple candidate background music based on the attribute information of multiple sampled frames. The target background music includes at least one music segment, and each music segment corresponds to one shot.

[0385] The third determining module 904 is specifically used for:

[0386] The target music segment is determined from at least one music segment based on the number of at least one cluster set, the number of multiple highlight shots, and the number of shots corresponding to each music segment.

[0387] The number of shots corresponding to the target music segment is determined as the number of shots in the final video.

[0388] The difference between the number of shots in the final video and the number of multiple highlight shots is determined as the number of regular shots included in the final video.

[0389] Since multiple sampling frames are obtained by sampling multiple clips, and the attribute information of multiple sampling frames can indicate the characteristics of the corresponding sampling frames, the best sampling frames that can form highlight segments in the clips can be determined from these multiple sampling frames based on the attribute information. These highlight sampling frames can also be determined, along with the highlight template matching the highlight sampling frame, thereby generating corresponding highlight video segments and regular video segments, ultimately producing the final video. In other words, the video editing method provided in this application embodiment can automatically generate a final video without requiring the user to have editing knowledge, thus effectively improving the efficiency of video editing. Furthermore, since the final video is generated based on regular video segments and highlight video segments, and the highlight segment is the most exciting video clip in the final video generated based on the target highlight template and highlight sampling frames, the final video also possesses rich aesthetics and expressiveness. Moreover, this application embodiment can also replace highlight segments and regular segments, that is, provide users with more choices, so that the final generated video can meet the user's preferences and requirements. Furthermore, in this embodiment, the video editing method described above can be implemented using only one device, without the need for interaction between the terminal and the server, thereby saving communication time and greatly improving the efficiency of video editing.

[0390] It should be noted that the video editing device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video editing device and the video editing method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0391] Figure 10 This is a structural block diagram of a terminal 1000 provided in an embodiment of this application. The terminal 1000 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1000 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0392] Typically, terminal 1000 includes a processor 1001 and a memory 1002.

[0393] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0394] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one instruction, which is executed by the processor 1001 to implement the control method of the smart device provided in the method embodiments of this application.

[0395] In some embodiments, the terminal 1000 may also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1004, a touch display screen 1005, a camera 1006, an audio circuit 1007, a positioning component 1008, and a power supply 1009.

[0396] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0397] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application embodiment.

[0398] Display screen 1005 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, serving as the front panel of terminal 1000; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of terminal 1000 or in a folded design; in still other embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of terminal 1000. Furthermore, display screen 1005 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0399] The camera assembly 1006 is used to acquire images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0400] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.

[0401] The positioning component 1008 is used to determine the current geographical location of the terminal 1000 in order to enable navigation or LBS (Location Based Service). The positioning component 1008 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.

[0402] Power supply 1009 is used to power the various components in terminal 1000. Power supply 1009 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1009 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0403] In some embodiments, the terminal 1000 further includes one or more sensors 1010. The one or more sensors 1010 include, but are not limited to: an accelerometer 1011, a gyroscope 1012, a pressure sensor 1013, a fingerprint sensor 1014, an optical sensor 1015, and a proximity sensor 1016.

[0404] Accelerometer 1011 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 1000. For example, accelerometer 1011 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1001 can control touchscreen 1005 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1011. Accelerometer 1011 can also be used for games or for acquiring user motion data.

[0405] The gyroscope sensor 1012 can detect the orientation and rotation angle of the terminal 1000. The gyroscope sensor 1012, in conjunction with the accelerometer sensor 1011, can collect 3D motion data from the user on the terminal 1000. Based on the data collected by the gyroscope sensor 1012, the processor 1001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0406] The pressure sensor 1013 can be disposed on the side bezel of the terminal 1000 and / or on the lower layer of the touch display screen 1005. When the pressure sensor 1013 is disposed on the side bezel of the terminal 1000, it can detect the user's grip signal on the terminal 1000, and the processor 1001 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1013. When the pressure sensor 1013 is disposed on the lower layer of the touch display screen 1005, the processor 1001 can control the operable controls on the UI interface based on the user's pressure operation on the touch display screen 1005. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0407] The fingerprint sensor 1014 is used to collect a user's fingerprint. The processor 1001 identifies the user based on the fingerprint collected by the fingerprint sensor 1014, or vice versa. When the user's identity is identified as trusted, the processor 1001 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 1014 can be located on the front, back, or side of the terminal 1000. When the terminal 1000 has physical buttons or a manufacturer's logo, the fingerprint sensor 1014 can be integrated with the physical buttons or manufacturer's logo.

[0408] An optical sensor 1015 is used to collect ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the touch screen 1005 based on the ambient light intensity collected by the optical sensor 1015. Specifically, when the ambient light intensity is high, the display brightness of the touch screen 1005 is increased; when the ambient light intensity is low, the display brightness of the touch screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 based on the ambient light intensity collected by the optical sensor 1015.

[0409] The proximity sensor 1016, also known as a distance sensor, is typically mounted on the front panel of the terminal 1000. The proximity sensor 1016 is used to detect the distance between the user and the front of the terminal 1000. In one embodiment, when the proximity sensor 1016 detects that the distance between the user and the front of the terminal 1000 is gradually decreasing, the processor 1001 controls the touchscreen display 1005 to switch from a screen-on state to a screen-off state; when the proximity sensor 1016 detects that the distance between the user and the front of the terminal 1000 is gradually increasing, the processor 1001 controls the touchscreen display 1005 to switch from a screen-off state to a screen-on state.

[0410] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on terminal 1000 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0411] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application. The server 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including random access memory (RAM) 1102 and read-only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The server 1100 also includes a basic input / output system (I / O system) 1106 that facilitates the transfer of information between various devices within the computer, and a mass storage device 1107 for storing the operating system 1113, application programs 1114, and other program modules 1115.

[0412] The basic input / output system 1106 includes a display 1108 for displaying information and an input device 1109 for user input, such as a mouse or keyboard. Both the display 1108 and the input device 1109 are connected to the central processing unit 1101 via an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include the input / output controller 1110 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides output to a display screen, printer, or other types of output devices.

[0413] Mass storage device 1107 is connected to central processing unit 1101 via a mass storage controller (not shown) connected to system bus 1105. Mass storage device 1107 and its associated computer-readable media provide non-volatile storage for server 1100. That is, mass storage device 1107 may include computer-readable media (not shown) such as hard disk or CD-ROM drive.

[0414] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technologies, CD-ROM, DVD or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1104 and mass storage device 1107 described above can be collectively referred to as memory.

[0415] According to various embodiments of this application, server 1100 can also be connected to a remote computer on a network, such as the Internet. That is, server 1100 can be connected to network 1112 via network interface unit 1111 connected to system bus 1105, or it can also use network interface unit 1111 to connect to other types of networks or remote computer systems (not shown).

[0416] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.

[0417] This application also provides a computer-readable storage medium storing instructions that, when executed on a terminal, cause the terminal to perform the steps of the video editing method described in the above embodiments.

[0418] This application also provides a computer-readable storage medium storing instructions that, when executed on a server, cause the server to perform the steps of the video editing method described in the above embodiments.

[0419] This application also provides a computer program product containing instructions that, when executed on a terminal, cause the terminal to perform the steps of the video editing method described in the above embodiments. Alternatively, a computer program is provided that, when executed on a terminal, causes the terminal to perform the steps of the video editing method described in the above embodiments.

[0420] This application also provides a computer program product containing instructions that, when executed on a server, cause the server to perform the steps of the video editing method described in the above embodiments. Alternatively, a computer program is provided that, when executed on a server, causes the server to perform the steps of the video editing method described in the above embodiments.

[0421] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.

[0422] It should be understood that "multiple" as mentioned herein refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., do not necessarily imply that they are different.

[0423] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, multiple editing materials involved in the embodiments of this application were obtained with full authorization.

[0424] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A video editing method, characterized in that, Applied to a terminal, the method includes: Multiple clips, including videos or images, are sampled to obtain multiple sample frames. Determine the attribute information for each of the sampled frames, wherein the attribute information indicates the characteristics of the sampled frames; The attribute information of the plurality of sampling frames is sent to the server so that the server determines the target highlight template and the highlight sampling frame corresponding to each highlight lens among the plurality of highlight lenses included in the target highlight template, wherein the highlight sampling frame is the sampling frame that matches the highlight lens among the plurality of sampling frames; Receive the target specular template and the specular sampling frames corresponding to the plurality of specular lenses sent by the server; Based on the attribute information of the multiple clip materials and the multiple sampled frames, the target highlight template and the highlight sampled frames corresponding to the multiple highlight shots, the regular video segments and highlight video segments corresponding to the multiple clip materials are determined; Based on the regular video segments and the highlight video segments, a final video is generated after video editing of the multiple clips.

2. The method as described in claim 1, characterized in that, The step of determining the regular video segments and highlight video segments corresponding to the multiple clip materials based on the attribute information of the multiple sampled frames, the target highlight template, and the highlight sampled frames corresponding to the multiple highlight shots includes: Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses; Determine the regular shot corresponding to each regular sampling frame in the at least one set of regular sampling frames from the plurality of clip materials; By combining the regular shots corresponding to the regular sampling frames in the same group of the at least one group of regular sampling frames into a regular video segment, at least one of the regular video segments is obtained. Based on the target highlight template and the highlight sampling frame corresponding to each highlight shot, at least one highlight video segment is determined from the plurality of edited materials.

3. The method as described in claim 2, characterized in that, Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses, including: Based on the attribute information of the multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses, and the image content of the multiple candidate sampling frames is different; Based on the attribute information of the multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one of the candidate sampling frames; Based on the number of the at least one cluster set and the number of the plurality of highlight shots, the number of regular shots included in the final video is determined; Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain the at least one set of regular sample frames.

4. The method as described in claim 3, characterized in that, After sending the attribute information of the plurality of sampled frames to the server, the method further includes: Receive target background music sent by the server, the target background music including at least one music segment, each music segment corresponding to one shot; Determining the number of regular shots included in the final video clip based on the number of the at least one cluster set and the number of the plurality of highlight shots includes: Based on the number of the at least one cluster set, the number of the plurality of highlight shots, and the number of shots corresponding to each of the music segments, a target music segment is determined from the at least one music segment; The number of shots corresponding to the target music segment is determined as the number of shots in the final video clip; The difference between the number of shots in the final video and the number of the multiple highlight shots is determined as the number of regular shots included in the final video.

5. The method as described in claim 2, characterized in that, After determining at least one set of regular sampling frames from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses, based on the attribute information of the multiple sampling frames and the number of the multiple highlight lenses, the method further includes: The at least one set of regular sampling frames is sent to the server so that the service determines the video effect information corresponding to the final video based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot. Receive the video effect information sent by the server; The process of generating a final video clip from the multiple edited materials, based on the regular video segments and the highlight video segments, includes: The final video clip is generated based on the regular video clip, the highlight video clip, and the video effect information.

6. The method according to any one of claims 1-5, characterized in that, The process of generating a final video clip from the multiple edited materials, based on the regular video segments and the highlight video segments, includes: Display a video preview interface, which includes the regular video segment and the highlight video segment; In response to an editing operation on a target video segment, the editing operation is performed on the target video segment, wherein the target video segment is either the regular video segment or the highlight video segment; The final video clip is generated based on the edited regular video segments and highlight video segments.

7. A video editing method, characterized in that, Applied to a server, the method includes: The receiving terminal sends attribute information for each of the multiple sampled frames, which are obtained by sampling multiple clip materials, and the attribute information indicates the characteristics of the sampled frames; Based on the attribute information of the multiple sampling frames, a target highlight template and a highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template are determined from multiple candidate highlight templates. The highlight sampling frame is the sampling frame that matches the highlight lens among the multiple sampling frames. The target highlight template and the highlight sampling frame corresponding to each highlight lens are sent to the terminal so that the terminal can determine the regular video segment and highlight video segment corresponding to the multiple editing materials, and generate the final video clip obtained after video editing of the multiple editing materials based on the regular video segment and the highlight video segment.

8. The method as described in claim 7, characterized in that, Each of the candidate specular templates has matching constraints; The step of determining the target highlight template and the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template, based on the attribute information of the multiple sampling frames, includes: Based on the attribute information of each sampled frame, a matching score is determined between each sampled frame and each highlight lens in each candidate highlight template; Based on the matching score between each sampled frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, a plurality of first highlight templates and highlight sampled frames corresponding to each highlight lens in each first highlight template are determined from the plurality of candidate highlight templates. At least one first highlight template is selected from the plurality of first highlight templates as the target highlight template, and the highlight sampling frame corresponding to each highlight lens in the at least one first highlight template is used as the highlight sampling frame corresponding to each highlight lens in the target highlight template.

9. The method as described in claim 8, characterized in that, The attribute information of each sampled frame includes the feature label and semantic feature vector of the sampled frame, and each highlight shot in each candidate highlight template has label requirements and semantic feature vector; The step of determining the matching score between each sampled frame and each highlight shot in each candidate highlight template based on the attribute information of each sampled frame includes: For each candidate highlight template, based on the feature label of each sampled frame and the label requirement of each highlight shot in the candidate highlight template, the label score between each sampled frame and each highlight shot in the candidate highlight template is determined; Based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight shot in the candidate highlight template, a similarity score is determined between each sampled frame and each highlight shot in the candidate highlight template. Based on the label score and similarity score between each sampled frame and each highlight shot in the candidate highlight template, a matching score is determined between each sampled frame and each highlight shot in the candidate highlight template.

10. The method as described in claim 8, characterized in that, The step of determining multiple first highlight templates and highlight sampling frames corresponding to each highlight lens in each first highlight template from the multiple candidate highlight templates, based on the matching score between each sampling frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, includes: For each candidate highlight template, based on the matching score between each sampled frame and each highlight shot in the candidate highlight template and the matching constraints of the candidate highlight template, it is determined whether each highlight shot in the candidate highlight template has a matching sampled frame. If each highlight lens in the candidate highlight template has a matching sampling frame, and the number of sampling frames matching the target highlight lens in the candidate highlight template is multiple, then based on the attribute information of the multiple sampling frames matching the target highlight lens, the highlight sampling frame corresponding to the target highlight lens is determined, and the candidate highlight template is used as the first highlight template, and the target highlight lens is at least one highlight lens in the candidate highlight template.

11. The method as described in claim 8, characterized in that, Each candidate highlight template is categorized into a beginning highlight template, a middle highlight template, or a ending highlight template; The step of selecting at least one first highlight template from the plurality of first highlight templates as the target highlight template, and using the highlight sampling frame corresponding to each highlight shot in the at least one first highlight template as the highlight sampling frame corresponding to each highlight shot in the target highlight template, includes: The plurality of first highlight templates are filtered to obtain a plurality of second highlight templates, wherein the highlight sampling frames corresponding to the highlight lenses in each pair of second highlight templates do not overlap; If there are at least two beginning highlight templates and / or at least two ending highlight templates among the plurality of second highlight templates, then the beginning highlight template and the ending highlight template with the highest priority among the plurality of second highlight templates are selected. The highest priority opening highlight template and ending highlight template, as well as the middle highlight template among the plurality of second highlight templates, are determined as the target highlight template. The highlight sampling frame corresponding to each highlight shot in the highest priority opening highlight template and ending highlight template, as well as the middle highlight template among the plurality of second highlight templates, is determined as the highlight sampling frame corresponding to each highlight shot in the target highlight template.

12. The method as described in claim 7, characterized in that, After the receiving terminal sends the attribute information of each of the multiple sampling frames, the method further includes: Based on the attribute information of the multiple sampled frames, a target background music is determined from multiple candidate background music. The target background music includes at least one music segment, and each music segment corresponds to one shot. The target background music is sent to the terminal.

13. The method as described in claim 7, characterized in that, The method further includes: Receive at least one set of regular sampling frames sent by the terminal; Based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot in the target highlight template, determine the video effect information corresponding to the finished video; The video effect information is sent to the terminal.

14. A video editing method, characterized in that, The method includes: Multiple clips, including videos or images, are sampled to obtain multiple sample frames. Determine the attribute information for each of the sampled frames, wherein the attribute information indicates the characteristics of the sampled frames; Based on the attribute information of the multiple sampling frames, a target highlight template and a highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template are determined from multiple candidate highlight templates. The highlight sampling frame is the sampling frame that matches the highlight lens among the multiple sampling frames. Based on the attribute information of the multiple clip materials and the multiple sampled frames, the target highlight template and the highlight sampled frames corresponding to the multiple highlight shots, the regular video segments and highlight video segments corresponding to the multiple clip materials are determined; Based on the regular video segments and the highlight video segments, a final video is generated after video editing of the multiple clips.

15. The method as described in claim 14, characterized in that, Each of the candidate specular templates has matching constraints; The step of determining the target highlight template and the highlight sampling frame corresponding to each highlight lens among the multiple highlight lenses included in the target highlight template, based on the attribute information of the multiple sampling frames, includes: Based on the attribute information of each sampled frame, a matching score is determined between each sampled frame and each highlight lens in each candidate highlight template; Based on the matching score between each sampled frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, a plurality of first highlight templates and highlight sampled frames corresponding to each highlight lens in each first highlight template are determined from the plurality of candidate highlight templates. At least one first highlight template is selected from the plurality of first highlight templates as the target highlight template, and the highlight sampling frame corresponding to each highlight lens in the at least one first highlight template is used as the highlight sampling frame corresponding to each highlight lens in the target highlight template.

16. The method as described in claim 14, characterized in that, The step of determining the regular video segments and highlight video segments corresponding to the multiple clip materials based on the attribute information of the multiple sampled frames, the target highlight template, and the highlight sampled frames corresponding to the multiple highlight shots includes: Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses; Determine the regular shot corresponding to each regular sampling frame in the at least one set of regular sampling frames from the plurality of clip materials; By combining the regular shots corresponding to the regular sampling frames in the same group of the at least one group of regular sampling frames into a regular video segment, at least one of the regular video segments is obtained. Based on the target highlight template and the highlight sampling frame corresponding to each highlight shot, at least one highlight video segment is determined from the plurality of edited materials.

17. The method as described in claim 16, characterized in that, Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses, including: Based on the attribute information of the multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses, and the image content of the multiple candidate sampling frames is different; Based on the attribute information of the multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one of the candidate sampling frames; Based on the number of the at least one cluster set and the number of the plurality of highlight shots, the number of regular shots included in the final video is determined; Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain the at least one set of regular sample frames.

18. The method as described in claim 17, characterized in that, After determining the attribute information of each of the sampled frames, the method further includes: Based on the attribute information of the multiple sampled frames, a target background music is determined from multiple candidate background music. The target background music includes at least one music segment, and each music segment corresponds to one shot. Determining the number of regular shots included in the final video clip based on the number of the at least one cluster set and the number of the plurality of highlight shots includes: Based on the number of the at least one cluster set, the number of the plurality of highlight shots, and the number of shots corresponding to each of the music segments, a target music segment is determined from the at least one music segment; The number of shots corresponding to the target music segment is determined as the number of shots in the final video clip; The difference between the number of shots in the final video and the number of the multiple highlight shots is determined as the number of regular shots included in the final video.

19. A video editing device, characterized in that, The video editing device includes: The sampling module is used to sample multiple clip materials to obtain multiple sample frames, wherein the clip materials include videos or images; The first determining module is used to determine the attribute information of each of the sampled frames, wherein the attribute information indicates the characteristics of the sampled frames; The first sending module is used to send the attribute information of the plurality of sampling frames to the server, so that the server determines the target highlight template and the highlight sampling frame corresponding to each highlight lens among the plurality of highlight lenses included in the target highlight template, wherein the highlight sampling frame is the sampling frame that matches the highlight lens among the plurality of sampling frames; The first receiving module is used to receive the target highlight template and the highlight sampling frames corresponding to the plurality of highlight lenses sent by the server; The second determining module is used to determine the regular video segment and the highlight video segment corresponding to the multiple editing materials based on the attribute information of the multiple sampling frames, the target highlight template and the highlight sampling frames corresponding to the multiple highlight shots; The generation module is used to generate a final video clip obtained after video editing of the multiple clips based on the regular video clips and the highlight video clips.

20. The video editing apparatus as described in claim 19, characterized in that, The second determining module is specifically used for: Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses; Determine the regular shot corresponding to each regular sampling frame in the at least one set of regular sampling frames from the plurality of clip materials; By combining the regular shots corresponding to the regular sampling frames in the same group of the at least one group of regular sampling frames into a regular video segment, at least one of the regular video segments is obtained. Based on the target highlight template and the highlight sampling frame corresponding to each highlight shot, at least one highlight video segment is determined from the plurality of edited materials.

21. The video editing apparatus as described in claim 20, characterized in that, The second determining module is specifically used for: Based on the attribute information of the multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses, and the image content of the multiple candidate sampling frames is different; Based on the attribute information of the multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one of the candidate sampling frames; Based on the number of the at least one cluster set and the number of the plurality of highlight shots, the number of regular shots included in the final video is determined; Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain the at least one set of regular sample frames.

22. The video editing apparatus as described in claim 21, characterized in that, The video editing device also includes: The second receiving module is used to receive the target background music sent by the server. The target background music includes at least one music segment, and each music segment corresponds to a number of shots. The second determining module is specifically used for: Based on the number of the at least one cluster set, the number of the plurality of highlight shots, and the number of shots corresponding to each of the music segments, a target music segment is determined from the at least one music segment; The number of shots corresponding to the target music segment is determined as the number of shots in the final video clip; The difference between the number of shots in the final video and the number of the multiple highlight shots is determined as the number of regular shots included in the final video.

23. The video editing apparatus as described in claim 20, characterized in that, The video editing device also includes: The second sending module is used to send the at least one set of regular sampling frames to the server, so that the server determines the video effect information corresponding to the finished video based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight shot; The third receiving module is used to receive the video effect information sent by the server; The generation module is specifically used for: The final video clip is generated based on the regular video clip, the highlight video clip, and the video effect information.

24. The video editing apparatus as described in any one of claims 19-23, characterized in that, The generation module is specifically used for: Display a video preview interface, which includes the regular video segment and the highlight video segment; In response to an editing operation on a target video segment, the editing operation is performed on the target video segment, wherein the target video segment is either the regular video segment or the highlight video segment; The final video clip is generated based on the edited regular video segments and highlight video segments.

25. A video editing device, characterized in that, The video editing device includes: The first receiving module is used to receive attribute information of each of the multiple sampling frames sent by the terminal. The multiple sampling frames are obtained by sampling multiple clip materials, and the attribute information indicates the characteristics of the sampling frames. The first determining module is used to determine, based on the attribute information of the plurality of sampling frames, a target highlight template and a highlight sampling frame corresponding to each highlight lens among the plurality of highlight lenses included in the target highlight template from the plurality of candidate highlight templates, wherein the highlight sampling frame is the sampling frame that matches the highlight lens among the plurality of sampling frames; The first sending module is used to send the target highlight template and the highlight sampling frame corresponding to each highlight lens to the terminal, so that the terminal can determine the regular video segment and highlight video segment corresponding to the multiple editing materials, and generate a video clip obtained after video editing of the multiple editing materials based on the regular video segment and the highlight video segment.

26. The video editing apparatus as described in claim 25, characterized in that, Each of the candidate specular templates has matching constraints; The first determining module is specifically used for: Based on the attribute information of each sampled frame, a matching score is determined between each sampled frame and each highlight lens in each candidate highlight template; Based on the matching score between each sampled frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, a plurality of first highlight templates and highlight sampled frames corresponding to each highlight lens in each first highlight template are determined from the plurality of candidate highlight templates. At least one first highlight template is selected from the plurality of first highlight templates as the target highlight template, and the highlight sampling frame corresponding to each highlight lens in the at least one first highlight template is used as the highlight sampling frame corresponding to each highlight lens in the target highlight template.

27. The video editing apparatus as described in claim 26, characterized in that, The attribute information of each sampled frame includes the feature label and semantic feature vector of the sampled frame, and each highlight shot in each candidate highlight template has label requirements and semantic feature vector; The first determining module is specifically used for: For each candidate highlight template, based on the feature label of each sampled frame and the label requirement of each highlight shot in the candidate highlight template, the label score between each sampled frame and each highlight shot in the candidate highlight template is determined; Based on the semantic feature vector of each sampled frame and the semantic feature vector of each highlight shot in the candidate highlight template, a similarity score is determined between each sampled frame and each highlight shot in the candidate highlight template. Based on the label score and similarity score between each sampled frame and each highlight shot in the candidate highlight template, a matching score is determined between each sampled frame and each highlight shot in the candidate highlight template.

28. The video editing apparatus as described in claim 26, characterized in that, The first determining module is specifically used for: For each candidate highlight template, based on the matching score between each sampled frame and each highlight shot in the candidate highlight template and the matching constraints of the candidate highlight template, it is determined whether each highlight shot in the candidate highlight template has a matching sampled frame. If each highlight lens in the candidate highlight template has a matching sampling frame, and the number of sampling frames matching the target highlight lens in the candidate highlight template is multiple, then based on the attribute information of the multiple sampling frames matching the target highlight lens, the highlight sampling frame corresponding to the target highlight lens is determined, and the candidate highlight template is used as the first highlight template, and the target highlight lens is at least one highlight lens in the candidate highlight template.

29. The video editing apparatus as described in claim 26, characterized in that, Each candidate highlight template is categorized into a beginning highlight template, a middle highlight template, or a ending highlight template; The first determining module is specifically used for: The plurality of first highlight templates are filtered to obtain a plurality of second highlight templates, wherein the highlight sampling frames corresponding to the highlight lenses in each pair of second highlight templates do not overlap; If there are at least two beginning highlight templates and / or at least two ending highlight templates among the plurality of second highlight templates, then the beginning highlight template and the ending highlight template with the highest priority among the plurality of second highlight templates are selected. The highest priority opening highlight template and ending highlight template, as well as the middle highlight template among the plurality of second highlight templates, are determined as the target highlight template. The highlight sampling frame corresponding to each highlight shot in the highest priority opening highlight template and ending highlight template, as well as the middle highlight template among the plurality of second highlight templates, is determined as the highlight sampling frame corresponding to each highlight shot in the target highlight template.

30. The video editing apparatus as described in claim 25, characterized in that, The video editing device also includes: The second determining module is used to determine the target background music from multiple candidate background music based on the attribute information of the multiple sampled frames. The target background music includes at least one music segment, and each music segment corresponds to one shot. The second sending module is used to send the target background music to the terminal.

31. The video editing apparatus as described in claim 25, characterized in that, The video editing device also includes: The second receiving module is used to receive at least one set of conventional sampling frames sent by the terminal; The third determining module is used to determine the video effect information corresponding to the finished video based on the at least one set of regular sampling frames and the highlight sampling frames corresponding to each highlight lens in the target highlight template. The third sending module is used to send the video effect information to the terminal.

32. A video editing device, characterized in that, The device includes: The sampling module is used to sample multiple clip materials to obtain multiple sample frames, wherein the clip materials include videos or images; The first determining module is used to determine the attribute information of each of the sampled frames, wherein the attribute information indicates the characteristics of the sampled frames; The second determining module is used to determine, based on the attribute information of the plurality of sampling frames, a target highlight template and a highlight sampling frame corresponding to each highlight lens among the plurality of highlight lenses included in the target highlight template from the plurality of candidate highlight templates, wherein the highlight sampling frame is the sampling frame that matches the highlight lens among the plurality of sampling frames; The third determining module is used to determine the regular video segment and the highlight video segment corresponding to the multiple editing materials based on the attribute information of the multiple sampling frames, the target highlight template and the highlight sampling frames corresponding to the multiple highlight shots; The generation module is used to generate a final video clip obtained after video editing of the multiple clips based on the regular video clips and the highlight video clips.

33. The apparatus as claimed in claim 32, characterized in that, Each of the candidate specular templates has matching constraints; The second determining module is specifically used for: Based on the attribute information of each sampled frame, a matching score is determined between each sampled frame and each highlight lens in each candidate highlight template; Based on the matching score between each sampled frame and each highlight lens in each candidate highlight template, and the matching constraints of each candidate highlight template, a plurality of first highlight templates and highlight sampled frames corresponding to each highlight lens in each first highlight template are determined from the plurality of candidate highlight templates. At least one first highlight template is selected from the plurality of first highlight templates as the target highlight template, and the highlight sampling frame corresponding to each highlight lens in the at least one first highlight template is used as the highlight sampling frame corresponding to each highlight lens in the target highlight template.

34. The apparatus as claimed in claim 32, characterized in that, The third determining module is specifically used for: Based on the attribute information of the plurality of sampling frames and the number of the plurality of highlight lenses, at least one set of regular sampling frames is determined from the sampling frames other than the highlight sampling frames corresponding to the plurality of highlight lenses; Determine the regular shot corresponding to each regular sampling frame in the at least one set of regular sampling frames from the plurality of clip materials; By combining the regular shots corresponding to the regular sampling frames in the same group of the at least one group of regular sampling frames into a regular video segment, at least one of the regular video segments is obtained. Based on the target highlight template and the highlight sampling frame corresponding to each highlight shot, at least one highlight video segment is determined from the plurality of edited materials.

35. The apparatus as claimed in claim 34, characterized in that, The third determining module is specifically used for: Based on the attribute information of the multiple sampling frames, multiple candidate sampling frames are determined from the sampling frames other than the highlight sampling frames corresponding to the multiple highlight lenses, and the image content of the multiple candidate sampling frames is different; Based on the attribute information of the multiple candidate sampling frames, the multiple candidate sampling frames are clustered to obtain at least one cluster set, and each cluster set includes at least one of the candidate sampling frames; Based on the number of the at least one cluster set and the number of the plurality of highlight shots, the number of regular shots included in the final video is determined; Based on the number of regular shots and the number of candidate sample frames in each cluster set, a set of regular sample frames is selected from each cluster set to obtain the at least one set of regular sample frames.

36. The apparatus as claimed in claim 35, characterized in that, The device further includes: The fourth determining module is used to determine the target background music from multiple candidate background music based on the attribute information of the multiple sampled frames. The target background music includes at least one music segment, and each music segment corresponds to one shot. The third determining module is specifically used for: Based on the number of the at least one cluster set, the number of the plurality of highlight shots, and the number of shots corresponding to each of the music segments, a target music segment is determined from the at least one music segment; The number of shots corresponding to the target music segment is determined as the number of shots in the final video clip; The difference between the number of shots in the final video and the number of the multiple highlight shots is determined as the number of regular shots included in the final video.

37. A terminal, characterized in that, The terminal includes a memory and a processor, the memory being used to store a computer program, and the processor being configured to execute the computer program stored in the memory to implement the steps of the method according to any one of claims 1-6 or 14-18.

38. A server, characterized in that, The server includes a memory and a processor, the memory being used to store a computer program, and the processor being configured to execute the computer program stored in the memory to implement the steps of the method according to any one of claims 7-13 or 14-18.

39. A computer-readable storage medium, characterized in that, The storage medium stores instructions that, when executed on the terminal, cause the terminal to perform the steps of the method described in any one of claims 1-6 or 14-18.

40. A computer-readable storage medium, characterized in that, The storage medium stores instructions that, when executed on the server, cause the server to perform the steps of the method described in any one of claims 7-13 or 14-18.

Citation Information

Patent Citations

  • Live stream editing method, apparatus and device

    CN108833969A

  • Video editing method and device, server and storage medium

    CN110996112A