Video editing method, electronic equipment and storage medium
By determining the story arrangement logic and performing video editing based on the image-text multimodal recognition model and the GPT model, the problem of discontinuous logic in video segments in existing technologies is solved, achieving natural transitions and personalized editing effects.
Patent Information
- Application Number
- CN202411486492.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, video editing methods cannot meet users' personalized needs, resulting in video clips that are logically discontinuous and have unnatural transitions.
By acquiring video footage data to be edited, the story arrangement logic is determined, and the video footage data is edited based on this logic, including the selection and synthesis of video segments. The image-text multimodal recognition model and GPT model are used to identify video attributes and generate editing scripts.
It achieves logical continuity and natural transitions between video clips in video editing, improves video quality, and meets users' personalized editing needs.
Smart Images

Figure CN121924306A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia technology, and more particularly to a video editing method, electronic device, and storage medium. Background Technology
[0002] In related technologies, automatic editing of video clips through software or platforms often involves sequentially splicing detected highlight segments as the final edited output. This results in the selected segments having discontinuous logic and unnatural transitions, failing to meet users' personalized editing needs. Summary of the Invention
[0003] In view of this, embodiments of this application provide a video editing method, an electronic device, and a storage medium, aimed at improving the video quality of edited videos.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] In a first aspect, embodiments of this application provide a video editing method, including:
[0006] Obtain the video footage data to be edited;
[0007] Determine the story's logical structure;
[0008] The video material data is edited based on the story arrangement logic to obtain an edited video.
[0009] Secondly, embodiments of this application also provide an electronic device, the electronic device comprising: a processor and a memory for storing a computer program capable of running on the processor, wherein, when the processor is used to run the computer program, it executes the steps of the method described in the first aspect of embodiments of this application.
[0010] Thirdly, embodiments of this application also provide a computer storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect of embodiments of this application.
[0011] The technical solution provided in this application involves acquiring video material data to be edited; determining the story arrangement logic; and editing the video material data based on the story arrangement logic to obtain an edited video. In this way, video material data can be edited based on the determined story arrangement logic, making the video segments logically continuous and naturally connected, thereby effectively improving the video quality of the edited video. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating the video editing method according to an embodiment of this application;
[0013] Figure 2 This is a schematic diagram illustrating the principle of the video editing method in the application embodiments of this application;
[0014] Figure 3 This is a schematic diagram of the two-stage video editing process in an application embodiment of this application;
[0015] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0016] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.
[0018] This application provides a video editing method that can be applied to electronic devices with data processing capabilities, such as terminal devices and servers, or can be implemented by the cooperation of terminal devices and servers. Specifically, the terminal device can be a computer, smartphone, personal digital assistant (PDA), etc.; the server can be an application server or a web server. In actual deployment, the server can be a standalone server or a cluster server.
[0019] like Figure 1 As shown, the method includes:
[0020] Step 101: Obtain the video footage data to be edited.
[0021] Here, the acquired video footage data to be edited can be video data shot by the user, video data downloaded from the Internet, or a combination of both; this application embodiment does not limit this. The video footage data can be understood as a collection of multiple video segments, and each video segment can be understood as an image sequence composed of multiple frames.
[0022] Step 102: Determine the story arrangement logic.
[0023] Here, story arrangement logic can be understood as the structural framework of video editing. It is used to combine with the textual content of the video footage data to construct the editing script, thereby obtaining a video with natural shot transitions. Story arrangement logic includes one or more of the following: story arrangement, beginning and ending, story focus, shooting intention, camera movement, transitions, special effects, theme, and events. Among them, [story arrangement] refers to the overall story logic and rhythm arrangement of the video, such as what the beginning, development, and ending of the video present, and the overall focus of the video; [beginning and ending] indicates whether the video has a special beginning and ending, such as a quick review of the overall content at the beginning, or a beginning and end echoing each other, or a landscape shot with blank space at the end, etc. [story focus] indicates the video style, such as showing a scene of sports, a family outing, or beautiful scenery, etc. [shooting intention] indicates what the shooting intention of the video segment is, such as wanting to shoot a video of sports. [camera movement] indicates what kind of shooting method the video segment uses, such as first-person or third-person, following or fixed shooting, etc. [transitions and special effects] indicates the transition effects in the video segment; [theme] indicates the theme around which the video segment revolves, etc.
[0024] In one embodiment, the story arrangement logic can be determined based on a template. For example, users can upload their favorite video clips as templates, and the logic can be determined based on learning from these templates. Alternatively, multiple story arrangement logics can be preset for users to choose from, or user-input story arrangement logic can be directly accepted.
[0025] Step 103: Based on the story arrangement logic, the video material data is edited to obtain an edited video.
[0026] It is understood that the embodiments of this application can edit video material data based on a determined story arrangement logic, which can make the logical continuity and natural connection of video segments in the edited video, that is, the shot connection of the edited video is natural, thereby effectively improving the video quality of the edited video.
[0027] For example, the determination of story arrangement logic includes:
[0028] Obtain reference video clips uploaded by users and use the reference video clips as editing reference templates;
[0029] The editing reference template is analyzed, and the story arrangement logic is determined based on the analysis results.
[0030] Here, the electronic device can acquire reference video clips uploaded by the user as editing templates. The device analyzes these templates and determines the story arrangement logic required by the user based on the analysis results. Furthermore, it can determine the story arrangement logic for the current edit based on the user's interactive operations, thus meeting the user's personalized editing needs. Specifically, learning is performed based on the user-uploaded reference video clips. The algorithm divides the template video into segments of preset duration (e.g., 5 seconds), then extracts frames from all the segments and inputs them into a pre-trained attribute recognition model to obtain all attributes of the segment, including [event, shooting intention, camera movement, theme], etc. Alternatively, the entire user-uploaded template video can be input into a holistic attribute recognition model to generate overall video attributes, including [story arrangement, beginning and ending, story focus]. All of these attributes are output as text by the algorithm.
[0031] For example, the step of analyzing the clip reference template and determining the story arrangement logic based on the analysis results includes:
[0032] Perform overall video attribute recognition on the editing reference template to obtain the overall video attribute recognition result;
[0033] The story arrangement logic is determined based on the overall attribute recognition results of the video.
[0034] Here, the electronic device can perform overall video attribute recognition on the editing reference template based on a pre-set AI (artificial intelligence) algorithm, obtain the overall video attribute recognition result, and determine the story arrangement logic based on the overall video attribute recognition result. The AI algorithm can employ a multimodal image-text recognition model or a GPT (Generative Pre-Trained Transformer) model, which can automatically recognize the overall video attribute recognition result of the video data; this embodiment does not limit this approach.
[0035] For example, the overall video attribute recognition result includes one or more of the following: a first descriptive text representing the story arrangement, a second descriptive text representing the story beginning and / or the story ending, and a third descriptive text representing the story's focus. In one application example, the overall video attribute recognition result includes the aforementioned first descriptive text, second descriptive text, and third descriptive text, thus determining the story arrangement logic, including the story arrangement, the story beginning and / or the story ending, and the story's focus.
[0036] For example, story arrangement can represent the overall narrative logic and pacing of the video, such as what is presented at the beginning, development, and end, and the overall focus of the video. The story beginning and / or ending can indicate whether the video has a special opening and / or closing, such as a quick overview of the overall content at the beginning, a beginning-and-end echoing each other, or a scenic shot with minimal background at the end. Story focus can represent the video's style, such as a significant portion of the video needing to highlight action scenes, family outings, or beautiful scenery.
[0037] In some embodiments, determining the story arrangement logic includes:
[0038] Receive story arrangement constraint information input by the user, and determine story arrangement logic based on the story arrangement constraint information; or...
[0039] The story arrangement logic is determined based on the preset arrangement logic.
[0040] In one application example, an electronic device can determine story arrangement logic based on story arrangement constraints input by the user. Here, the story arrangement constraints input by the user can be the content of voice or text input. The story arrangement constraints can include at least one of the aforementioned story arrangement, story beginning and / or story ending, and story emphasis. In other words, the electronic device can determine one or more of the aforementioned story arrangement, story beginning and / or story ending, and story emphasis based on the content of the user's voice or text input, and then generate story arrangement logic.
[0041] In another application example, if the electronic device does not obtain the story arrangement constraint information input by the user, nor does it obtain the reference video clip uploaded by the user, the electronic device can also determine the story arrangement logic of the current video clip based on the preset story arrangement logic. For example, the story arrangement logic of the previous video clip can be saved as the default story arrangement logic. In the case that the user has not input personalized story arrangement logic information, the video clip can be edited based on the default story arrangement logic. In this way, the video clipping is highly intelligent and can take into account the editing needs of different users. For example, it can be compatible with the editing needs of users who have set the default story arrangement logic.
[0042] For example, the step of editing the video material data based on the story arrangement logic to obtain an edited video includes:
[0043] Obtain the structured text data of the video material;
[0044] The editing script is determined based on the story arrangement logic and the structured text data;
[0045] The video footage data is edited based on the editing script to obtain an edited video.
[0046] Here, electronic devices can convert video footage data into structured text data based on algorithms that recognize video content. For example, they can use image-text multimodal recognition models or GPT models to identify the text content corresponding to the video footage data, obtaining predefined structured text data. Based on this, the electronic devices can determine an editing script based on the current story arrangement logic and the structured text data, and then edit the video footage data based on the editing script to obtain an edited video. The editing script can be understood as a description file obtained by processing the structured text data according to the story arrangement logic based on a text story arrangement algorithm. This editing script can provide text descriptions of each shot in the edited video. Thus, based on the editing script, video segments can be selected and sequentially connected, making the logical continuity and natural transitions of the video segments in the edited video possible.
[0047] For example, the editing script includes: text descriptions of multiple shots, and the arrangement order of the multiple shots; the editing process of the video material data based on the editing script includes:
[0048] Matching the text description and structured text data of each shot to determine the target structured text fragment that matches the text description of each shot.
[0049] Obtain the target video segment corresponding to each target structured text segment;
[0050] The target video segments are synthesized according to the arrangement order to obtain an edited video.
[0051] Understandably, electronic devices can match the text descriptions of each shot in the editing script with structured text data to obtain target structured text fragments that match the text descriptions of each shot in the editing script. Then, the target structured text fragments are mapped to the corresponding target video fragments, and the video fragments are synthesized according to the arrangement order to obtain the edited video.
[0052] For example, the step of matching the text description and structured text data of each shot to determine the target structured text fragment that matches the text description of each shot includes:
[0053] Semantic matching is performed between the text description of each shot and the text description in the structured text data to obtain the target structured text fragment that matches the text description of each shot. The timestamp information corresponding to the target structured text fragment is determined. The timestamp information refers to the start time and end time of the target structured text fragment in the structured text data.
[0054] The step of obtaining the target video segment corresponding to each target structured text segment includes:
[0055] Based on the timestamp information, the target video segment corresponding to each target structured text segment is extracted from the video material data.
[0056] Understandably, the electronic device performs semantic matching between the text description of each shot and the text description in the structured text data. After obtaining the target structured text fragment that matches the text description of each shot, it can also extract the timestamp information corresponding to the target structured text fragment. Thus, the electronic device can extract the target video fragment corresponding to each target structured text fragment from the video footage data based on this timestamp information. It should be noted that the structured text data of the video footage data includes multiple structured text fragments, where each structured text fragment corresponds to a video fragment, and each video fragment has corresponding timestamp information.
[0057] For example, determining the edited script based on the story arrangement logic and the structured text data includes:
[0058] The story arrangement logic and the structured text data are used as inputs to a pre-trained script arrangement model to obtain the edited script output by the script arrangement model.
[0059] Here, the story arrangement algorithm can be a pre-trained script arrangement model. The electronic device will use the story arrangement logic and the structured text data as input to the pre-trained script arrangement model to obtain the edited script output by the script arrangement model. The pre-trained script arrangement model can be trained based on a pre-constructed training sample set, which includes the story arrangement logic, structured text data, and corresponding edited videos.
[0060] Exemplarily, the method further includes:
[0061] The video material data is converted into multiple video segments based on a set duration;
[0062] The multiple video segments are subjected to segment filtering processing to obtain at least one video segment library;
[0063] The process of editing the video material data based on the story arrangement logic to obtain an edited video includes:
[0064] Based on the story arrangement logic, the at least one video clip library is subjected to video synthesis processing to obtain an edited video.
[0065] Here, before editing the video footage, the electronic device can also perform segment filtering on the video clips in the video footage data. This can improve editing efficiency and better meet the user's personalized editing needs.
[0066] For example, the step of performing segment filtering processing on the plurality of video segments to obtain at least one video segment library includes:
[0067] Perform segment attribute recognition on the multiple video segments to obtain the segment attribute recognition results for each video segment;
[0068] Based on the segment attribute recognition results of each video segment, waste segment filtering and / or highlight segment extraction are performed on the multiple video segments to obtain at least one video segment library.
[0069] Here, after the electronic device divides the video footage data to be edited into multiple video segments, it can perform segment attribute recognition on each video segment and, based on the segment attribute recognition results, filter out unusable segments and / or extract highlight segments, thereby obtaining a video segment library composed of video segments that more closely match the editing needs. This reduces the amount of subsequent data processing and improves editing efficiency. This segment attribute recognition is based on image feature recognition of image data. Different image features are pre-learned to correspond to attribute categories, and subsequent processing is determined based on these attribute categories. Attribute categories can include: occlusion categories, image blur categories, jitter categories, dirt categories, etc. In another embodiment, attribute categories can be further divided into: unusable segment categories, highlight categories, theme categories, etc. The specific division method can be customized as needed.
[0070] For example, the step of performing segment attribute recognition on the plurality of video segments to obtain segment attribute recognition results for each video segment includes:
[0071] The image data of each video segment in the multiple video segments are used to identify attributes based on a pre-trained image-text multimodal model to obtain the segment attribute identification results of each video segment.
[0072] The image-text multimodal model is configured with predefined attribute categories and corresponding text attribute definition information. The image-text multimodal model is used to encode the image data to obtain image encoding features, and output fragment attribute recognition results based on the image encoding features and the text encoding features of the text attribute definition information.
[0073] It is understandable that since electronic devices need to perform attribute recognition on image data, the use of a text-image multimodal model can achieve cross-modal semantic understanding. This allows for cross-modal understanding based on the image encoding features of image data and the text encoding features of predefined attribute categories, thereby obtaining attribute recognition results. In this way, the attribute categories of image data can be defined based on text attribute definition information. Traditional image / video-image / video matching models often require matching against an image / video standard library (covering all image / video content in multiple cases for each type). During inference, the hardware needs to pre-store all image / video encoding content of the entire standard library. Based on this, the text-image multimodal model introduced in this application only needs to pre-store the unique text description encoding for each attribute type in the hardware, improving the convenience of hardware storage and retrieval. Furthermore, compared to previous classification models that extract features from videos and use classifiers for judgment, which fix the number of model categories and discrimination criteria after training, the image-text multimodal model introduced in this application can not only modify the current attribute discrimination criteria by directly modifying the text description encoding for discrimination, but also easily expand and add the discrimination ability of other attributes, thus having the characteristics of easy expansion and adjustment.
[0074] For example, the image data is multi-frame image data obtained by extracting frames from a video segment of a set duration, and the step of encoding the image data to obtain image encoding features includes:
[0075] Based on the image encoder, the multi-frame image data is encoded as a whole to obtain the image coding features of the multi-frame image data, or...
[0076] The image encoder encodes each frame of the multi-frame image data separately to obtain the image coding features of each frame.
[0077] Here, each video segment can be divided based on a set duration, and the image data corresponding to each video segment can be multi-frame image data obtained after frame extraction of the video segment.
[0078] Here, a set of images after frame extraction is used as the input data of the model and input into the image feature encoder for image feature extraction. Two extraction methods can be selected: Method 1, extract unique image features for all the image data after frame extraction in the set, that is, obtain unique image coding features for multi-frame image data; Method 2, extract image features for each frame of all the image data after frame extraction in the set, that is, obtain image coding features for each frame of image data after frame extraction.
[0079] For example, the text encoding feature output attribute recognition result based on the image encoding features and the text attribute definition information includes:
[0080] Based on the similarity between the image coding features of the multi-frame image data or the image coding features of each frame image data and the text coding features of each attribute category, the attribute recognition result is output; or,
[0081] Based on the autoregressive generator, the image coding features of the multi-frame image data or the image coding features of each frame image data and the text coding features of each attribute category are processed to output the attribute recognition results.
[0082] It is understood that, based on the multimodal model of the present application, the segment attribute recognition results of each video segment in the video material data to be edited can be identified, and on this basis, waste clip filtering and / or highlight extraction can be performed to obtain at least one video segment library. Thus, the editing method of the present application can be understood as including two-stage editing, namely, the first stage of editing including segment selection and the second stage of editing including story arrangement of the video segment library after segment selection. The number of video segments to be edited can be reasonably controlled, and video editing based on the video segment library can improve the quality of video editing.
[0083] It should be noted that, in order to meet the diverse needs of users for video editing, a collection of video clips with multiple themes can be extracted from the video material data to be edited, thereby satisfying diverse editing requirements. Based on this, in some embodiments, a step of clustering the video clips based on theme attributes can also be introduced; that is, clip selection can also include clustering processing.
[0084] For example, prior to performing waste filtering and / or highlight extraction on the plurality of video clips, the method further includes:
[0085] Clustering processing is performed on the multiple video segments;
[0086] The step of filtering out unusable video clips and / or extracting highlight clips from the plurality of video clips to obtain at least one video clip library includes:
[0087] The clustered video clip sets are then subjected to waste clip filtering and / or highlight clip extraction to obtain at least one video clip library.
[0088] Understandably, in this application example, the video footage data to be edited is first clustered into video segments, and then the fragment sets of each cluster are filtered for unusable segments and / or highlight segments are extracted to obtain at least one video segment library.
[0089] For example, after performing waste filtering and / or highlight extraction on the plurality of video clips, the method further includes:
[0090] After filtering out the waste footage and / or extracting the highlight segments, the video segments are clustered to obtain at least one video segment library.
[0091] Understandably, in this application example, video clips in the video material data to be edited can first be filtered for defective clips and / or have highlights extracted, and then the processed video clips can be clustered to obtain at least one video clip library.
[0092] For example, the fragment attribute recognition result includes: a first recognition result characterizing the topic attribute; the clustering process includes:
[0093] Based on the first identification result, each video segment is clustered to obtain a set of video segments with at least one clustered theme and a set of video segments without a theme.
[0094] Here, the theme attribute represents the theme surrounding the video segment, such as including but not limited to the following theme categories: scene, event, person, time, etc. Each theme category includes at least one theme attribute. Thus, video segments with different theme attributes can be divided based on the first recognition result.
[0095] For example, if the electronic device does not collect parameters indicating the thematic tendency of the video clips, it can perform clustering processing on the video segments based on the first recognition result to obtain a set of video segments without a theme and a set of video segments with at least one clustered theme. The number and thematic tendency of the at least one clustered theme can be determined using preset parameters. In this way, multiple video segment libraries can be output for users to subsequently select target video segment libraries for editing, thereby enriching the diversity of video editing and better meeting users' personalized editing needs.
[0096] For example, the fragment attribute recognition result includes: a first recognition result characterizing the topic attribute; the method further includes:
[0097] Get the first parameter that indicates the thematic tendency of the video clip;
[0098] The target topic tendency is determined based on the first parameter;
[0099] The clustering process includes:
[0100] Based on the first identification result, a set of target video segments related to the target theme is extracted.
[0101] Understandably, if an electronic device collects a first parameter indicating the thematic tendency of a video clip—that is, if the user inputs a first parameter indicating the thematic tendency of the video clip—the electronic device can determine the target thematic tendency based on this first parameter, and then perform clustering processing on the video segments based on this first recognition result to extract a set of target video segments related to the target thematic tendency, thus obtaining a video segment library corresponding to the target thematic tendency. In this way, a video clip corresponding to the thematic tendency desired by the user can be directly obtained.
[0102] For example, the step of performing video synthesis processing on the at least one video clip library based on the story arrangement logic to obtain an edited video includes:
[0103] Select a target video clip library from at least one of the video clip libraries;
[0104] Obtain the structured text data of the target video clip library;
[0105] The editing script is determined based on the story arrangement logic and the structured text data;
[0106] Based on the editing script, the target video clip library is edited to obtain an edited video.
[0107] Understandably, if the user has already input a first parameter indicating the thematic tendency of the video clip, the electronic device can directly obtain the target video clip library based on clip filtering. If the electronic device has not acquired this first parameter, it can output indication information for multiple generated video clip libraries and receive interactive operations input by the user based on this indication information, selecting the target video clip library based on the interactive operations. The electronic device then acquires the structured text data of the target video clip library, determines the editing script based on the story arrangement logic and this structured cultural data, and edits the target video clip library based on the editing script to obtain the edited video. The specific process of determining the editing script and editing the target video clip library can be referred to the aforementioned description of editing video material data based on story arrangement logic, and will not be repeated here.
[0108] For example, the segment attribute recognition result of a video clip may further include: a second recognition result characterizing the attributes of discarded clips and a third recognition result characterizing the text description. The aforementioned discarded clip filtering and / or highlight segment extraction based on the segment attribute recognition result includes:
[0109] For multiple video clips, waste clip filtering and / or highlight clip extraction are performed based on at least one of the first recognition result, the second recognition result, and the third recognition result.
[0110] Here, the electronic device can identify defective video clips from multiple video segments based on the second recognition result, filter the identified defective video clips, and then extract highlight segments from the filtered video segments based on the first and / or third recognition results. For example, the electronic device first performs clustering processing on multiple video clips in the video material data to be edited, obtaining a set of clips corresponding to multiple theme categories, and then performs defective clip filtering and / or highlight segment extraction on each set of clips to obtain a video clip library corresponding to each theme category.
[0111] The present application will be further described in detail below with reference to application examples.
[0112] This application example provides a personalized video editing method, referencing... Figure 2 The method includes an editing preparation stage and an editing execution stage. The editing preparation stage includes receiving the user-uploaded video material library (i.e., the video material data to be edited) and determining the story arrangement logic. The editing execution stage includes processing the video material data based on the story arrangement logic to obtain the edited video.
[0113] In this application embodiment, after receiving the video material library uploaded by the user, the electronic device can perform segment attribute recognition on each video segment in the video material library (e.g., dividing the video segment into segments every 3 seconds) to obtain a first recognition result, a second recognition result, and a third recognition result for each video segment. The first recognition result corresponds to a theme attribute, the second recognition result corresponds to a discarded segment attribute, and the third recognition result corresponds to a text description. Based on the second recognition result, the electronic device automatically determines the discarded segment attribute to obtain a usable segment material library. For example, based on the second recognition result, video segments with occlusion, jitter, or blurriness are filtered out as discarded segments from the initial video segments to obtain a usable segment material library.
[0114] It's important to note that during the editing preparation stage, the electronic device can also receive input from the user via touch, text, or voice interaction, including: the desired final cut time, and suggestions for a thematic cut (e.g., wanting to edit around a specific scene or event, providing a script, specifying scenes or content to be removed, or requiring the addition of a scene or content). Furthermore, it can receive pre-edited clips uploaded by the user for style reference. This involves receiving reference video clips and obtaining style reference information based on them. This style reference information includes the following: story structure, story beginnings and / or endings, and story emphasis. Understandably, the electronic device can determine the story's logical structure based on this style reference information.
[0115] It should be noted that the electronic device can replace all available segments in the available selection material library with algorithmically generated text descriptions for subsequent steps. Furthermore, all personalized user input will also be converted into text instructions for the next stage, including converting interactive commands into text and uploaded videos into text descriptions of the video's narrative, development, and specific content. This application example does not limit the text conversion; it can use any video captioning algorithm or speech-to-text algorithm.
[0116] In this application embodiment, if the user inputs the aforementioned preference prompt: then a topic search is first performed on all video clips, and video clips related to all topics are retrieved before video editing. If the user does not input the aforementioned preference prompt: then the topic attributes of all video clips are first clustered to generate a set of video clips without a topic and a set of video clips with at least one clustered topic, and then topic-free video editing and topic-related video editing are performed respectively.
[0117] In this application embodiment, video editing includes segment selection and story arrangement of the selected video segment library. The following example illustrates this process using the scenario where the user has not entered a theme preference:
[0118] Reference Figure 3 First, the user-uploaded video files are divided into segments to obtain multiple video fragments; then, attribute recognition is performed on each video fragment based on its attributes to obtain the first recognition result (corresponding to...). Figure 3 The useful part), the second recognition result (corresponding to) Figure 3 (in the context of "useless") and the third identification result (corresponding to) Figure 3 (Caption in the image). Based on the first identification result, attribute clustering is performed to obtain a set of video clips without a theme and a set of video clips with multiple clustered themes. Next, a two-stage story arrangement is performed on each set of video clips. The specific process is as follows:
[0119] Stage 1: Based on the first, second, and third recognition results, each video clip set is filtered for invalid clips and highlights are extracted, forming the video clip library for Stage 2. Here, if the user uploads a template video (i.e., a reference video clip), the clip attributes of the template video are used as the basis for highlight selection. If no template video is available, default highlight selection criteria are used (e.g., including but not limited to: focusing on moving scenes rather than still images, focusing on interactions between people, including hugs, kisses, entertainment, etc.).
[0120] Stage 2: Using a story arrangement algorithm, the video material text selected in Stage 1 is arranged into a complete story editing script according to the order of the beginning, development and ending of the complete story. Then, the segments are selected in sequence and the selected segments are video synthesized to obtain the edited video.
[0121] Specifically, in Phase 1, the video clips within each cluster are first arranged chronologically. For example, ten video clips are grouped together, resulting in n groups. The algorithm then filters out unwanted clips and selects highlight clips for each group. Finally, the selected video clip sets from all groups are merged. If the total number of clips in the set does not meet the condition at this point, and it exceeds a preset maximum value, it means that there are too many usable video clips. Phase 1 needs to be run multiple times until it falls within the preset maximum value range (this strategy is to prevent excessive user uploads or a large number of duplicate highlight clips from being selected, leading to too many clips being input for Phase 2 editing and making it impossible to consider all clips). If it is less than a preset minimum value, it means that there are too few usable video clips. In this case, Phase 1 is restarted, without filtering out unwanted clips, only highlight selection is performed, and then the result is directly output to Phase 2 (this strategy is to prevent the unwanted clip attribute from blocking a large number of videos, resulting in a very small number of editable video clips).
[0122] Understandably, the first stage described above can be performed multiple times, thereby effectively reducing the number of segments in the candidate segment set and thus effectively controlling the length of the final video. The second stage, based on the video segments output from the first stage, can be edited according to the story arrangement logic, making the logical continuity and natural transitions of the edited video segments, thereby effectively improving the video quality of the edited video.
[0123] It should be noted that during video editing, users can input personalized editing instructions using interactive methods such as touch, text, or voice. For example, such as... Figure 2As shown, users can input editing commands (such as deleting or adding video clips) for various video clip sets (with or without a theme) through human-computer interaction. The commands are then used to edit the video based on a storytelling logic. For example, a user can select the original video clips to output and then delete unsatisfactory clips or add unselected clips through content retrieval, such as adding a clip about a girl skiing. Users can then search for or manually select desired clips from all automatically edited videos to add to the final edit. Alternatively, they can search for clips from all available video clips to add to the final product. Users can also input voice or text to have the algorithm automatically recommend the top 5 most suitable clips for addition, from which they can then select the desired clip. All interactive content is converted into text by the algorithm. By analyzing the user's interactive text, the algorithm transforms all personalized inputs into basic operations such as adding, deleting, modifying, and querying, calling the corresponding operation programs to complete the personalized interactive service. The editing algorithm takes into account the user-selected negatives, deleted segments, and retrieved added segments, considering the logic and rationality of the video. Following the chronological order and the sequence of events, it rationally plans and outputs the final interactive, automatically edited video without changing the content of the negatives.
[0124] It is understood that the method in this application embodiment, which performs video editing based on story arrangement logic, can incorporate the story-related arrangement logic into the editing process, compared to sequentially splicing together the highlight segments detected by the algorithm as the final edited video output. This results in logical continuity and natural transitions between video segments in the edited video. Furthermore, it supports outputting clips without a theme or multiple sets of clips with different themes for users to choose from, making the themes more diverse and better meeting users' editing needs. Moreover, by introducing human-computer interaction before and / or during editing, the editing style and content of the automatic editing can be personalized for users, better meeting the specific needs of a broad user base.
[0125] To implement the methods of the embodiments of this application, the embodiments of this application also provide an electronic device. Figure 4 The diagram shows only an exemplary structure of the electronic device, not the entire structure; implementation is possible as needed. Figure 4 The structure shown may be part or all of the structure.
[0126] like Figure 4As shown, the electronic device 400 provided in this application embodiment includes: at least one processor 401, a memory 402, a user interface 403, and at least one network interface 404. The various components in the electronic device 400 are coupled together via a bus system 405. It can be understood that the bus system 405 is used to implement communication between these components. In addition to a data bus, the bus system 405 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 4 The general designated all buses as Bus System 405.
[0127] The user interface 403 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.
[0128] The memory 402 in this embodiment is used to store various types of data to support the operation of the electronic device. Examples of such data include any computer program used to operate on the electronic device.
[0129] The video editing method disclosed in this application can be applied to or implemented by the processor 401. The processor 401 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the video editing method can be completed by the integrated logic circuitry of the hardware in the processor 401 or by instructions in software form. The processor 401 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 401 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in a storage medium, which is located in the memory 402. The processor 401 reads the information in the memory 402 and, in conjunction with its hardware, completes the steps of the video editing method provided in the embodiments of this application.
[0130] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.
[0131] It is understood that memory 402 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.
[0132] In an exemplary embodiment, this application also provides a computer storage medium, specifically a computer-readable storage medium, such as a memory 402 storing a computer program, which can be executed by a processor 401 of an electronic device to complete the steps described in the method of this application embodiment. The computer-readable storage medium can be a ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0133] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a processor 401 of an electronic device 400 to perform the steps described in the method of this application embodiment.
[0134] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0135] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0136] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video editing method, characterized in that, include: Obtain the video footage data to be edited; Determine the story's logical structure; The video material data is edited based on the story arrangement logic to obtain an edited video.
2. The method according to claim 1, characterized in that, The determination of the story arrangement logic includes: Obtain reference video clips uploaded by users and use the reference video clips as editing reference templates; The editing reference template is analyzed, and the story arrangement logic is determined based on the analysis results.
3. The method according to claim 2, characterized in that, The analysis of the editing reference template and the determination of the story arrangement logic based on the analysis results include: Perform overall video attribute recognition on the editing reference template to obtain the overall video attribute recognition result; The story arrangement logic is determined based on the overall attribute recognition results of the video.
4. The method according to claim 3, characterized in that, The overall video attribute recognition results include one or more of the following: a first descriptive text representing the story arrangement, a second descriptive text representing the beginning and / or end of the story, and a third descriptive text representing the focus of the story.
5. The method according to claim 1, characterized in that, The process of editing the video material data based on the story arrangement logic to obtain an edited video includes: Obtain the structured text data of the video material; The editing script is determined based on the story arrangement logic and the structured text data; The video footage data is edited based on the editing script to obtain an edited video.
6. The method according to claim 5, characterized in that, The editing script includes: text descriptions of multiple shots, and the order in which the shots are arranged; The editing process of the video material data based on the editing script includes: Matching the text description and structured text data of each shot to determine the target structured text fragment that matches the text description of each shot. Obtain the target video segment corresponding to each target structured text segment; The target video segments are synthesized according to the arrangement order to obtain an edited video.
7. The method according to claim 6, characterized in that, The matching of text descriptions and structured text data for each shot to determine the target structured text fragment that matches the text description of each shot includes: Semantic matching is performed between the text description of each shot and the text description in the structured text data to obtain the target structured text fragment that matches the text description of each shot. The timestamp information corresponding to the target structured text fragment is determined. The timestamp information refers to the start time and end time of the target structured text fragment in the structured text data. The step of obtaining the target video segment corresponding to each target structured text segment includes: Based on the timestamp information, the target video segment corresponding to each target structured text segment is extracted from the video material data.
8. The method according to claim 5, characterized in that, The process of determining the edited script based on the story arrangement logic and the structured text data includes: The story arrangement logic and the structured text data are used as inputs to a pre-trained script arrangement model to obtain the edited script output by the script arrangement model.
9. The method according to claim 1, characterized in that, The determination of the story arrangement logic includes: Receive story arrangement constraint information input by the user, and determine story arrangement logic based on the story arrangement constraint information; or... The story arrangement logic is determined based on the preset arrangement logic.
10. The method according to claim 1, characterized in that, The method further includes: The video material data is converted into multiple video segments based on a set duration; The multiple video segments are subjected to segment filtering processing to obtain at least one video segment library; The process of editing the video material data based on the story arrangement logic to obtain an edited video includes: Based on the story arrangement logic, the at least one video clip library is subjected to video synthesis processing to obtain an edited video.
11. The method according to claim 10, characterized in that, The step of performing segment filtering on the plurality of video segments to obtain at least one video segment library includes: Perform segment attribute recognition on the multiple video segments to obtain the segment attribute recognition results for each video segment; Based on the segment attribute recognition results of each video segment, waste segment filtering and / or highlight segment extraction are performed on the multiple video segments to obtain at least one video segment library.
12. The method according to claim 11, characterized in that, Before performing waste filtering and / or highlight extraction on the plurality of video clips, the method further includes: Clustering processing is performed on the multiple video segments; The step of filtering out unusable video clips and / or extracting highlight clips from the plurality of video clips to obtain at least one video clip library includes: For each video segment set after clustering, perform defective segment filtering and / or highlight segment extraction to obtain at least one video segment library; or, After performing waste filtering and / or highlight extraction on the plurality of video clips, the method further includes: After filtering out the waste footage and / or extracting the highlight segments, the video segments are clustered to obtain at least one video segment library.
13. The method according to claim 12, characterized in that, The fragment attribute recognition result includes: a first recognition result representing the topic attribute; the clustering process includes: Based on the first identification result, each video segment is clustered to obtain a set of video segments with at least one clustered theme and a set of video segments without a theme.
14. The method according to claim 12, characterized in that, The fragment attribute recognition result includes: a first recognition result characterizing the topic attribute; the method further includes: Get the first parameter that indicates the thematic tendency of the video clip; The target topic tendency is determined based on the first parameter; The clustering process includes: Based on the first identification result, a set of target video segments related to the target theme is extracted.
15. The method according to claim 11, characterized in that, The step of performing segment attribute recognition on the multiple video segments to obtain the segment attribute recognition result for each video segment includes: The image data of each video segment in the multiple video segments are used to identify attributes based on a pre-trained image-text multimodal model to obtain the segment attribute identification results of each video segment. The image-text multimodal model is configured with predefined attribute categories and corresponding text attribute definition information. The image-text multimodal model is used to encode the image data to obtain image encoding features, and output fragment attribute recognition results based on the image encoding features and the text encoding features of the text attribute definition information.
16. The method according to claim 10, characterized in that, The step of performing video synthesis processing on the at least one video clip library based on the story arrangement logic to obtain an edited video includes: Select a target video clip library from at least one of the video clip libraries; Obtain the structured text data of the target video clip library; The editing script is determined based on the story arrangement logic and the structured text data; Based on the editing script, the target video clip library is edited to obtain an edited video.
17. An electronic device, characterized in that, The electronic device includes: a processor and a memory for storing computer programs capable of running on the processor, wherein, The processor, when running a computer program, performs the steps of the method according to any one of claims 1 to 16.
18. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.