Video editing method and device, storage medium and program product

Automatically generate subtitles and accurately extract keywords through the AI video editing model, solving the problem of poor linkage of video editing effects in the existing technology, achieving efficient video editing and keyword special effects synchronization, and improving the attractiveness and user experience of the video.

CN120455805APending Publication Date: 2025-08-08BEIJING 58 INFORMATION TTECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510806352.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In existing video editing software, the display of keywords is easily out of sync with subtitle dynamic special effects, specific sound effects and video screen content, resulting in poor linkage of effects and inaccurate keyword selection, which affects user experience and editing efficiency.

Method used

The network architecture adopts the AI video editing model to automatically realize the automatic generation of subtitle files and the precise extraction of target keywords. Through the special effect generation network, dynamic special effect images and audio are generated based on keyword drivers to ensure that the special effects and keywords are played simultaneously.

Benefits of technology

It realizes automation of the video editing process, improves editing efficiency, enhances the synergy between video content and subtitles and keyword special effects, and enhances the overall attractiveness and user experience of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455805A_ABST
    Figure CN120455805A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video editing method and device, a storage medium and a program product. In the embodiment of the invention, by applying the network architecture of the AI video editing model, the automatic generation of the subtitle file, the accurate extraction of the target keyword and the special effect generation logic driven based on the keywords are automatically realized. The process involves creating dynamic special effect images and special effect audios related to each target keyword, and setting playing time intervals of the dynamic special effect images and the special effect audios with the aim of synchronous playing of the dynamic special effects and the target keywords. Therefore, automation of a video editing process can be realized, the requirements of users on rapid editing and batch processing of videos are met, the collaborative effect among video picture contents, subtitles and keyword special effects of the video picture contents is enhanced, and the overall attraction of the videos is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a video editing method, device, storage medium and program product. Background Art

[0002] With the rapid development of software technology, the complexity and functional diversity of video editing software are constantly increasing. At present, video editing software can add subtitles, subtitle dynamic effects and / or specific sound effects to audio and video files according to the actual needs of users. Generally, the method of adding subtitles, subtitle dynamic effects and / or specific sound effects to audio and video is as follows: the subtitle recognition function provided by the video software is used to identify the voice text in the video and add the voice text to the audio and video file as the subtitle of the audio and video file; the keywords contained in the subtitles can also be manually determined, and subtitle dynamic effects and specific sound effects can be manually set for the keywords to enrich the video effect.

[0003] However, with the above-mentioned video editing method, there may be a synchronization problem between the display of keywords and the keyword subtitle dynamic effects, specific sound effects and video screen content, resulting in poor linkage between the edited audio and video subtitles, subtitle dynamic effects, specific sound effects and video screen content, affecting the user experience; moreover, the selection of keywords is affected by subjective factors and may not be accurate enough, making the audio and video with the addition of keyword effects not eye-catching enough; in addition, the above-mentioned video editing method is less efficient and cannot meet the user's needs for fast editing of audio and video and large-scale batch editing of audio and video. Summary of the Invention

[0004] Multiple aspects of the present application provide a video editing method, device, storage medium and program product for automating the video editing process, which not only meets the user's needs for fast video editing and batch processing, improves the video editing efficiency, but also realizes the accurate extraction of keywords, enhances the synergy between video screen content, subtitles and their keyword special effects, and improves the overall attractiveness of the video and the user experience.

[0005] An embodiment of the present application provides a video editing method, comprising: in response to a video editing request, obtaining at least one initial video material, each initial video material having its own subject type; inputting each initial video material into a subtitle generation network in an AI video editing model, identifying the video content and / or audio clip of each initial video material, and generating a subtitle file for the initial video material based on the video content and / or audio clip, wherein each subtitle file includes: multiple subtitle line texts, a start time and an end time for each subtitle line text display; inputting the subject type of each initial video material and the subtitle file into a keyword extraction network in the AI video editing model, performing a keyword extraction operation on the subtitle file of each initial video material, and obtaining target keywords adapted to the subject type and the start time and end time of each target keyword displayed in the corresponding subtitle line text. End time, each target keyword corresponds to its own video screen; each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text are input into the special effect generation network in the AI video editing model, and the special effect generation logic driven by the target keyword is executed. Based on the video screen corresponding to each target keyword, the dynamic special effect parameters and audio special effect parameters corresponding to each target keyword are generated, and the dynamic special effect image and special effect audio associated with the target keyword are generated according to the dynamic special effect parameters and audio special effect parameters. Based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, the playback time interval of the dynamic special effect image and special effect audio is configured with the goal of playing the dynamic special effect image and special effect audio in conjunction with the target keyword.

[0006] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in the above-mentioned class inheritance relationship parsing method or code analysis method based on program execution.

[0007] An embodiment of the present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor implements each step of the above-mentioned class inheritance relationship parsing method or code analysis method based on program execution.

[0008] In an embodiment of the present application, by using the network architecture of the AI video editing model, the automatic generation of subtitle files, the accurate extraction of target keywords, and the special effects generation logic driven by these keywords are automatically realized. This process involves creating dynamic special effects images and special effects audio related to each target keyword, and setting the playback time interval of the dynamic special effects images and special effects audio with the purpose of playing these dynamic special effects synchronously with the target keywords. In this way, the automation of the video editing process can be achieved, which not only meets the user's demand for fast video editing and batch processing, improves the editing efficiency of the video, but also realizes the accurate extraction of keywords, enhances the synergy between the video picture content, subtitles and their keyword special effects, and improves the overall attractiveness of the video and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0010] Figure 1 A flowchart of a video editing method provided by an exemplary embodiment of the present application;

[0011] Figure 2 A schematic diagram of special effects of target keywords provided by an exemplary embodiment of the present application;

[0012] Figure 3 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0013] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0014] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0015] In addition, it should be noted that when the embodiments of the present application involve user interaction operations or triggering operations, the user interaction operations or triggering operations involved in the embodiments of the present application include but are not limited to: touch operations, gesture operations, voice operations, head movement operations, eye movement operations and other interactive operations in various ways; among which, touch operations include but are not limited to: click operations, double-click operations, long press operations, sliding operations, pinch operations or mouse hover operations, etc. Sliding operations include but are not limited to: straight sliding, curved sliding, etc.

[0016] Continuing from the above background technology, there may be a problem of asynchrony between the display of keywords and the subtitle dynamic special effects, specific sound effects and video screen content of the keywords, resulting in poor linkage between the edited audio and video subtitles, subtitle dynamic special effects, specific sound effects and video screen content, affecting the user experience; moreover, the selection of keywords is affected by subjective factors and may not be accurate enough, making the audio and video after adding keyword special effects not eye-catching enough; in addition, the above video editing method is inefficient and cannot meet the technical problems of users' needs for fast editing of audio and video and large-scale batch editing of audio and video. In the embodiment of the present application, by using the network architecture of the AI video editing model, the automatic generation of subtitle files, the accurate extraction of target keywords and the special effects generation logic driven by these keywords are automatically realized. This process involves creating dynamic special effects images and special effects audio related to each target keyword, and setting the playback time interval of dynamic special effects images and special effects audio with the purpose of playing these dynamic special effects synchronously with the target keywords. In this way, the video editing process can be automated, which not only meets the user's needs for fast video editing and batch processing, improves the video editing efficiency, but also realizes the accurate extraction of keywords, enhances the synergy between video content, subtitles and their keyword effects, and improves the overall appeal of the video and the user experience.

[0017] The above-mentioned solution provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings.

[0018] Figure 1 This is a flow chart of a video editing method provided in an embodiment of the present application. Figure 1 As shown, the method includes:

[0019] 101. In response to a video editing request, obtain at least one initial video material, each initial video material having a respective subject type;

[0020] 102. Input each initial video material into a subtitle generation network in the AI video editing model, identify the video content and / or audio clip of each initial video material, and generate a subtitle file for the initial video material based on the video content and / or audio clip. Each subtitle file includes: multiple subtitle lines, and the start and end time of each subtitle line display.

[0021] 103. Input the theme type and subtitle file of each initial video material into the keyword extraction network of the AI video editing model, perform keyword extraction on the subtitle file of each initial video material, obtain target keywords adapted to the theme type, and the start time and end time of each target keyword displayed in the corresponding subtitle line text, and each target keyword corresponds to its own video screen;

[0022] 104. Input each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text into the special effect generation network in the AI video editing model, execute the special effect generation logic driven by the target keyword, generate the dynamic special effect parameters and audio special effect parameters corresponding to each target keyword based on the video screen corresponding to each target keyword, generate the dynamic special effect image and special effect audio associated with the target keyword according to the dynamic special effect parameters and audio special effect parameters, and based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, configure the playback time interval of the dynamic special effect image and special effect audio with the goal of playing the dynamic special effect image and special effect audio in conjunction with the target keyword.

[0023] In an optional embodiment, the above-described video editing method can be applied to a video editing application. This embodiment does not limit the implementation form of the video editing application. The implementation form of the video editing application can be a standalone app, a mini-program that relies on the app to run, or a webpage.

[0024] In this embodiment, the video editing application can be, for example, Shotcut software, which supports video, audio, and image editing in multiple formats. Shotcut software is built on the Multimedia Lightweight Toolkit (MLT) framework, using the MLT framework as its video processing engine. Leveraging its powerful multimedia processing capabilities, it implements various video editing functions, allowing users to perform nonlinear editing via a timeline and supporting keyframes, filters, transitions, and other features. In this embodiment, Shotcut software is primarily used for local video editing and exports video editing parameters via MLT templates for subsequent batch video generation. The timeline is a visual tool used to display the temporal order of media materials such as video, audio, and images. Users can perform operations such as editing, moving, and reordering materials on the timeline. Nonlinear editing allows users to edit videos, add special effects, and adjust the order of videos at any point in time, without having to edit them in a fixed order. This editing method is more flexible and can greatly improve editing efficiency. The specific operations in Shotcut are as follows: users can freely edit video clips, add audio and image materials, and adjust the order and duration of materials via the timeline. For example, the user can cut video clip A into two parts, and then insert another video clip B into the middle of A.

[0025] In actual applications, a video editing application can provide a video editing configuration page, and users can submit video editing requests to the video editing application based on the video editing configuration page; accordingly, the video editing application performs corresponding video editing operations in response to the video editing request. For example, the video editing configuration page includes an upload entry for video materials, and users can upload at least one initial video material through the entry, and submit a video editing request to the video editing application in response to a trigger operation on an "upload" or "confirm" control. Alternatively, other applications different from the video editing application have an associated relationship with the video editing application, and the other applications are connected to the video editing application in communication, and users can submit video editing requests to the video editing application based on the other applications; the corresponding video editing application performs corresponding video editing operations in response to the video editing request. For example, a video editing configuration page is configured in the other application, and the video editing configuration page has an associated relationship with the video editing page. Based on the video material upload entry of the video editing configuration page, at least one initial video material can be uploaded to send a video editing request to the video editing application.

[0026] In another optional embodiment, the above video editing method can also be applied to the server side of a video editing application, where the front-end video editing application is used to provide a video editing configuration page, or other applications can directly send video editing requests to the server side of the video editing application, etc. The above embodiments are merely illustrative and do not limit the technical solutions of this application.

[0027] In an embodiment of the present application, in response to a video editing request, at least one video material can be obtained, and each video material has its own subject type. For the sake of distinction and description, the video material can be referred to as the initial video material. Among them, the initial video material refers to the various basic materials used for video production and editing. These materials can be pre-recorded videos, audio files, images, animations, subtitles, special effects, etc. Video material is an indispensable part of the video creation process. It provides a rich source of content for video editing, allowing creators to construct rich and colorful video works. In the following examples of this application, the initial video material focuses on video. In addition, this application does not limit the subject type of the initial video material. For example, the subject types include but are not limited to: advertising, education, entertainment, news and documentary. Furthermore, the subject types can be divided more finely. For example, the advertising category includes but is not limited to: housekeeping advertisements, housing advertisements and product advertisements.

[0028] In order to automate the video editing process and improve the efficiency of video editing, an AI video editing model is introduced in an embodiment of the present application, and the above-mentioned video editing method is performed by the AI video editing model. The AI video editing model includes but is not limited to: a subtitle generation network, a keyword extraction network, and a special effects generation network. The working principle of each network and its specific implementation method can be found in the relevant description of the following embodiments, which will not be repeated here. The AI video editing model is obtained by training the initial video editing model based on a large amount of sample data in the field of video editing, or by fine-tuning the pre-trained model. More specifically, it is obtained by training each network contained in the initial video editing model or by fine-tuning each network contained in the pre-trained model.

[0029] Taking the AI video editing model as an example, the process of training the initial video editing model based on sample data to obtain the AI video editing model is as follows: obtain a sample data set, the sample data set includes multiple sample video materials, sample video content and / or sample audio clips of each sample video material, sample subtitle files, sample keywords, and the sample start time and sample end time of each sample keyword displayed in the corresponding sample subtitle line text, the sample video screen corresponding to each sample keyword, the sample dynamic special effect parameters and sample audio special effect parameters corresponding to each sample keyword, and each sample keyword. The sample dynamic special effects images and sample special effects audio associated with the word, the sample playback time interval of the sample dynamic special effects images and sample special effects audio, etc.; further, the sample video content and / or sample audio clip of each sample video material will be identified, and the intermediate state subtitle file of the sample video material will be generated based on the intermediate state video content and / or intermediate state audio clip, and each intermediate state subtitle file will include: multiple intermediate state subtitle line texts, the intermediate state start time and the intermediate state end time displayed by each intermediate state subtitle line text; the intermediate state theme type of each sample video material and the intermediate state subtitle file will be input into the keyword extraction network in the initial video editing model, for The intermediate state subtitle file of each sample video material is subjected to keyword extraction operation to obtain the intermediate state keywords adapted to the intermediate state theme type and the intermediate state start time and intermediate state end time of each intermediate state keyword displayed in the corresponding intermediate state subtitle line text, and each intermediate state keyword corresponds to its own intermediate state video screen; each intermediate state keyword in the intermediate state subtitle file of each intermediate state video material, the intermediate state video screen corresponding to each intermediate keyword, and the intermediate state start time and intermediate state end time displayed in the corresponding sample subtitle line text are input into the special effect generation network in the initial video editing model, and the intermediate state keyword-based special effect generation network is executed. The word-driven special effects generation logic generates intermediate state dynamic special effects parameters and intermediate state audio special effects parameters corresponding to each intermediate state keyword based on the intermediate state video screen corresponding to each intermediate state keyword, generates intermediate state dynamic special effects images and intermediate state special effects audio associated with the intermediate state keyword according to the intermediate state dynamic special effects parameters and intermediate state audio special effects parameters, and configures the intermediate state playback time interval of the intermediate state dynamic special effects image and intermediate state special effects audio based on the intermediate state start time and intermediate state end time displayed in the corresponding intermediate state subtitle line text of the intermediate state keyword, with the goal of playing the intermediate state dynamic special effects image and intermediate state special effects audio in conjunction with the intermediate state keyword.Further, based on obtaining the sample data set, the sample video content and / or sample audio clip of each sample video material and the intermediate video content and / or intermediate audio clip, as well as the sample subtitle file and the intermediate subtitle file, as well as the sample keyword and the sample start time and sample end time of each sample keyword displayed in the corresponding sample subtitle line text and the intermediate state keyword and the intermediate state keyword displayed in the corresponding intermediate state subtitle line text, as well as the sample video picture corresponding to each sample keyword and the intermediate state video picture corresponding to each intermediate state keyword, as well as the sample dynamic special effect parameters and sample audio special effect parameters corresponding to each sample keyword. The model loss function is calculated based on the intermediate state dynamic special effect parameters and intermediate state audio special effect parameters corresponding to each intermediate state keyword, the sample dynamic special effect image and sample special effect audio associated with each sample keyword, the intermediate state dynamic special effect image and intermediate state special effect audio associated with each intermediate state keyword, and the sample playback time interval of the sample dynamic special effect image and sample special effect audio and the intermediate state playback time interval of the intermediate state dynamic special effect image and intermediate state special effect audio. If the model loss function does not meet the model training termination condition, the initial business opportunity analysis model continues to be trained until the model loss function meets the model training termination condition, thereby obtaining the AI video editing model.

[0030] Optionally, based on obtaining a sample data set, the sample video content and / or sample audio clip of each sample video material and the intermediate video content and / or intermediate audio clip, as well as the sample subtitle file and the intermediate subtitle file, as well as the sample keyword and the sample start time and sample end time of each sample keyword displayed in the corresponding sample subtitle line text and the intermediate state keyword and the intermediate state start time and intermediate state end time of the intermediate state keyword displayed in the corresponding intermediate state subtitle line text, as well as the sample video picture corresponding to each sample keyword and the intermediate state video picture corresponding to each intermediate state keyword, as well as the sample dynamic special effect parameters and sample audio special effect parameters corresponding to each sample keyword and the intermediate state dynamic special effect parameters and intermediate state audio special effect parameters corresponding to each intermediate state keyword, and the sample dynamic special effects images and sample special effects audio associated with each sample keyword, the intermediate dynamic special effects images and intermediate special effects audio associated with each intermediate keyword, as well as the sample playback time interval of the sample dynamic special effects images and sample special effects audio and the intermediate playback time interval of the intermediate dynamic special effects images and intermediate special effects audio, calculate the model loss function, and when the model loss function does not meet the model training termination condition, continue to train the initial business opportunity analysis model until the model loss function meets the model training termination condition, and obtain the AI video editing model, including: calculating the loss function of each network layer respectively, and taking these loss functions as the model loss function; or, performing weighted summation on the loss function calculated for each network layer respectively to obtain the model loss function.

[0031] In an embodiment of the present application, each initial video material can be input into the subtitle generation network in the AI video editing model, and the video content and / or audio clip of each initial video material can be identified. Based on the video content and / or audio clip, a subtitle file of the initial video material is generated. Each subtitle file includes: multiple subtitle line texts, the start time and end time of each subtitle line text display. Among them, the audio clip is a non-silent audio clip. By generating a subtitle file corresponding to each initial video material through the subtitle generation network in the AI video editing model, the accuracy of the content in the subtitle file can be improved.

[0032] In some embodiments, each initial video material is input into the subtitle generation network in the AI video editing model, and the picture content and / or audio clips of each initial video material are identified. The specific implementation of generating the subtitle file of the initial video material based on the picture content and / or audio clips includes but is not limited to the following methods.

[0033] Method 1:

[0034] Each initial video material is input into the subtitle generation network in the AI video editing model, and the audio segment and the video segment corresponding to each audio segment of each initial video material are identified to obtain multiple audio segments and the video segment corresponding to each audio segment. The video segment is a video segment with picture content in the video screen it contains, and each audio segment and its corresponding video segment have their own playback start time and end time; according to the picture content of the video segment corresponding to each audio segment, the initial subtitle text of the audio segment is optimized to obtain the target subtitle text; based on the subtitle line segmentation rule, the target subtitle text is segmented into at least one subtitle line text that is synchronously adapted to the audio content of the audio segment and / or the picture content of the corresponding video segment; based on each subtitle line text and its subtitle line sequence number and the start time and end time of each subtitle line text display, a subtitle file corresponding to the initial video material is generated.

[0035] In an optional embodiment, each initial video material is input into the subtitle generation network in the AI video editing model, and the audio segments and the video segments corresponding to each audio segment are identified for each initial video material to obtain multiple audio segments and the video segments corresponding to each audio segment, including: inputting each initial video material into the subtitle generation network in the AI video editing model, identifying the complete audio and video from each initial video material, and the audio contains non-silent segments and silent segments; wherein the non-silent segment refers to the part of the audio with obvious sound, including but not limited to: speech content, music, sound effects or any other perceptible sound, and the non-silent segment is the audio. The main part of the audio usually contains the core content of the audio. The focus of this application is on the speaking content of non-silent segments; silent segments refer to audio with no sound or very weak sound, including but not limited to: speaking pauses; identifying each silent period from the audio, and segmenting the audio based on the two endpoints of each silent period to obtain multiple segmented segments; screening non-silent segments from the multiple segmented segments as audio segments, each audio segment corresponding to its own playback time interval; based on the playback time interval corresponding to each audio segment, segmenting the video segment played in each playback time interval from the complete video as the video segment corresponding to each audio segment.

[0036] Furthermore, when segmenting the audio clip, the tone, pause duration, maximum number of words in the subtitle line, and video screen content can also be comprehensively considered to segment the audio clip. Among them, the tone parameter, pause parameter, maximum number of words in the subtitle line parameter, and video screen content are comprehensively considered to segment the audio clip. Among them, the tone parameter is a description of the tone of voice. For example, the questioning tone is usually at the end of a sentence and can be broken. The pause parameter is a description of the pause duration. If you speak fast, this parameter can be set larger, and if you speak slowly, this parameter can be set smaller. The maximum number of words in the subtitle line parameter is a description of the maximum number of words that can be displayed in a subtitle line. Therefore, the tone parameter, pause parameter, maximum number of words in the subtitle line parameter, and video screen content can be used as subtitle line segmentation rules.

[0037] In an optional embodiment, the initial subtitle text of each audio segment is optimized based on the screen content of the video segment corresponding to the audio segment to obtain the target subtitle text, including: text parsing the audio segment to obtain the initial subtitle text of the audio segment; parsing the video screen of the video segment corresponding to the audio segment to obtain the screen content of the video screen; generating subtitle guide words based on the screen content of the video screen; and optimizing the initial subtitle text of the audio segment under the guidance of the subtitle guide words to obtain the target subtitle text. The subtitle guide words include but are not limited to: guide words for guiding the adaptation of the subtitle text to the screen content of the video screen, guide words for correcting inaccurate descriptions, guide words for guiding the structure of the subtitle text, guide words for guiding the correct and standardized use of terms, and guide words for guiding the addition of corresponding modal particles based on the subject type.

[0038] In an optional embodiment, the number of characters in each subtitle line that the screen can carry is limited. The target subtitle text corresponding to an audio clip may have a large number of characters, and the screen cannot display these texts at the same time. In this case, the text needs to be segmented. The segmentation process corresponds to a subtitle line segmentation rule. Based on the subtitle line segmentation rule, the target subtitle text is segmented into at least one subtitle line text that is synchronously adapted to the audio content of the audio clip and / or the picture content of the corresponding video clip, including: identifying the semantic information of the target subtitle text, the semantic information contains the expression of meaning at different stages; based on the expression of meaning at different stages contained in the semantic information, the picture content of the corresponding video clip, and the maximum number of characters in the subtitle line that can be displayed on a single screen, the target subtitle text is segmented into at least one subtitle line text that is synchronously adapted to the audio content of the audio clip and / or the picture content of the corresponding video clip, and each subtitle line text corresponds to the expression of meaning at one stage. Wherein, a subtitle line text refers to a line of subtitle text displayed on a single screen. When segmenting the target subtitle text, the meaning of the target subtitle text at different stages, the picture content of the corresponding video clip and the maximum number of text words that can be displayed on the screen are taken into consideration, which can improve the accuracy of subtitle line text segmentation.

[0039] Method 2:

[0040] Each initial video material is input into the subtitle generation network in the AI video editing model, and video clips are identified based on the video file of each initial video material to obtain multiple video clips, each of which has its own playback start time and end time; based on the picture content of each video clip, subtitle text of each video clip is generated; based on the subtitle line segmentation rule, each subtitle text is segmented into at least one subtitle line text that is synchronized with the picture content of the video clip; based on each subtitle line text and its subtitle line sequence number and the start time and end time of each subtitle line text display, the subtitle file of the initial video material is generated.

[0041] The specific implementation methods of each step involved in Method 2 can be found in the relevant descriptions of each real-time method in Method 1, and will not be repeated here.

[0042] Method 3:

[0043] Each initial video material is input into the subtitle generation network in the AI video editing model. The audio segments are identified based on the audio file of the initial video material to obtain multiple audio segments, each of which has its own playback start and end time. Based on the subtitle line segmentation rule, the subtitle text of each audio segment is segmented into at least one subtitle line text that is synchronized with the audio content of the audio segment. The subtitle file of the initial video material is generated based on each subtitle line text, its subtitle line sequence number, and the start and end time of each subtitle line text display.

[0044] The specific implementation methods of each step involved in Method 3 can be found in the relevant descriptions of each real-time method in Method 1, and will not be repeated here.

[0045] After obtaining the subtitle file of each initial video material, in order to accurately identify the keywords contained in each subtitle file, the theme type of each initial video material and the subtitle file can be input into the keyword extraction network in the AI video editing model, and a keyword extraction operation is performed on the subtitle file of each initial video material to obtain the target keywords adapted to the theme type and the start time and end time of each target keyword displayed in the corresponding subtitle line text. Each target keyword corresponds to its own video screen.

[0046] It should be noted that, based on precise segmentation, the start time and end time of each audio segment are obtained by adding up the start time and end time of the previous audio segment and silent segment. The same is true for the start time and end time of the video segment and subtitle line text. This can improve the accuracy of determining each start time and end time, thereby improving the accuracy of subtitle line text and keywords.

[0047] In this embodiment, the theme type and subtitle file of each initial video material are input into the keyword extraction network in the AI video editing model, and a keyword extraction operation is performed on the subtitle file of each initial video material to obtain target keywords adapted to the theme type and the start time and end time of each target keyword displayed in the corresponding subtitle line text, including: inputting the theme type and subtitle file of each initial video material into the keyword extraction network in the AI video editing model, and determining the candidate keywords of each subtitle line text in each subtitle file based on the key semantic information of each subtitle line text in each subtitle file; from the candidate keywords of each subtitle line text, selecting the keyword adapted to the theme type of the corresponding initial video material as the target keyword for each subtitle line text in each subtitle file; and determining the start time and end time of each target keyword displayed in the corresponding subtitle line text according to the position information of each target keyword in the corresponding subtitle line text and the start time and end time of the subtitle line text display.

[0048] In an optional embodiment, the subject type of each initial video material and the subtitle file are input into the keyword extraction network in the AI video editing model, and based on the key semantic information of each subtitle line text in each subtitle file, the candidate keywords of each subtitle line text in each subtitle file are determined, including: inputting the subject type of each initial video material and the subtitle file into the keyword extraction network in the AI video editing model, performing a semantic information extraction operation on each subtitle line text in each subtitle file to obtain the key semantic information of each subtitle line text; selecting keywords from each subtitle line text in each subtitle file whose key semantic information is moderately higher than a preset threshold as candidate keywords.

[0049] In an optional embodiment, from the candidate keywords of each subtitle line text, keywords that are adapted to the theme type of the corresponding initial video material are selected as target keywords for each subtitle line text in each subtitle file, including: calculating the semantic similarity between each candidate keyword and the theme type of the corresponding initial video material; and using the candidate keywords with semantic similarity greater than a preset threshold as the target keywords for each subtitle line text in each subtitle file.

[0050] It should be noted that not all subtitle line texts contain keywords. When a subtitle long text contains keywords, the subtitle line text may contain one or more keywords.

[0051] After obtaining the target keywords of each subtitle line text in each subtitle file of each initial video file, in order to add dynamic special effects or audio special effects to each keyword and improve the linkage between each special effect and the corresponding video screen, each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword and the start time and end time of each target keyword displayed in the corresponding subtitle line text can be input into the special effects generation network in the AI video editing model, and the special effects generation logic driven by the target keyword is executed. Based on the video screen corresponding to each target keyword, the dynamic special effects parameters and audio special effects parameters corresponding to each target keyword are generated. According to the dynamic special effects parameters and audio special effects parameters, the dynamic special effects image and special effects audio associated with the target keyword are respectively generated. Based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, the playback time interval of the dynamic special effects image and special effects audio is configured with the goal of playing the dynamic special effects image and special effects audio in linkage with the target keyword. Among them, executing the special effect generation logic driven by the target keyword, that is, based on the video screen corresponding to each target keyword, generating the dynamic special effect parameters and audio special effect parameters corresponding to each target keyword, generating the dynamic special effect image and special effect audio associated with the target keyword according to the dynamic special effect parameters and audio special effect parameters, and based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, with the dynamic special effect image and special effect audio being played in conjunction with the target keyword, configuring the playback time interval of the dynamic special effect image and special effect audio. Among them, the dynamic effects can be, for example: magnification, flashing, rotation, etc. Special effect audio includes but is not limited to: reward sound, celebration sound, warning sound or falling sound.

[0052] In addition, the display position of the dynamic special effects of the target keyword can also be set. For example, the display position of the dynamic special effects can be set to the upper left corner, upper right corner, lower left corner, lower right corner or middle position of the screen. The specific position depends on the needs and is not limited in this embodiment. Figure 2 This is a diagram showing the dynamic effect of the keyword "30%".

[0053] In this embodiment, each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text are input into the special effect generation network in the AI video editing model, and the special effect generation logic driven by the target keyword is executed. Based on the video screen corresponding to each target keyword, the dynamic special effect parameters and audio special effect parameters corresponding to each target keyword are generated, and the dynamic special effect image and special effect audio associated with the target keyword are respectively generated according to the dynamic special effect parameters and audio special effect parameters, including: each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text are input into the special effect generation network in the AI video editing model, and according to each The emotional attribute information of the target keyword and the picture content attribute information of its corresponding video picture are used to determine the dynamic special effects template and audio special effects template of the target keyword; according to the dynamic special effects template and audio special effects template corresponding to the target keyword, the dynamic special effects parameters and audio special effects parameters corresponding to the target keyword are generated, the dynamic special effects parameters include multi-frame special effects parameters, and the audio special effects parameters include multi-note special effects parameters; the multi-frame special effects parameters corresponding to the target keyword are input into the image drawing tool, and under the guidance of the multi-frame special effects parameters, a multi-frame special effects image associated with the target keyword is generated, and based on the multi-frame special effects image, a dynamic special effects image associated with the target keyword is generated using a dynamic special effects graph processing tool; and the multi-note special effects parameters corresponding to the target keyword are input into the audio special effects processing tool, and under the guidance of the multi-note special effects parameters, special effects audio associated with the target keyword is generated.

[0054] In an optional embodiment, each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text are input into the special effect generation network in the AI video editing model, and the dynamic special effect template and audio special effect template of the target keyword are determined according to the emotional attribute information of each target keyword and the picture content attribute information of its corresponding video screen, including: determining the picture content of the video screen corresponding to each target keyword according to the start time and end time displayed by each target keyword; performing attribute analysis on each target keyword and the picture content of its corresponding video screen respectively to obtain the emotional attribute information of each target keyword and the picture content attribute information of the video screen corresponding to the target keyword; wherein the emotional attribute information includes: positive keywords and negative keywords words, positive keywords refer to words with positive meanings, including but not limited to: happy, reward, celebrate, cheer, etc.; negative keywords refer to words with negative meanings, including but not limited to: unlucky, sad, bad luck, etc.; picture content attribute information includes: positive picture content and negative picture content, positive picture content can be, for example, pictures with an atmosphere of sunshine, happiness, reward, celebration, cheer, etc., and negative picture content can be, for example, pictures with an atmosphere of dim light, unlucky, sad, bad luck, etc.; further, according to the emotional attribute information of each target keyword and the picture content attribute information of the video picture corresponding to the target keyword, determine the dynamic special effect attribute information and audio special effect attribute information adapted to the target keyword; further, according to the dynamic special effect attribute information and audio special effect attribute information of the target keyword, determine the dynamic special effect template and audio special effect template of the target keyword.

[0055] Optionally, dynamic special effects parameters and audio special effects parameters corresponding to the target keyword are generated based on the dynamic special effects template and audio special effects template corresponding to the target keyword, including: performing a frame-by-frame special effects parameter parsing operation based on the dynamic special effects template corresponding to the target keyword to generate multi-frame dynamic special effects parameters; and performing a note-by-note parameter parsing operation based on the audio special effects template corresponding to the target keyword to generate multi-note special effects parameters.

[0056] More specifically, based on the dynamic special effects template corresponding to the target keyword, a frame-by-frame special effects parameter parsing operation is performed to generate multi-frame dynamic special effects parameters, including: determining the target dynamic special effects library corresponding to the theme type according to the theme type of the initial video material and the theme type-dynamic special effects library mapping table; selecting a dynamic special effects template that is adapted to the emotional attribute information from the target dynamic special effects library according to the emotional attribute information of the target keyword, the dynamic special effects template containing multi-frame special effects; performing a frame-by-frame special effects parameter parsing operation on the dynamic special effects template to generate multi-frame dynamic special effects parameters.

[0057] More specifically, a note-by-note parameter parsing operation is performed based on the audio special effects template corresponding to the target keyword to generate multi-note audio special effects parameters, including: determining the target audio special effects library corresponding to the theme type based on the theme type of the initial video material and a theme type-audio special effects library mapping table; selecting an audio special effects template that is compatible with the emotional attribute information of the target keyword from the target audio special effects library, the audio special effects template containing multiple note special effects; and performing a note-by-note special effects parameter parsing operation on the audio special effects template to generate multi-note special effects parameters. Positive audio, such as reward sounds or celebratory sounds, can be inserted into positive keywords; negative audio, such as warning sounds or falling prompt sounds, can be inserted into negative keywords.

[0058] In an optional embodiment, the above-mentioned image drawing tool can be a canvas, and the dynamic special effects image processing tool can be an ffmpage tool; the multi-frame special effects parameters corresponding to the target keyword are input into the image drawing tool, and under the guidance of the multi-frame special effects parameters, a multi-frame special effects image associated with the target keyword is generated, and based on the multi-frame special effects image, a dynamic special effects image processing tool is used to generate a dynamic special effects image associated with the target keyword, including: inputting the multi-frame special effects parameters corresponding to the target keyword into the image drawing tool, and under the guidance of the multi-frame special effects parameters, generating a multi-frame special effects image associated with the target keyword according to each frame of special effects parameters in turn, and in the process of generating each frame of special effects image, using the ffmpage tool to record each frame of special effects image in turn to obtain a dynamic special effects image associated with the target keyword. This embodiment does not limit the format of the dynamic special effects image, for example, it can be a GIF format, MP4 format, APNG format, etc.

[0059] In an optional embodiment, the audio special effects processing tool may be Soundify, and the multi-note special effects parameters corresponding to the target keyword are input into the audio special effects processing tool, and under the guidance of the multi-note special effects parameters, special effects audio associated with the target keyword is generated, including: inputting the multi-note special effects parameters corresponding to the target keyword into the Soundify tool, and under the guidance of the multi-note special effects parameters, special effects audio associated with the target keyword is generated. Alternatively, the audio special effects processing tool may be an AI-based audio representation model, such as SeedFoley, and the multi-note special effects parameters corresponding to the target keyword are input into the audio special effects processing tool, and under the guidance of the multi-note special effects parameters, special effects audio associated with the target keyword is generated, including: inputting the multi-note special effects parameters corresponding to the target keyword into the SeedFoley model, and under the guidance of the multi-note special effects parameters, special effects audio associated with the target keyword is generated.

[0060] In an optional embodiment, each initial video source is equipped with a video track, a subtitle track, a dynamic special effects track, an audio special effects track, and a time track. These tracks are interconnected, and the start and end points of each track are aligned. As time progresses on the time track, the corresponding subtitle text or special effects are activated on the corresponding track. The dynamic special effects track can be a Lottie dynamic special effects track. Based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, with the goal of playing the dynamic special effects image and special effects audio in conjunction with the target keyword, the playback time interval of the dynamic special effects image and special effects audio is configured, including: in the display time interval formed by the start time and end time of the display of the target keyword on the time track, the playback time interval occupied by the video picture adapted to the dynamic special effects image and / or special effects audio on the time track is configured as the playback time interval of the dynamic special effects image and / or special effects audio of the target keyword, and the playback time interval includes the playback start time and end time; based on the playback time interval of the dynamic special effects image and / or special effects audio of the target keyword on the time track, in the vertical dimension, the dynamic special effects image and / or special effects audio of the target keyword are respectively configured in the corresponding position intervals of the dynamic special effects track and / or the audio special effects track as the activation position interval of the dynamic special effects image and / or the special effects audio.

[0061] Accordingly, the activation position interval of the video screen on the video screen playback track can be determined based on the time information provided by the time track equipped with the initial video material and the start playback time and end playback time of the video screen of the initial video material, so that the video screen can be activated and started to be played when the current time information on the time track represents the playback start time of the video screen, and the video screen can be stopped when the end time of the live video screen arrives; based on the start time and end time of the display of each subtitle line text, the activation position interval of the subtitle line text on the subtitle line track can be determined, so that the subtitle line text can be activated and started to be displayed when the current time information on the time track represents the playback start time of the subtitle line text, and the subtitle line text can be stopped when the end time of the live video screen arrives. Stop displaying the subtitle line text when the end time of the subtitle line text is reached; determine the activation position interval of the dynamic special effects image and / or special effects audio of the target keyword in the dynamic special effects track and / or audio special effects track respectively based on the start time and end time of the playback of the dynamic special effects image and / or special effects audio of each target keyword, so that the dynamic special effects image is activated and played when the current time information on the time track represents the playback start time of the dynamic special effects image, and the dynamic special effects image is stopped when the end time of the dynamic special effects image is reached; and / or, the special effects audio is activated and played when the current time information on the time track represents the playback start time of the special effects audio, and the special effects audio is stopped when the end time of the special effects audio is reached.

[0062] Additionally, you can set static effects for each subtitle line, such as font type, size, or color, so that each subtitle line displays the corresponding static effect. For more information on selecting font type, size, or color, refer to the selection methods for dynamic effects templates and audio effects templates, and will not be repeated here.

[0063] In this embodiment, the subtitle file is typically in SRT format, a widely used subtitle file format with a simple and clear structure. It contains a subtitle line sequence number, a start time, an end time, and the corresponding subtitle line text. The subtitle line sequence number is a number that identifies the order of each subtitle line, starting from 1 and increasing in sequence. It can be used to help the player display subtitles correctly in sequence and facilitate positioning and management when editing subtitle files. The start time and end time respectively indicate the start and end time of the subtitle line display. The timestamp format is HH:MM:SS, mmm, where HH represents hours, MM represents minutes, SS represents seconds, and mmm represents milliseconds. The player can use these timestamps to accurately display and hide subtitles during video playback. The subtitle text line is the text portion of the subtitle, typically a sentence or paragraph, which is displayed to the viewer during video playback to help the viewer understand the dialogue or content in the video. However, the SRT format itself does not support text styles (such as color, font size) or special effects and is only used for timeline synchronization and plain text display.

[0064] In this regard, the subtitle file can be converted into a format that supports text styles or special effects, such as converted into a rich text format. Correspondingly, the AI video editing model also includes a rich text generation network, and the subtitle file and each special effect file of each initial video material can be input into the rich text generation network, and the subtitle file and special effect file of each initial video material are converted into a rich text format to obtain an intermediate subtitle file and an intermediate special effect file corresponding to each initial video material. The intermediate subtitle file and intermediate special effect file in rich text format can be in JOSN format, but is not limited to this. In addition, the rich text format conversion can also be performed before identifying each start time and end time.

[0065] In an embodiment of the present application, after obtaining the intermediate state file corresponding to each initial video material, under the guidance of the MLT template, a target video file corresponding to each initial video material can be generated according to each initial video material, the intermediate state subtitle file corresponding to each initial video material, the dynamic special effect image of each target keyword corresponding to each initial video material and its playback time interval, the special effect audio and its playback time interval. The target video file is used to render the target video material with linked subtitles, dynamic special effect images and special effect audio.

[0066] MLT templates refer to the MLT framework's ability to save video editing parameters and settings as template files (usually in XML format). These template files contain all editing parameters, including video cut points, filter settings, transition effects, and audio processing. Exporting parameters allows users to export their currently edited video project as an MLT template file. This allows users to open the template file on other computers using the same video editing software and continue editing or batch generate videos. This export method is particularly suitable for scenarios where batch video generation is required across multiple computers. For example, within a content creation team, one editor can export the edited template, and other members can use it to generate multiple similar videos. Batch generation means that using MLT templates, users can standardize video editing parameters and then batch generate multiple videos using automated scripts or other tools. This is extremely useful when generating a large number of similar videos (such as advertisements or instructional videos), significantly improving work efficiency. In short, an MLT template is a predefined configuration file that instructs the MLT framework on how to combine video assets, subtitle files, dynamic special effects graphics, and special effects audio to generate the final target video file. By using MLT templates, you can automate the video editing process, improve production efficiency, and ensure the consistency and high quality of video content.

[0067] In an optional embodiment, under the guidance of the MLT template, a target video file corresponding to each initial video material is generated according to each initial video material, the intermediate subtitle file corresponding to each initial video material, the dynamic special effect image of each keyword corresponding to each initial video material and its playback time interval, and the audio special effect and its playback time interval, including: obtaining the first storage path information of each initial video material, the second path information of the subtitle file corresponding to each initial video material, the third storage path information of the dynamic special effect image and the fourth position information of the audio special effect; under the guidance of the MLT template rule, a target video file corresponding to each video material is generated according to the first storage path information of each initial video material and its playback time interval, the second path information of the subtitle file corresponding to each initial video material, the intermediate subtitle file corresponding to each initial video material, the third storage path information of the dynamic special effect image and its playback time interval, the fourth position information of the audio special effect and its playback time interval, and the configuration information of each track.

[0068] Furthermore, a video rendering engine can be used to render the target video file corresponding to each video source to obtain the corresponding video. When batch rendering multiple videos, multiple rendering engines can be called to render multiple target video files in parallel, thereby improving video rendering efficiency. If the user is dissatisfied with the rendered video effect, for example, the target keyword position does not meet the user's requirements, this can be resolved by fine-tuning the MLT template parameters.

[0069] The technical solutions provided by the above-mentioned embodiments of this application automatically realize the automatic generation of subtitle files, the precise extraction of target keywords, and the special effects generation logic driven by these keywords by using the network architecture of the AI video editing model. This process involves creating dynamic special effects images and special effects audio related to each target keyword, and setting the playback time interval of the dynamic special effects images and special effects audio with the purpose of playing these dynamic special effects synchronously with the target keywords. In this way, the automation of the video editing process can be achieved, which not only meets the user's needs for fast video editing and batch processing, improves the editing efficiency of the video, but also realizes the precise extraction of keywords, enhances the synergy between the video screen content, subtitles and their keyword special effects, and improves the overall appeal of the video and the user experience.

[0070] Figure 3 This is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Figure 3 As shown, it includes: a memory 30a and a processor 30b; the memory 30a is used to store computer programs; the processor 30b is coupled with the memory 30a and is used to execute the computer programs to perform the steps in the above method embodiment.

[0071] The detailed implementation and beneficial effects of each module in the embodiments of the present application have been described in detail in the aforementioned embodiments and will not be elaborated here.

[0072] Further, if Figure 3 As shown, the electronic device also includes: a communication component 30c, a display 30d, a power component 30e, an audio component 30f and other components. Figure 3 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 3 In addition, Figure 3 The components in the dotted box are optional components, not mandatory components, and the specific components may depend on the product form of the electronic device. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT device, or a server device such as a conventional server, a cloud server or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it may include Figure 3 If the electronic device of this embodiment is implemented as a conventional server, cloud server or server array and other server-side devices, it may not include Figure 3 Components within the dotted box.

[0073] The above-mentioned memory can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0074] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 6G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra wide band (UWB) technology, Bluetooth (BT) technology and other technologies.

[0075] The above-mentioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundary of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0076] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.

[0077] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as call mode, recording mode, and voice recognition mode, the microphone is configured to receive external audio signals. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0078] Accordingly, an exemplary embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the above method.

[0079] Accordingly, exemplary embodiments of the present application further provide a computer program product, which includes a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to perform the steps in the above method.

[0080] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) that contain computer-usable program code.

[0081] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0082] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0084] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.

[0085] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0086] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0087] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0088] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A video editing method, characterized in that: include: In response to a video editing request, obtaining at least one initial video material, each initial video material having a respective subject type; Input each initial video material into the subtitle generation network in the AI video editing model, identify the video content and / or audio clips of each initial video material, and generate a subtitle file for the initial video material based on the video content and / or audio clips, where each subtitle file includes: multiple subtitle lines, and the start and end time of each subtitle line display; Input the theme type and subtitle file of each initial video material into the keyword extraction network in the AI video editing model, perform a keyword extraction operation on the subtitle file of each initial video material, obtain the target keyword adapted to the theme type and the start and end time of each target keyword displayed in the corresponding subtitle line text, and each target keyword corresponds to its own video screen; The target keywords in the subtitle file of each initial video material, the video pictures corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text are input into the special effects generation network in the AI video editing model, and the special effects generation logic driven by the target keywords is executed. Based on the video pictures corresponding to each target keyword, the dynamic special effects parameters and audio special effects parameters corresponding to each target keyword are generated. According to the dynamic special effects parameters and audio special effects parameters, the dynamic special effects image and special effects audio associated with the target keyword are respectively generated. Based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, the playback time interval of the dynamic special effects image and special effects audio is configured with the goal of playing the dynamic special effects image and special effects audio in conjunction with the target keyword.

2. The method according to claim 1, characterized in that The video clip is a video clip with picture content, and the audio clip is a non-silent audio clip. Each initial video material is input into the subtitle generation network in the AI video editing model, the picture content and / or audio clip of each initial video material is identified, and a subtitle file of the initial video material is generated based on the picture content and / or audio clip, including: Input each initial video material into the subtitle generation network in the AI video editing model, identify the audio segments of each initial video material and the video segments corresponding to each audio segment, and obtain multiple audio segments and the video segments corresponding to each audio segment. Each audio segment and its corresponding video segment have their own playback start time and end time; Optimizing the initial subtitle text of each audio segment according to the screen content of the video segment corresponding to the audio segment to obtain a target subtitle text; Based on a subtitle line segmentation rule, segment the target subtitle text into at least one subtitle line text that is synchronously adapted to the audio content of the audio segment and / or the picture content of the corresponding video segment; A subtitle file corresponding to the initial video material is generated based on each subtitle line text and its subtitle line sequence number as well as the start time and end time of display of each subtitle line text.

3. The method according to claim 1, characterized in that Input the theme type and subtitle file of each initial video material into the keyword extraction network in the AI video editing model, perform keyword extraction on the subtitle file of each initial video material, obtain target keywords adapted to the theme type, and the start and end time of each target keyword displayed in the corresponding subtitle line text, including: Inputting the subject type of each initial video material and the subtitle file into the keyword extraction network in the AI video editing model, and determining candidate keywords for each subtitle line text in each subtitle file based on the key semantic information of each subtitle line text in each subtitle file; Selecting keywords that match the subject type of the corresponding initial video material from the candidate keywords of each subtitle line text as target keywords for each subtitle line text in each subtitle file; as well as The start time and end time of display of each target keyword in the corresponding subtitle line text are determined according to the position information of each target keyword in the corresponding subtitle line text and the start time and end time of display of the subtitle line text.

4. The method according to any one of claims 1 to 3, characterized in that Input each target keyword in the subtitle file of each initial video material, the video image corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text into the special effect generation network in the AI video editing model, execute the special effect generation logic driven by the target keyword, generate dynamic special effect parameters and audio special effect parameters corresponding to each target keyword based on the video image corresponding to each target keyword, and generate dynamic special effect images and special effect audio associated with the target keyword according to the dynamic special effect parameters and audio special effect parameters, including: Input each target keyword in the subtitle file of each initial video material, the video screen corresponding to each target keyword, and the start time and end time of each target keyword displayed in the corresponding subtitle line text into the special effects generation network in the AI video editing model, and determine the dynamic special effects template and audio special effects template of the target keyword based on the emotional attribute information of each target keyword and the picture content attribute information of its corresponding video screen; Generate dynamic special effect parameters and audio special effect parameters corresponding to the target keyword according to the dynamic special effect template and audio special effect template corresponding to the target keyword, wherein the dynamic special effect parameters include multi-frame special effect parameters and the audio special effect parameters include multi-note special effect parameters; Inputting multi-frame special effect parameters corresponding to the target keyword into an image drawing tool, generating a multi-frame special effect image associated with the target keyword under the guidance of the multi-frame special effect parameters, and generating a dynamic special effect image associated with the target keyword based on the multi-frame special effect image using a dynamic special effect graph processing tool; as well as The multi-note special effect parameters corresponding to the target keyword are input into the audio special effect processing tool, and under the guidance of the multi-note special effect parameters, the special effect audio associated with the target keyword is generated.

5. The method according to claim 4, characterized in that According to the emotional attribute information of each target keyword and the picture content attribute information of its corresponding video picture, a dynamic special effect template and an audio special effect template for the target keyword are determined, including: Determining the screen content of the video screen corresponding to each target keyword according to the start time and end time of each target keyword display; Performing attribute analysis on each target keyword and the picture content of the corresponding video picture to obtain emotional attribute information of each target keyword and picture content attribute information of the video picture corresponding to the target keyword; Determining dynamic special effect attribute information and audio special effect attribute information adapted to the target keyword based on the emotional attribute information of each target keyword and the picture content attribute information of the video picture corresponding to the target keyword; According to the dynamic special effect attribute information and the audio special effect attribute information of the target keyword, the dynamic special effect template and the audio special effect template of the target keyword are determined.

6. The method according to claim 4, characterized in that Generate dynamic special effect parameters and audio special effect parameters corresponding to the target keyword according to the dynamic special effect template and audio special effect template corresponding to the target keyword, including: According to the dynamic special effect template corresponding to the target keyword, a frame-by-frame special effect parameter parsing operation is performed to generate multi-frame dynamic special effect parameters; as well as According to the audio special effect template corresponding to the target keyword, a note-by-note parameter parsing operation is performed to generate multi-note audio special effect parameters.

7. The method according to claim 4, characterized in that Each initial video material is equipped with a video image playback track, a subtitle line display track, a dynamic special effect playback track, an audio special effect playback track, and a time track. These tracks are interconnected and the start and end points of each track are aligned with each other, so that as time on the time track gradually advances, the corresponding subtitle line text or special effect is activated on the corresponding track; based on the start time and end time displayed in the corresponding subtitle line text of the target keyword, with the goal of playing the dynamic special effect image and special effect audio in conjunction with the target keyword, the playback time interval of the dynamic special effect image and special effect audio is configured, including: The display time interval formed by the start time and end time of the target keyword display on the time track, and the playback time interval occupied by the video picture adapted to the dynamic special effect image and / or special effect audio on the time track, are configured as the playback time interval of the dynamic special effect image and / or special effect audio of the target keyword, and the playback time interval includes the playback start time and end time; Based on the playback time interval of the dynamic special effects image and / or special effects audio of the target keyword on the time track, in the vertical dimension, the dynamic special effects image and / or special effects audio of the target keyword are respectively configured in the corresponding position interval of the dynamic special effects track and / or the audio special effects track as the activation position interval of the dynamic special effects image and / or the special effects audio.

8. The method according to claim 7, characterized in that The method further comprises: Based on the time information provided by the time track equipped with the initial video material and the start time and end time of the video frame of the initial video material, determining the activation position interval of the video frame on the video frame playback track, so as to activate and start playing the video frame when the playback start time of the video frame represented by the current time information on the time track is reached, and stop playing the video frame when the end time of the live video frame is reached; Based on the start time and end time of display of each subtitle line text, determining the activation position interval of the subtitle line text on the subtitle line track, so that the subtitle line text is activated and starts to be displayed when the playback start time of the subtitle line text represented by the current time information on the time track is reached, and the subtitle line text is stopped from being displayed when the end time of the subtitle line text is reached; Based on the start time and end time of the playback of the dynamic special effects image and / or special effects audio of each target keyword, the activation position interval of the dynamic special effects image and / or special effects audio of the target keyword in the dynamic special effects track and / or audio special effects track is determined respectively, so that the dynamic special effects image is activated and played when the playback start time of the dynamic special effects image represented by the current time information on the time track is reached, and the dynamic special effects image is stopped when the end time of the dynamic special effects image is reached; and / or, the special effects audio is activated and played when the playback start time of the special effects audio represented by the current time information on the time track is reached, and the special effects audio is stopped when the end time of the special effects audio is reached.

9. The method according to claim 7 or 8, characterized in that The AI video editing model also includes a rich text generation network, and the method further includes: Inputting the subtitle file of each initial video material into the rich text generation network, converting the subtitle file of each initial video material into a rich text format, and obtaining an intermediate subtitle file corresponding to each initial video material; Under the guidance of the MLT template, a target video file corresponding to each initial video material is generated based on each initial video material, the intermediate subtitle file corresponding to each initial video material, the dynamic special effect image and its playback time interval of each target keyword corresponding to each initial video material, and the special effect audio and its playback time interval. The target video file is used to render the target video material with linked subtitles, dynamic special effect images and special effect audio.

10. The method according to claim 9, characterized in that Under the guidance of the MLT template rules, based on each initial video material, the intermediate subtitle file corresponding to each initial video material, the dynamic special effect image of each keyword corresponding to each initial video material and its playback time interval, the audio special effect and its playback time interval, the target video file corresponding to each initial video material is generated, including: Obtaining first storage path information of each initial video material, second path information of a subtitle file corresponding to each initial video material, third storage path information of a dynamic special effect image, and fourth location information of an audio special effect; Under the guidance of the MLT template rules, the target video file corresponding to each video material is generated based on the first storage path information and playback time interval of each initial video material, the second path information of the subtitle file corresponding to each initial video material, the intermediate subtitle file corresponding to each initial video material, the third storage path information and playback time interval of the dynamic special effects image, the fourth position information of the audio special effects and playback time interval, and the configuration information of each track.

11. An electronic device, characterized in that: include: memory and processor; The memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in any one of the methods of claims 1-10.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to implement the steps in any one of the methods of claims 1 to 10.

13. A computer program product, characterized in that The computer program product comprises a computer program / instruction, which, when executed by a processor, enables the processor to implement the steps of any one of the methods of claims 1 to 10.

Citation Information

Patent Citations

  • Video processing method and device, computer equipment and storage medium

    CN114095782A

  • Video recording method and electronic equipment

    CN114390341A

  • Video editing method and device, computer equipment and storage medium

    CN114449310A

  • Video processing method, video processing device, electronic equipment and storage medium

    CN115883919A

  • Information display method and device, electronic equipment and storage medium

    CN115952319A