Video editing method, device, equipment and storage medium

By generating subtitles and audio text card timelines, the problems of low efficiency and lack of accuracy in traditional video editing are solved, and fast and accurate video editing effects are achieved.

CN119450156BActive Publication Date: 2025-09-23CHINA OPEN UNIV PRESS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411605242.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-09-23
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The traditional video editing process is inefficient and prone to missing important information, and the editing accuracy is insufficient.

Method used

By extracting subtitles from video keyframes, generating subtitle text cards and constructing a subtitle text card timeline, the video is edited based on this axis, and multi-dimensional text card fusion is performed with audio text cards to improve editing accuracy and efficiency.

Benefits of technology

It can quickly and accurately locate the position of specific knowledge points, improving the accuracy and efficiency of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119450156B_ABST
    Figure CN119450156B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video editing method, apparatus, device, and storage medium. The method comprises: obtaining a video to be edited; extracting key frames from the video, extracting and analyzing the subtitles in each key frame, and generating a subtitle text card for each key frame; generating a subtitle text card timeline for the video based on the subtitle text cards for each key frame; and editing the video based on the subtitle text card timeline. In the present disclosure, the subtitle text cards can be used to quickly and accurately locate specific knowledge points, and the clips to be edited containing the specific knowledge points can be edited, thereby improving the accuracy and efficiency of video editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video editing method, apparatus, device, and storage medium. Background Art

[0002] The traditional video editing process requires the editor to watch the original video content multiple times and replay it repeatedly to determine the appropriate splitting time. After determining the splitting time, the editor must repeatedly drag and drop on the traditional timeline to select the appropriate video in and out points to ensure the integrity and accuracy of the edited content. This traditional video editing process is not only inefficient but also prone to missing important information. Therefore, how to improve the accuracy and efficiency of video editing is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0003] In view of this, the present disclosure proposes a video editing method, apparatus, device and storage medium, which can improve the accuracy and efficiency of video editing.

[0004] According to a first aspect of the present disclosure, a video editing method is provided, comprising:

[0005] Get the video to be edited;

[0006] Extracting key frames from the video, extracting and analyzing the subtitles in each key frame, and generating subtitle text cards for each key frame;

[0007] Generating a subtitle text card timeline of the video based on the subtitle text cards of each key frame;

[0008] The video is edited based on the subtitle text card timeline.

[0009] In a possible implementation, when extracting and analyzing the subtitles in the key frame to generate the subtitle text card of the key frame, the process includes:

[0010] extracting subtitles from the key frames;

[0011] Sentence processing is performed on the subtitles to obtain all sentences constituting the subtitles;

[0012] Traversing all sentences constituting the subtitles, determining, for a current sentence traversed, an appearance time of the current sentence in the video, and generating a subtitle text card corresponding to the current sentence based on the current sentence and the appearance time of the current sentence in the video;

[0013] After the traversal is completed, the subtitle text cards corresponding to all the sentences constituting the subtitle are used as the subtitle text cards of the key frame.

[0014] In a possible implementation, when determining the appearance time of the traversed current sentence in the video, the method further includes:

[0015] Determining a display position of the current sentence in the key frame;

[0016] After generating the subtitle text card corresponding to the current sentence, the method further includes:

[0017] The display position is associated with the subtitle text card corresponding to the current sentence.

[0018] In a possible implementation, when generating a subtitle text card timeline of the video based on the subtitle text cards of each key frame, the method includes:

[0019] Arranging the subtitle text cards of the key frames in the order of their appearance time in the video;

[0020] The arranged subtitle text cards are sequentially associated with the corresponding time on the video timeline to obtain the subtitle text card timeline of the video.

[0021] In a possible implementation, when editing the video based on the subtitle text card timeline, the process includes:

[0022] Obtaining a screening condition of the clip to be edited, and filtering out a subtitle text card that meets the screening condition from the subtitle text card timeline as a target subtitle text card, and displaying the target subtitle text card on the subtitle text card timeline;

[0023] Based on each target subtitle text card on the subtitle text card timeline, setting the in point position and the out point position of the to-be-cut segment;

[0024] Based on the in point position and the out point position, the segment to be edited is cut out from the video.

[0025] In a possible implementation, the screening condition includes at least one of a keyword, a preset word, and a display position.

[0026] In a possible implementation, the method further includes:

[0027] Extracting an audio file from the video;

[0028] Based on the audio file, generate an audio text card for the video;

[0029] Based on the audio and text cards, generating an audio and text card timeline of the video;

[0030] The audio text card timeline and the subtitle text card timeline are merged to obtain a multi-dimensional text card timeline.

[0031] According to a second aspect of the present disclosure, there is provided a video editing device, comprising:

[0032] A video acquisition module is used to acquire the video to be edited;

[0033] A subtitle text card generation module is used to extract key frames from the video, extract and analyze the subtitles in each key frame, and generate a subtitle text card for each key frame;

[0034] A subtitle text card timeline generation module, configured to generate a subtitle text card timeline for the video based on the subtitle text cards of each key frame;

[0035] The video editing module is used to edit the video based on the subtitle text card timeline.

[0036] According to a third aspect of the present disclosure, a video editing device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in the first aspect of the present disclosure.

[0037] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions, when executed by a processor, implement the method described in the first aspect of the present disclosure.

[0038] The present disclosure provides a video editing method, apparatus, device, and storage medium. The method comprises: obtaining a video to be edited; extracting key frames from the video, extracting and analyzing the subtitles in each key frame, and generating a subtitle text card for each key frame; generating a subtitle text card timeline for the video based on the subtitle text cards for each key frame; and editing the video based on the subtitle text card timeline. In the present disclosure, the subtitle text cards can be used to quickly and accurately locate specific knowledge points, and the clips to be edited containing the specific knowledge points can be edited, thereby improving the accuracy and efficiency of video editing.

[0039] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0041] Figure 1 A flowchart of a video editing method according to an embodiment of the present disclosure is shown;

[0042] Figure 2 A diagram showing a video editing interface according to an embodiment of the present disclosure is shown;

[0043] Figure 3 A diagram showing a video editing interface according to another embodiment of the present disclosure is shown;

[0044] Figure 4 A schematic block diagram of a video editing device according to an embodiment of the present disclosure is shown;

[0045] Figure 5 A schematic block diagram of a video editing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0046] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0047] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0048] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0049] <Method Example>

[0050] Figure 1 FIG. 1 is a flow chart showing a video editing method according to an embodiment of the present disclosure. Figure 1 As shown, the method includes steps S1100-S1400.

[0051] S1100: Obtain the video to be edited. The video to be edited is the video currently being edited. For ease of description, the video to be edited will be referred to as the video below. When editing a video, the video must first be uploaded to an editing system that implements the disclosed video editing method. This allows the editing system to retrieve the video and proceed to step S1200.

[0052] S1200 extracts key frames from the video and extracts and analyzes the subtitles within each key frame to generate a subtitle text card for each key frame. A key frame is a frame that contains a key action in the motion of a character or object, equivalent to the original image in a two-dimensional animation. The subtitles within a key frame refer to all text within the key frame image. This subtitle may include at least one of the presentation text presented within the key frame and lyrics used to explain the content of the key frame.

[0053] For each extracted key frame, the subtitles in the key frame need to be extracted and analyzed to generate a subtitle text card for each key frame. Since the generation process of the subtitle text card for each key frame is the same, the following key frame is used as an example to illustrate the generation process of the subtitle text card for the key frame. Specifically, when extracting and analyzing the subtitles in the key frame and generating the subtitle text card for the key frame, the following steps can be included:

[0054] First, subtitles are extracted from the key frames. Specifically, when extracting subtitles from the key frames, it can be implemented based on an OCR recognition algorithm.

[0055] Second, the subtitles are processed into sentences to obtain all the sentences that make up the subtitles. Specifically, the recognized subtitles can be input into a natural language processing model to perform sentence processing on the subtitles through the natural language processing model to obtain all the sentences that make up the subtitles.

[0056] Third, traverse all the sentences that make up the subtitles, and for the current sentence that has been traversed, determine the appearance time of the current sentence in the video, and generate the subtitle text card corresponding to the current sentence based on the current sentence and the appearance time of the current sentence in the video. Specifically, when traversing to the current sentence, you can first determine the appearance time of the frame where the current sentence is located in the video based on the frame rate of the video. Then, fill the current sentence and the appearance time of the current sentence in the video into the preset text card rendering template, and then generate the subtitle text card corresponding to the current sentence based on the filled text card rendering template. Among them, the subtitle text card corresponding to the current sentence includes the current sentence and the appearance time of the current sentence in the video (the appearance time is composed of the start time and the end time). Through the subtitle text card corresponding to the current sentence, the user can quickly understand the appearance time of the knowledge content involved in the current sentence in the video clip.

[0057] In one possible implementation, when the subtitles in the same keyframe are divided into multiple sentences, the time at which the keyframe appears in the video can be used as the time at which each sentence in the keyframe appears in the video. In another possible implementation, all sentences appearing in the same keyframe can be merged into a subtitle text card for display, and the time at which the keyframe appears in the video can be used as the time at which the merged subtitle text card appears in the video.

[0058] Fourth, at the end of the traversal, a subtitle file card for each sentence is obtained, and the subtitle text cards for all the sentences obtained are used as the subtitle text cards for the keyframe. In other words, the keyframe subtitles contain several sentences, and the keyframe has several letter text cards. For example, when the subtitles of the keyframe are divided into sentences, four sentences are obtained. Four subtitle text cards will be generated for each of the four divided sentences. These four subtitle text cards constitute the subtitle text card for the keyframe.

[0059] In a possible implementation, when determining the appearance time of the current sentence in the video for the current sentence that has been traversed, it also includes: determining the display position of the current sentence in the key frame. Specifically, when extracting subtitles from the key frame, the display position of each character in the subtitle in the key frame will also be identified and recorded at the same time. In this way, when traversing the current sentence, the display position of each character that makes up the current sentence in the key frame can be obtained, and then the display position of the current sentence in the key frame can be determined based on the display position of each character in the key frame. In this implementable method, after generating the subtitle text card corresponding to the current sentence, the display position of the current sentence is also associated with the subtitle text card corresponding to the current sentence. This is to facilitate the subsequent screening of subtitle text cards according to the display position set by the user.

[0060] By processing each key frame in the video in accordance with the above method, the subtitle text card of each key frame can be obtained.

[0061] After obtaining the subtitle text cards of each key frame, step S1300 may be executed to generate a subtitle text card timeline of the video based on the subtitle text cards of each key frame.

[0062] In a possible implementation, when generating a subtitle text card timeline for a video based on the subtitle text cards of each key frame, the following steps may be included:

[0063] First, arrange the subtitle text cards for each keyframe in the order of their appearance in the video. Specifically, each subtitle text card includes the recorded time of the sentence's appearance in the video. The recorded time of the sentence's appearance in the video is the appearance time of the subtitle text card in the video. After determining the appearance time of each subtitle text card in the video, they can be arranged in the order of their appearance time.

[0064] Second, associate each arranged subtitle text card with the corresponding time on the video timeline in turn to obtain the subtitle text card timeline of the video. Specifically, first generate the video timeline, and then associate each subtitle text card to the corresponding position of the video timeline according to the time when each subtitle text card appears in the video, thereby obtaining a subtitle text card timeline composed of the video timeline and the subtitle text cards bound to the video timeline. After obtaining the subtitle text card timeline, slide on the video timeline, locate the time point of the video according to the sliding position, and locate the subtitle text card that appears at the time point according to the located video time point, so that the subtitles that appear at the time point in the video can be understood through the subtitle text cards that appear at the time point.

[0065] In a possible implementation, the interval associated with the subtitle text card on the video timeline can also be highlighted to facilitate users to quickly locate the valid area associated with the subtitle text card on the video timeline.

[0066] It should be noted here that the video timeline generated in this step is a new timeline independent of the original video timeline. After associating each subtitle text card to the newly generated video timeline, a new subtitle text card timeline independent of the original video timeline will be generated. That is to say, after executing this step, two timelines will be presented in the editing interface, one is the original video timeline, and the other is the subtitle text card timeline including multiple subtitle text cards. In a specific video editing example, the generated subtitle text card timeline can be seen in Figure 2 3. Subtitle Timeline. This way, when you subsequently edit a video, after determining the in and out points of the clip to be edited using the subtitle text cards on the subtitle timeline, the corresponding in and out points can be displayed on the original video timeline, making it easier for users to view the location information of the clip to be edited in the video.

[0067] After the subtitle text card timeline is generated, step S1400 can be executed to edit the video based on the subtitle text card timeline. Specifically, the following steps may be included when editing the video:

[0068] First, a filter condition for the clip to be edited is obtained, and subtitle text cards that meet the filter condition are selected from the subtitle text card timeline as target subtitle text cards. The target subtitle text cards are then displayed on the subtitle text card timeline to increase the visibility of valid subtitle text cards on the subtitle text card timeline. The filter condition may include at least one of a keyword, a preset word, a display position, and hidden lyrics.

[0069] When filtering based on keywords, users can Figure 3 Enter a keyword in the keyword filtering area shown in , so that when filtering the target subtitle text cards, the subtitle text cards including the keyword will be filtered out from all subtitle text cards as the target subtitle text cards, and the filtered target subtitle text cards will be displayed in sequence on the subtitle text card timeline, and other subtitle text cards that do not meet the filtering conditions will be hidden.

[0070] When filtering based on preset words, it is necessary to first obtain the preset word information configured by the user. When obtaining the preset word information configured by the user, it includes: responding to the triggering of the preset word setting control, pushing and displaying the preset word configuration interface, and obtaining the preset word information configured by the user through the preset word configuration interface. The preset word information may include a preset word and a preset word type. The preset word may be a text specified by the user (such as a conclusion) or a symbol (such as "()"). The type of the preset word is used to indicate the position of the preset word in the sentence. For example, the type of the preset word may include at least one of starting with the preset word and ending with the preset word.

[0071] In a possible implementation, the preset word configuration interface displayed by the push is as follows: Figure 3 As shown in the window interface in the middle, it includes a preset word information display area on the upper side and a preset word configuration area on the lower side. The user can enter the preset word to be searched in the preset word configuration area on the lower side and set the type of the preset word. After completing the input of the preset word and configuring the preset word type, click the Add control to add the newly configured preset word information to the preset word information display area. After completing the configuration of the preset word, close the preset word configuration interface, and the editing system can obtain all the preset word information configured by the user through the preset word configuration interface.

[0072] After obtaining the preset word information configured by the user, the preset word will be extracted from the preset word information and pushed to the user's Figure 3The preset word filtering area is shown. In this way, when the user performs preset word filtering, he can select the desired preset word in the preset word filtering area. When the preset word is selected, the preset word and the type of the preset word will be obtained, and then the subtitle text cards that include the selected preset word and whose position in the sentence is consistent with the type of the preset word will be filtered out from all the subtitle text cards as the target subtitle text cards, and the filtered target subtitle text cards will be displayed in order on the subtitle text card timeline, and other subtitle text cards that do not meet the filtering conditions will be hidden.

[0073] In addition, when the user needs to adjust the preset words, he can trigger the preset word setting control again to call out the preset word configuration interface, and add, delete and modify the preset words based on the preset word configuration interface.

[0074] When filtering based on display location, users can Figure 3 Enter the specified display position in the coordinate filtering area shown in , so that when filtering the target subtitle text cards, the subtitle text cards whose display position of the bound sentence is the same as the display position specified by the user will be filtered out from all the subtitle text cards as the target subtitle text cards. It should be noted here that the display position of the key knowledge points in the video frame is different in different videos. Therefore, the user can filter out the target subtitle text cards including the key knowledge points by setting the display position of the key knowledge points, and display the filtered target subtitle text cards in sequence on the subtitle text card timeline. Other subtitle text cards that do not meet the filtering conditions will be hidden, thereby improving the accuracy of the target subtitle text card filtering.

[0075] When filtering based on hidden lyrics, users can Figure 3 In the hidden lyrics area shown, it is set whether to hide the lyrics. If the user chooses to hide the lyrics, the pre-configured lyrics display position will be obtained, and all subtitle text cards with the same display position of the bound sentences as the lyrics display position will be hidden, and other non-hidden subtitle text cards will be displayed as target subtitle text cards. It should be noted here that the sentences contained in the lyrics text cards are the content of the keyframe that explains the demonstration courseware (i.e., the subtitles of the audio commentary). When the courseware content displayed in the keyframe is used as a knowledge point for video editing, you can choose to hide the lyrics, and then hide the subtitle text cards including the lyrics, so that only the subtitle text cards related to the displayed courseware are displayed, so that users can easily perform video editing based on the knowledge points in the displayed courseware content.

[0076] After filtering out the target subtitle text cards, you can proceed to the second step.

[0077] Second, based on each target subtitle text card on the subtitle text card timeline, set the in point position and out point position of the clip to be edited. Figure 3 As shown, the filtered target subtitle text cards are displayed on the subtitle text card timeline in the order of their appearance. In addition to the sentence and the appearance time of the sentence, each subtitle text card also includes an entry point position setting control (i.e., a circular control including a left arrow) and an exit point position setting control (i.e., a circular control including a right arrow). The user can determine the start subtitle text card and the end subtitle text card corresponding to the segment to be edited from the target subtitle text card, and then click the entry point position setting control on the start subtitle text card. When the entry point position control is clicked, the start time of the appearance time will be extracted from the start subtitle text card as the entry point position of the segment to be edited. Click the exit point position setting control on the end subtitle text card. When the exit point position setting control is clicked, the end time of the appearance time will be extracted from the end subtitle text card as the exit point position of the segment to be edited.

[0078] In one possible implementation, in order to clearly display the entry point and exit point of the clip to be edited, the corresponding entry point and exit point will be marked on the original video timeline, and the area between the entry point and exit point will be highlighted, so that the user can clearly understand the position of the clip to be edited in the video.

[0079] Third, based on the in point position and the out point position, the clip to be edited is cut out from the video. Specifically, how to cut out the clip to be edited from the video based on the in point position and the out point position is common knowledge in the art and will not be described in detail here.

[0080] In one possible implementation, after selecting target subtitle text cards, subtitle text cards containing useless sentences can be deleted from the target subtitle text cards. After deleting the subtitle text cards, the corresponding time segments in the video are deleted based on the appearance time of the deleted subtitle text cards. In this way, after identifying the entry point and exit point of the segment to be edited based on the remaining target text cards, the video segment between the entry point and the exit point can be cut out from the edited video as the segment to be edited, thereby improving the simplicity of the segment to be edited.

[0081] In one possible implementation, after filtering out the target subtitle text cards, you can also set the order of appearance of the target subtitle text cards. When the order of appearance of the target subtitle text cards is set, corresponding video clips will be edited from the video according to the appearance time of each target subtitle text card, and then the edited video clips will be combined in sequence according to the order of appearance of each target subtitle text card, so as to obtain an edited video that conforms to the order of appearance of each target subtitle text card.

[0082] In a possible implementation, while generating the subtitle text card timeline of the video, the following steps are also included:

[0083] First, extract the audio file from the video.

[0084] Second, based on the audio file, audio and text cards for the video are generated. Specifically, the ASR algorithm is first used to recognize the audio text. The audio text is then input into a natural language processing model to segment the audio text into sentences, obtaining the individual sentences that make up the audio text. A corresponding audio and text card is then generated for each sentence that makes up the audio text. All the resulting audio and text cards constitute the audio and text cards for the video.

[0085] Third, based on the audio text card, generate the audio text card timeline of the video. The specific generation method refers to the generation process of the subtitle text card timeline and will not be repeated here.

[0086] Fourth, merge the audio text card timeline with the subtitle text card timeline to obtain a multi-dimensional text card timeline including subtitle text cards and audio text cards. Specifically, merge the timeline of the audio text card and the timeline of the subtitle text card into one timeline, associate the subtitle text cards and audio text cards that appear at the same time point through one timeline, and highlight the positions on the timeline associated with the subtitle text cards and audio text cards. When sliding on the multi-dimensional text card timeline, the subtitle text cards and audio text cards associated with the video time point will be linked and displayed according to the time point in the video at the sliding position and the located video time point. Based on the associated subtitle text cards and audio text cards, the subtitles and voices that appear in the video at the located time point can be clearly understood, which is convenient for users to screen subtitle and voice knowledge points at the same time. In a specific example, the multi-dimensional text card timeline can be as follows Figure 3 shown.

[0087] After generating the multi-dimensional text card timeline, users can choose to edit the video based on the subtitle text card or the audio text card, further improving the flexibility of video editing. The method of editing the video based on the audio text card is similar to the process of editing the video based on the subtitle text card, which will not be repeated here.

[0088] The present disclosure provides a video editing method, comprising: obtaining a video to be edited; extracting key frames from the video, extracting and analyzing the subtitles in each key frame, and generating a subtitle text card for each key frame; generating a subtitle text card timeline for the video based on the subtitle text cards for each key frame; and editing the video based on the subtitle text card timeline. In the present disclosure, the subtitle text cards can be used to quickly and accurately locate specific knowledge points, and the clips to be edited containing the specific knowledge points can be edited, thereby improving the accuracy and efficiency of video editing.

[0089] <Device Example>

[0090] Figure 4 FIG. 1 is a schematic block diagram of a video editing device according to an embodiment of the present disclosure. Figure 4 As shown, the device 100 includes:

[0091] The video acquisition module 110 is used to acquire the video to be edited;

[0092] The subtitle text card generation module 120 is used to extract key frames from the video, extract and analyze the subtitles in each key frame, and generate a subtitle text card for each key frame;

[0093] A subtitle text card timeline generation module 130 is used to generate a subtitle text card timeline for a video based on the subtitle text cards of each key frame;

[0094] The video editing module 140 is used to edit the video based on the subtitle text card timeline.

[0095] <Equipment Example>

[0096] Figure 5 FIG. 1 shows a schematic block diagram of a video editing device according to an embodiment of the present disclosure. Figure 5 As shown, the video editing device 200 includes: a processor 210 and a memory 220 for storing executable instructions of the processor 210. The processor 210 is configured to implement any of the above-mentioned video editing methods when executing the executable instructions.

[0097] It should be noted that there may be one or more processors 210. Furthermore, the video editing device 200 of the present embodiment may also include an input device 230 and an output device 240. The processor 210, memory 220, input device 230, and output device 240 may be connected via a bus or other means, which are not specifically limited herein.

[0098] Memory 220, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and various modules, such as the program or module corresponding to the video editing method of the present disclosure. Processor 210 executes the software programs or modules stored in memory 220 to perform various functional applications and data processing of video editing device 200.

[0099] The input device 230 may be used to receive input numbers or signals. The signals may be key signals related to user settings and function control of the device / terminal / server. The output device 240 may include a display device such as a display screen.

[0100] <Storage Medium Embodiment>

[0101] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is further provided, on which computer program instructions are stored. When the computer program instructions are executed by the processor 210, any of the above-mentioned video editing methods is implemented.

[0102] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A video editing method, characterized in that: include: Get the video to be edited; Extracting key frames from the video, extracting and analyzing the subtitles in each key frame, and generating subtitle text cards for each key frame; Generating a subtitle text card timeline of the video based on the subtitle text cards of each key frame; Editing the video based on the subtitle text card timeline; When generating a subtitle text card timeline of the video based on the subtitle text cards of each key frame, the method includes: Arranging the subtitle text cards of the key frames in the order of their appearance time in the video; Associating the arranged subtitle text cards with corresponding times on the video timeline in sequence to obtain the subtitle text card timeline of the video, wherein the video timeline is a new video timeline independent of the original video timeline; When editing the video based on the subtitle text card timeline, the method includes: Obtaining a screening condition of the clip to be edited, and filtering out a subtitle text card that meets the screening condition from the subtitle text card timeline as a target subtitle text card, and displaying the target subtitle text card on the subtitle text card timeline; Based on each target subtitle text card on the subtitle text card timeline, setting the in point position and the out point position of the to-be-cut segment; Cutting the to-be-edited segment from the video based on the in-point position and the out-point position; After selecting the target subtitle text cards, the method further includes: deleting the subtitle text cards including useless sentences in the target subtitle text cards; after deleting the subtitle text cards including the useless sentences, deleting the corresponding time segments in the video based on the appearance time of the deleted subtitle text cards; After deleting subtitle text cards including useless sentences from the target subtitle text cards, setting the appearance order of the target subtitle text cards. When the appearance order of the target subtitle text cards is set, corresponding video clips are clipped from the video according to the appearance time of each target subtitle text card, and then the clipped video clips are sequentially combined according to the appearance order of each target subtitle text card, so as to obtain a clipped video that conforms to the appearance order of each target subtitle text card. When extracting and analyzing the subtitles in the key frame to generate the subtitle text card of the key frame, the method includes: extracting subtitles from the key frames; Sentence processing is performed on the subtitles to obtain all sentences constituting the subtitles; Traversing all sentences constituting the subtitles, determining, for a current sentence traversed, an appearance time of the current sentence in the video, generating a subtitle text card corresponding to the current sentence based on the current sentence and the appearance time of the current sentence in the video, and determining a display position of the current sentence in the key frame, and associating the display position with the subtitle text card corresponding to the current sentence; After the traversal is completed, the subtitle text cards corresponding to all the sentences constituting the subtitle are used as the subtitle text cards of the key frame; The screening condition includes at least one of a keyword, a preset word, and a display position.

2. The method according to claim 1, characterized in that Also includes: Extracting an audio file from the video; Based on the audio file, generate an audio text card for the video; Based on the audio and text cards, generating an audio and text card timeline of the video; The audio text card timeline and the subtitle text card timeline are merged to obtain a multi-dimensional text card timeline.

3. A video editing device, characterized in that: include: A video acquisition module is used to acquire the video to be edited; A subtitle text card generation module is used to extract key frames from the video, extract and analyze the subtitles in each key frame, and generate a subtitle text card for each key frame; A subtitle text card timeline generation module, configured to generate a subtitle text card timeline for the video based on the subtitle text cards of each key frame; A video editing module, configured to edit the video based on the subtitle text card timeline; The subtitle text card timeline generation module is specifically used to generate the subtitle text card timeline of the video based on the subtitle text cards of each key frame: Arranging the subtitle text cards of the key frames in the order of their appearance time in the video; Associating the arranged subtitle text cards with corresponding times on the video timeline in sequence to obtain the subtitle text card timeline of the video, wherein the video timeline is a new video timeline independent of the original video timeline; The video editing module is specifically used to edit the video based on the subtitle text card timeline: Obtaining a screening condition of the clip to be edited, and filtering out a subtitle text card that meets the screening condition from the subtitle text card timeline as a target subtitle text card, and displaying the target subtitle text card on the subtitle text card timeline; Based on each target subtitle text card on the subtitle text card timeline, setting the in point position and the out point position of the to-be-cut segment; Cutting the to-be-edited segment from the video based on the in-point position and the out-point position; After selecting the target subtitle text cards, the method further includes: deleting the subtitle text cards including useless sentences in the target subtitle text cards; after deleting the subtitle text cards including the useless sentences, deleting the corresponding time segments in the video based on the appearance time of the deleted subtitle text cards; After deleting subtitle text cards including useless sentences from the target subtitle text cards, setting the appearance order of the target subtitle text cards. When the appearance order of the target subtitle text cards is set, corresponding video clips are clipped from the video according to the appearance time of each target subtitle text card, and then the clipped video clips are sequentially combined according to the appearance order of each target subtitle text card, so as to obtain a clipped video that conforms to the appearance order of each target subtitle text card. When extracting and analyzing the subtitles in the key frame to generate the subtitle text card of the key frame, the method includes: extracting subtitles from the key frames; Sentence processing is performed on the subtitles to obtain all sentences constituting the subtitles; Traversing all sentences constituting the subtitles, determining, for a current sentence traversed, an appearance time of the current sentence in the video, generating a subtitle text card corresponding to the current sentence based on the current sentence and the appearance time of the current sentence in the video, and determining a display position of the current sentence in the key frame, and associating the display position with the subtitle text card corresponding to the current sentence; After the traversal is completed, the subtitle text cards corresponding to all the sentences constituting the subtitle are used as the subtitle text cards of the key frame; The screening condition includes at least one of a keyword, a preset word, and a display position.

4. A video editing device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 2 when executing the executable instructions.

5. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 2 is implemented.

Citation Information

Patent Citations

  • Video editing method and system and storage medium

    CN110401878A