Video editing method, device, electronic device and storage medium

The method improves video editing efficiency by using voice identification to automatically detect and delete unnecessary clips in selfie videos, addressing the inefficiencies of traditional audio waveform-based methods.

JP7680578B2Active Publication Date: 2025-05-20BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023578854
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-01-19
Filing Date
2023-12-12
Publication Date
2025-05-20
Estimated Expiration
2043-12-12

AI Technical Summary

Technical Problem

Traditional video editing methods for selfie videos are time-consuming and inefficient due to the need to manually identify and delete unnecessary clips such as pauses and idle talk, which are not effectively addressed by current audio waveform-based techniques.

Method used

A method and apparatus for video editing that utilizes voice identification to determine invalid text in a video, marks the corresponding timeline sections, and allows for the adjustment and deletion of these sections, improving editing efficiency.

Benefits of technology

The method enhances video editing efficiency by automatically identifying and removing unnecessary clips, thereby streamlining the editing process and improving the quality of selfie videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680578000001
    Figure 0007680578000001
  • Figure 0007680578000002
    Figure 0007680578000002
  • Figure 0007680578000003
    Figure 0007680578000003
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a storage medium for editing a video. The method includes: determining an invalid text in the voice text of the target video and a timeline position of the invalid text by performing voice identification of an audio in the target video; displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text in the editing track clip according to the timeline position of the invalid text; adjusting the timeline section of the invalid text in the editing track clip of the target video in response to an adjustment operation of the invalid text; and deleting a video clip within the timeline section of the invalid text in the editing track clip from the target video in response to a video clip deletion operation of the invalid text. The method can delete the invalid clip in response to the adjustment operation of the invalid text and the video clip deletion operation by a user, thereby improving the efficiency of editing the video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] TECHNICAL FIELD The embodiments of the present disclosure relate to the field of computer technology, and in particular to a video editing method, device, electronic device, and storage medium. [Background technology]

[0002] With the rapid development of science and technology, applications such as video emerge. For example, users can take selfie videos and upload them to video applications. The selfie or audio in front of the camera contains the core content of the selfie video, so how to better edit the selfie video is currently an urgent problem to be solved.

[0003] Selfie video editing mainly requires processing clips that are unnecessary for users, such as pauses, stuttering, and idle talk in the video. Traditional editing methods generally identify invalid clips by audio waveforms in the video, but this method is time-consuming and has relatively low editing efficiency. Summary of the Invention [Problem to be solved by the invention]

[0004] SUMMARY OF THE DISCLOSURE The embodiments of the present disclosure provide a video editing method, device, electronic device and storage medium to improve video editing efficiency. [Means for solving the problem]

[0005] In a first aspect, an embodiment of the present disclosure provides a method for editing a video, the method comprising: determining invalid text in the audio text of the target video and a position of the invalid text in a timeline (hereinafter, "timeline position") by audio identifying audio in the target video, the timeline position of the invalid text being for indicating a time of appearance of the audio audio of the invalid text in the target video; displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text on the editing track clip according to a timeline position of the invalid text, the section in the timeline of the invalid text (hereinafter, the "timeline section") being a time section during which the voice audio of the invalid text appears in the target video; adjusting a timeline interval of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; In response to a video clip deletion operation of the invalid text, deleting the video clip within the timeline section of the invalid text in the edit track clip from the target video.

[0006] In a second aspect, an embodiment of the present disclosure further provides a video editing apparatus, comprising: a text determination module for determining invalid text of a voice text of the target video and a timeline location of the invalid text by voice identifying audio in the target video, the timeline location of the invalid text being for indicating a time of appearance of the voice audio of the invalid text in the target video; a section marking module for displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text on the editing track clip according to a timeline position of the invalid text, the timeline section of the invalid text being a time section during which the voice audio of the invalid text appears in the target video; a segment adjustment module for adjusting a timeline segment of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; a clip deletion module for deleting, in response to a video clip deletion operation of the invalid text, a video clip within a timeline interval of the invalid text in the edit track clip from the target video.

[0007] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising: one or more processors; a memory for storing one or more programs; The one or more programs, when executed by the one or more processors, cause the one or more processors to implement the video editing method described in the embodiments of the present disclosure.

[0008] In a fourth aspect, the embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored therein, the computer-readable storage medium being capable of implementing the video editing method according to the embodiment of the present disclosure when the computer-readable storage medium is executed by a processor. Effect of the Invention

[0009] The embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a storage medium for editing a video, the method including: determining invalid text in an audio text of the target video and a timeline position of the invalid text by performing audio identification on an audio in the target video, the timeline position of the invalid text being for indicating an appearance time of the audio audio of the invalid text in the target video; displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline interval of the invalid text in the editing track clip according to the timeline position of the invalid text, the timeline interval of the invalid text being a time interval during which the audio audio of the invalid text appears in the target video; adjusting the timeline interval of the invalid text in the editing track clip of the target video in response to an adjustment operation of the invalid text; and deleting a video clip within the timeline interval of the invalid text in the editing track clip from the target video in response to a video clip deletion operation of the invalid text. According to the above technical solution, the invalid text and the location of the invalid text in the target video can be determined by voice identifying the audio in the target video, and the timeline section of the invalid text can be marked in the editing track clip of the target video, so that the invalid clip can be deleted in response to the user's operation of adjusting the invalid text and deleting the video clip, thereby completing the video editing and improving the efficiency of the video editing. [Brief description of the drawings]

[0010] The above and other features, advantages and aspects of each embodiment of the present disclosure will become more apparent with reference to the following specific embodiments in conjunction with the drawings. Throughout the drawings, like or similar drawing symbols indicate like or similar elements. It should be understood that the drawings are schematic and that objects and elements are not necessarily drawn to scale.

[0011] [Figure 1] 1 is a schematic flow chart of a video editing method according to an embodiment of the present disclosure;

[0012] [Diagram 2] FIG. 2 is a schematic diagram of triggering voice recognition according to an embodiment of the present disclosure.

[0013] [Diagram 3] FIG. 13 is a schematic diagram of an editing screen according to an embodiment of the present disclosure.

[0014] [Figure 4] 1 is a schematic flowchart of another video editing method according to an embodiment of the present disclosure.

[0015] [Diagram 5] FIG. 13 is a schematic diagram of another editing screen according to an embodiment of the present disclosure.

[0016] [Figure 6] FIG. 13 is a schematic diagram of updating a target text according to an embodiment of the present disclosure.

[0017] [Figure 7] FIG. 13 is another schematic diagram of updating a target text according to an embodiment of the present disclosure.

[0018] [Figure 8] FIG. 13 is a further schematic diagram of updating a target text according to an embodiment of the present disclosure.

[0019] [Figure 9] FIG. 1 is a structural schematic diagram of a video editing device according to an embodiment of the present disclosure.

[0020] [Figure 10] FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0021] Hereinafter, the embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the drawings show some embodiments of the present disclosure, it should be understood that the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments discussed herein. On the contrary, the purpose of providing these embodiments is to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely illustrative and are not intended to limit the scope of protection of the present disclosure.

[0022] It should be understood that the steps described in the method embodiments of the present disclosure may be performed in a different order and / or simultaneously. Also, the method embodiments may include additional steps and / or omit the performance of steps shown. The scope of the present disclosure is not limited in this respect.

[0023] As used herein, the term "comprises" and variations thereof are open inclusive, i.e., "including, but not limited to." The term "based on" means "based at least in part on." The term "in one embodiment" refers to "at least one embodiment," the term "in another embodiment" refers to "at least one other embodiment," and the term "in some embodiments" refers to "at least some embodiments." Relevant definitions of other terms are provided in the description below.

[0024] It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are merely intended to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of functions performed by these devices, modules or units.

[0025] It should be noted that the modifications "a" and "a plurality" referred to in the present disclosure are exemplary rather than limiting, and should be understood as "one or more" unless the context clearly indicates otherwise, as would be understood by one of ordinary skill in the art.

[0026] The names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0027] 1 is a schematic flow chart of a video editing method according to an embodiment of the present disclosure. The method is applied to a video editing case and can be executed by a video editing device. The device is realized by software and / or hardware and is usually integrated into an electronic device. In this embodiment, the electronic device includes, but is not limited to, devices such as a mobile phone or a computer.

[0028] As shown in FIG. 1, the video editing method according to the embodiment of the present disclosure includes the following steps S110 to S140.

[0029] At S110, by voice identifying the audio in the target video, invalid text in the voice text of the target video and a timeline position of the invalid text are determined, the timeline position of the invalid text is for indicating an appearance time of the voice audio of the invalid text in the target video.

[0030] The target video may refer to a video to be edited. For example, the target video may be a video uploaded by a user or a video taken online by a user. In this embodiment, the target video includes audio, and the voice text may be regarded as text information corresponding to the audio of the target video, and may be obtained by voice identifying the audio of the target video. Exemplarily, the voice text includes invalid text, and the invalid text may include text corresponding to clips in the target video that are unnecessary for a user, such as pauses, repetitions, and dead talk. The timeline position of the invalid text may be used to indicate the appearance time of the voice audio of the invalid text in the target video.

[0031] In one embodiment, the target video is a continuous video, for example, the target video may be a continuous and complete video or a clip of a complete video.

[0032] In this embodiment, the audio in the target video can be voice-identified to determine the invalid text and the timeline location of the invalid text in the audio text of the target video, so as to perform a subsequent editing step. This step does not limit the specific manner of voice identification and the trigger timing of voice identification, for example, the voice identification for the audio in the target video can be triggered by triggering the identification control on the screen. In this way, the invalid text and the timeline location of the invalid text in the target video can be automatically identified, and the position and type of the identification control are not limited and can be set according to the actual page situation.

[0033] In one embodiment, the criteria for identifying invalid text include that clips with mute time greater than a predefined time (e.g., 120 ms) are recognized as paused clips by default, and invalid words (e.g., repetition, emphatic words, etc.) are identified by logic that reuses a predefined algorithm, which may be deployed in the current application or may be deployed in a server.

[0034] 2 is a schematic diagram of triggering voice recognition according to an embodiment of the present disclosure. As shown in FIG. 2, after selecting a video (e.g., a target video), a voice recognition for the audio in the target video can be triggered by right-clicking the control 1 in a pop-up window. Furthermore, by clicking the shortcut control 2 in the current screen, voice recognition for the audio in the target video can be realized.

[0035] In step S120, an editing track clip of the target video is displayed on an editing screen of the target video, and a timeline interval of the invalid text is marked on the editing track clip according to a timeline position of the invalid text, the timeline interval being a time interval during which the voice audio of the invalid text appears in the target video.

[0036] The editing screen is used to edit clips in the target video. The editing track clip of the target video may be understood to be an editing track clip corresponding to the target video, for example, an editing track clip whose starting point corresponds to the starting point of the target video and whose ending point corresponds to the ending point of the target video. Exemplarily, when the target video is a continuous and complete video, the editing track clip may be an editing track of the continuous and complete video. When the target video is a clip of a complete video, the editing track clip may be a track clip corresponding to the target video in the editing track of the complete video. The editing track clip may be an audio track clip of the target video, and the audio track clip may refer to the audio track clip corresponding to the target video. The timeline interval of the invalid text may be regarded as a time interval during which the voice audio of the invalid text appears in the target video.

[0037] In one embodiment, after determining the invalid text and the timeline position of the invalid text in the target video by performing voice recognition on the audio in the target video, the editing track clip of the target video may be displayed on the editing screen of the target video, and the timeline section of the invalid text may be marked in the editing track clip according to the timeline position of the invalid text. The means for displaying the editing track clip and marking the timeline section of the invalid text is not limited, and may be preset by a related person according to the actual situation of the editing screen, for example, a predetermined identifier may be used to mark the timeline section of the invalid text in the editing track clip of the target video. The predetermined identifier may be set according to a specific actual situation. Exemplarily, the predetermined identifier may be a rectangular frame. In this case, the rectangular frame may box-select the timeline section of the invalid text in the target video in the editing track clip of the target video, or the timeline section of the invalid text may be marked in the editing track clip with a first display type, while the timeline section of other text may be marked in the editing track clip with a second display type. The first display type is different from the second display type. For example, the first display type and the second display type may be distinguished by different colors or by different lines.

[0038] In one embodiment, in order to distinguish between the timeline intervals of invalid text and the timeline intervals of other text, all the timeline intervals of invalid text identified and acquired in the edit track clip may be initially marked, or a set number of timeline intervals of invalid text may be initially marked in the edit track clip. The set number may be preset by a person concerned. The user may adjust the timeline intervals of invalid text as necessary.

[0039] 3 is a schematic diagram of an editing screen according to an embodiment of the present disclosure. As shown in FIG. 3, the editing screen of the target video may display an editing track clip 3 of the target video, and mark a timeline section of invalid text in the editing track clip 3 according to the timeline position of the invalid text identified in the previous step. For example, in FIG. 3, the timeline section 4 of invalid text is marked with a black rectangular frame.

[0040] At S130, in response to the adjustment operation of the invalid text, a timeline interval of the invalid text is adjusted in an edit track clip of the target video.

[0041] The invalid text adjustment operation may refer to an operation for triggering an adjustment of a timeline section of the invalid text. In this embodiment, the specific operation manner of the invalid text adjustment operation is not limited. For example, the invalid text adjustment operation may be an operation of clicking a control in an editing screen or a predetermined trigger operation acting on an editing track clip, or an operation of performing a predetermined gesture on an editing screen. The predetermined gesture may be a preset gesture. For example, the predetermined gesture may be a gesture of sliding right, etc.

[0042] In this embodiment, the trigger position of the adjustment operation is not limited, and may be, for example, act on the timeline section of the invalid text, or act on other positions. It is sufficient to be able to trigger the adjustment of the timeline section of the invalid text.

[0043] Specifically, the embodiment may adjust the timeline interval of the invalid text in the editing track clip of the target video in response to the adjustment operation of the invalid text. The specific manner of adjusting the timeline interval may be different according to different adjustment operations. For example, if the adjustment operation of the invalid text is to trigger a control to change the state of the invalid text from a selected state to an unselected state, the number of timeline intervals of the invalid text may be adjusted in the editing track clip of the target video. For example, the text may be adjusted to invalid text by adding a mark of a timeline interval of the text in the editing track clip, or the text may be adjusted to non-invalid text by canceling the mark of a time interval of the text in the editing track clip.

[0044] At S140, in response to a video clip deletion operation of the invalid text, the video clip within the timeline section of the invalid text in the edit track clip is deleted from the target video.

[0045] The invalid text video clip delete operation may refer to an operation for triggering the deletion of the video clip corresponding to the invalid text in the edit track clip, for example, the invalid text video clip delete operation may be an operation of clicking a delete control in an edit screen.

[0046] This step may delete the video clip in the timeline section of the invalid text in the edit track clip from the target video in response to the invalid text video clip deletion operation. The specific deletion manner is not limited, for example, the trace corresponding to the timeline section of the invalid text in the edit track clip may be kept or removed according to the user's need. As shown in FIG. 3, the triggering of the control 5 by the user may be regarded as a video clip deletion operation. After the user triggers the control 5, the electronic device may delete the video clip corresponding to the timeline section 6 of the invalid text from the target video in response to the invalid text video clip deletion operation.

[0047] In one embodiment, deleting the video clips within the timeline section of the invalid text in the edit track clip from the target video includes: Delete the video clips in the timeline section of the invalid text in the edit track clip from the target video, and divide the target video into a plurality of independent video clips using the deleted video clips as division points; or deleting the video clips in the timeline section of the invalid text in the editing track clip from the target video, and integrating the target video after deleting the video clips into one video; Includes.

[0048] In one embodiment, the video clips in the timeline section of the invalid text in the edit track clip may be deleted from the target video, and the deleted video clips may be used as division points to divide the target video into a plurality of independent video clips, which can facilitate the user to operate each independent video clip as needed.

[0049] In one embodiment, the video clips in the timeline section of the invalid text in the edit track clip may be deleted from the target video, and the target video after the video clips are deleted may be integrated into one video. Based on this, the deletion operation of the invalid clips in the target video is realized, and the quality of the target video is improved. The invalid clips are clips corresponding to the timeline section of the invalid text, that is, the text corresponding to the timeline section marked in the edit track clip in the audio text of the target video.

[0050] In a video editing method according to an embodiment of the present disclosure, an invalid text in the audio text of the target video and a timeline position of the invalid text are determined by voice-identifying audio in the target video, the timeline position of the invalid text is for indicating an appearance time of the audio audio of the invalid text in the target video, an editing track clip of the target video is displayed on an editing screen of the target video, and a timeline section of the invalid text is marked on the editing track clip according to the timeline position of the invalid text, the timeline section of the invalid text is a time section during which the audio audio of the invalid text appears in the target video, the timeline section of the invalid text is adjusted in the editing track clip of the target video in response to an adjustment operation of the invalid text, and a video clip within the timeline section of the invalid text in the editing track clip is deleted from the target video in response to a video clip deletion operation of the invalid text. The method can determine the invalid text and the position of the invalid text in the target video by voice-identifying audio in the target video. Furthermore, by marking the timeline section of the invalid text in the edit track clip of the target video, the invalid clip can be deleted in response to the user's operation of adjusting the invalid text and deleting the video clip, thereby completing the video editing and improving the efficiency of the video editing.

[0051] In one embodiment, the method further comprises: playing the target video in a video area of ​​the editing screen; The method further includes, when the target video is played to a start position of a second timeline section, adjusting a playback progress of the target video to a first time node, and continuing to play the target video from the first time node. The second timeline interval is a timeline interval of the invalid text, and the first time node is a time node corresponding to an end position of the second timeline interval in the target video.

[0052] The video area may be regarded as an area in an editing screen for playing a target video, and the specific position and size of the video area may be set by a related person according to the situation of the editing screen. The second timeline section may be a timeline section of any one of the invalid texts. The first time node may be understood as a time node corresponding to the end position of the second timeline section in the target video.

[0053] In one embodiment, the target video may be played in the video area of ​​the editing screen. The timing of playing the target video is not limited, for example, the target video may be automatically played when the editing screen is displayed, or the target video may be played in the video area of ​​the editing screen when the user triggers a certain playback control in the editing screen after the editing screen is displayed, or the target video may be automatically played after the editing screen is displayed for a predetermined time, and the present embodiment is not limited thereto.

[0054] When the target video is played back to the start position of the second timeline section, this embodiment may automatically adjust the playback progress of the target video to the first time node, and continue to play the target video from the first time node. Based on this, invalid clips are automatically skipped when playing the target video, so that the user can preview the edited video to better edit the target video.

[0055] In one embodiment, after playing the target video in the video area of ​​the editing screen, The method further includes adjusting a playback progress of the target video to a second time node in response to a trigger operation on a fifth text, and continuing to play the target video from the second time node, wherein the fifth text is displayed within a text area of ​​the editing screen, the fifth text being invalid text or non-invalid text, and the second time node being a time node corresponding to a start position of the fifth text in the target video.

[0056] In this embodiment, the fifth text may be a text triggered by a user, or may be an invalid text or a non-invalid text. The non-invalid text is a text other than the invalid text, for example, a valid text in the audio text of the target video. The fifth text may be displayed in a text area of ​​the editing screen. The text area may be regarded as an area in the editing screen, which is used to display the audio text of the target video, for example, the fifth text in the audio text of the target video. The specific position and size of the text area may be set by the relevant person according to the situation of the editing screen.

[0057] In one embodiment, the disabled text in the text area may remain as a paragraph, for example with a separate line of "pause" text and the duration of the pause.

[0058] A trigger operation on the fifth text may trigger adjusting the playback progress of the target video to a second time node. For example, the trigger operation on the fifth text may be an operation of clicking on the fifth text. The second time node is a time node corresponding to the start position of the fifth text in the target video.

[0059] Specifically, after the target video is played in the video area of ​​the editing screen, the playback progress of the target video may be adjusted to the second time node in response to the trigger operation of the fifth text, and the target video may continue to be played from the second time node. On this basis, the playback progress of the target video is adjusted in response to the trigger operation of the fifth text, thereby realizing the positioning of the corresponding text in the target video, and facilitating the user to watch the video clip corresponding to the corresponding text.

[0060] In one embodiment, a time identifier is further displayed in the edit track clip. After playing the target video in the video area of ​​the edit screen, determining, in response to an justification operation on a target feature, a third timeline segment in which a target position of the target feature is located, the target position being a position of the target feature in the edit track clip upon completion of performance of the justification operation; if the third timeline interval is the timeline interval of the invalid text, adjust a playback progress of the target video to a third time node, and continue to play the target video from the third time node, where the third time node is a time node corresponding to a start position or an end position of the third timeline interval in the target video, or a time node until the target video was last played before receiving the position adjustment operation; if the third timeline interval is not the timeline interval of the invalid text, adjust a playback progress of the target video to a fourth time node, and continue to play the target video from the fourth time node, where the fourth time node is a time node corresponding to the target position in the target video; Further includes:

[0061] The time identifier is for representing a time node until the target video is currently played in the edit track clip of the target video. In other words, the time identifier can indicate a time node corresponding to the current playback progress of the target video. When the current application plays the target video in the video area of ​​the editing screen, the position of the target video in the edit track clip of the target mark can be updated in real time based on the playback progress of the target video. The position adjustment operation of the target mark may be understood as an operation of adjusting the position of the target mark. For example, the position adjustment operation of the target mark may be an operation of adjusting the target mark to a certain position a in the edit track clip by clicking the position a, or an operation of adjusting the target mark to the position a by dragging the target mark. The target position may be considered to be the position of the target mark in the edit track clip when the execution of the position adjustment operation is completed. The third timeline section is a timeline section in which the target position of the target mark is located.

[0062] The third time node may be a time node corresponding to a start position of the third timeline interval in the target video, or a time node corresponding to an end position of the third timeline interval in the target video, or a time node corresponding to the last playing of the target video before receiving a position advance adjustment operation. The fourth time node may be considered to be a time node corresponding to a target position in the target video.

[0063] In this step, after the target video is played back in the video area of ​​the editing screen, the position of the target mark may be adjusted in response to a position adjustment operation for the target mark.

[0064] For example, the target mark position adjustment operation may include adjusting the playback progress of the target video according to the position of the time identifier in the edit track clip, and when the position adjustment operation is completed, determining a target position of the target mark, determining a third timeline section in which the target position is located, and making different adjustments to the playback progress of the target video according to different third timeline sections.

[0065] In an embodiment, if the third timeline section is a timeline section of invalid text, the playback progress of the target video may be adjusted to the fourth time node, and the target video may continue to be played from the fourth time node. For example, when a target marker is dragged into the timeline section A of a certain invalid text, the embodiment may automatically adjust the playback progress of the target video to the end position of the timeline section A, and continue to play the target video from the end position of the timeline section A, or may return the position of the target marker to the position where the target marker was last located before receiving the position adjustment operation, and continue to play the target video from the position. When a position in the timeline section A marked on the edit track clip of the target video is clicked, the embodiment may automatically adjust the playback progress of the target video to the start position of the timeline section A, and continue to play the target video from the start position of the timeline section A.

[0066] In an embodiment, if the third timeline section is not the timeline section of the invalid text, the playback progress of the target video may be adjusted to a time node corresponding to the target position in the target video, and the target video may continue to be played from the time node. For example, if the target position b until the target mark is dragged is not located in the timeline section of the invalid text, or the target position b until the target mark is moved by clicking is not located in the timeline section of the invalid text, the embodiment may adjust the playback progress of the target video to a time node c corresponding to the target position b in the target video, and the target video may continue to be played from the time node c. Based on this, in response to the position adjustment operation for the target mark, if the adjusted position of the target mark is located in the invalid clip, the video clip corresponding to the invalid text can be automatically skipped, or the video clip corresponding to the invalid text can be played completely from the start position of the video clip corresponding to the invalid text, but if the adjusted position of the target mark is not located in the invalid clip, the target video is played from the time node corresponding to this position.

[0067] FIG. 4 is a schematic flow chart of another video editing method according to an embodiment of the present disclosure. The technical solution of this embodiment can be combined with one or more optional technical solutions of the above embodiments. Adjusting the timeline interval of the invalid text in the editing track clip of the target video in response to the adjustment operation of the invalid text may include at least one of: adjusting a length of the timeline interval of the invalid text in the editing track clip of the target video in response to a first adjustment operation of the invalid text; and adjusting a number of timeline intervals marked in the editing track clip of the target video in response to a second adjustment operation of the invalid text. As shown in FIG. 4, the method includes the following steps S210 to S250.

[0068] In S210, by voice identifying the audio in the target video, invalid text in the voice text of the target video and a timeline position of the invalid text are determined, the timeline position of the invalid text is for indicating an appearance time of the voice audio of the invalid text in the target video.

[0069] In step S220, an editing track clip of the target video is displayed on an editing screen of the target video, and a timeline interval of the invalid text is marked on the editing track clip according to a timeline position of the invalid text, the timeline interval being a time interval during which the voice audio of the invalid text appears in the target video.

[0070] At S230, in response to the first adjustment operation of the invalid text, a length of a timeline section of the invalid text in an edit track clip of the target video is adjusted.

[0071] The first adjustment operation may be for adjusting a length of a timeline section of the invalid text. The first adjustment operation may act on an editing track clip of the target video. Exemplarily, the first adjustment operation may be an operation for stretching or compressing a timeline section B of the invalid text in the editing track clip of the target video, such as dragging a start position of the timeline section B to the left.

[0072] This step may adjust a length of a timeline section of the invalid text in the edit track clip of the target video in response to the first adjustment operation of the invalid text. For example, the section length of a timeline section may be extended when the left boundary of the timeline section is dragged left or the right boundary of the timeline section is dragged right. The section length of a timeline section may be shortened when the left boundary of the timeline section is dragged right or the right boundary of the timeline section is dragged right.

[0073] At S240, in response to the second adjusting operation of the invalid text, the number of timeline segments marked in the edit track clip of the target video is adjusted.

[0074] The second adjustment operation may be for adjusting the number of timeline segments marked in the edit track clip. For example, the second adjustment operation may act on an adjustment control. Each invalid text may correspond to one adjustment control. The adjustment control is for controlling whether the timeline segment corresponding to the invalid text is marked in the edit track clip. The second adjustment operation may act on a timeline segment of the edit track clip.

[0075] This step may adjust the number of timeline segments marked in the edit track clip of the target video in response to the second adjustment operation of the invalid text. Although the details of the adjustment are not limited in this embodiment, it is sufficient that the number of timeline segments marked in the edit track clip can be adjusted.

[0076] At S250, in response to a video clip deletion operation of the invalid text, the video clip within the timeline section of the invalid text in the edit track clip is deleted from the target video.

[0077] In the video editing method according to the embodiment of the present disclosure, the length and number of timeline segments of the invalid text in the editing track clip of the target video can be adjusted in response to the first and second adjustment operations of the invalid text, thereby realizing accurate editing of the target video according to the needs of a user.

[0078] In one embodiment, displaying the edit track clip of the target video on the edit screen of the target video comprises: displaying an edit track clip of the target video in a track area of ​​the editing screen, and displaying invalid text and non-invalid text of the audio text in a text area of ​​the editing screen, the invalid text being displayed in a selected state on the editing screen, and the non-invalid text being displayed in a non-selected state on the editing screen.

[0079] The track area is an area within the editing screen that may be used to display the editing track clips of the target video. The selected state may be understood as a state in which some text is selected, and this text is going to be deleted. The unselected state may be understood as a state in which some text is not selected, and this text is not going to be deleted.

[0080] In this embodiment, displaying the edit track clip of the target video on the edit screen of the target video includes displaying the edit track clip of the target video in the track area of ​​the edit screen, and displaying the invalid text and non-invalid text of the voice text in the text area of ​​the edit screen. Here, the invalid text may be displayed in a selected state on the edit screen, and the non-invalid text may be displayed in a non-selected state on the edit screen. Based on this, displaying the edit track clip is realized, and the invalid text and non-invalid text of the voice text are displayed in the text area of ​​the edit screen. In addition, the invalid text is displayed in a selected state by default. The invalid text can be presented to the user so that it can be deleted. In addition, it is avoided that the user must perform a corresponding selection operation to select invalid text such as pause or repeat before it can be deleted, and the operation required for the user to delete the invalid text in the target video is further simplified.

[0081] 5 is a schematic diagram of another editing screen according to an embodiment of the present disclosure. As shown in FIG. 5, an editing track clip of a target video is displayed in a track area 7 of the editing screen, and invalid text (e.g., "pause 0.12 s" or "repeat") and non-invalid text (e.g., "my") of the audio text are displayed in a text area 8 of the editing screen. Here, the invalid text is displayed in a selected state on the editing screen, and the non-invalid text is displayed in a non-selected state on the editing screen.

[0082] In one embodiment, adjusting the length of the timeline segment of the invalid text in the edit track clip of the target video includes: and adjusting a length of a first timeline segment of a first text in an edit track clip of the target video and updating a target text displayed in the text area, the target text including the first text and / or a second text related to the first text.

[0083] The first text may be invalid text, for example, text to which the timeline interval targeted by the first trigger operation belongs. The first timeline interval is the timeline interval of the first text. The target text may refer to text to be adjusted depending on the situation. For example, the target text may include the first text and / or a second text related to the first text. The second text may be text corresponding to a timeline interval adjacent to the timeline interval of the first text, or may be text corresponding to a new timeline interval after reducing the timeline interval adjacent to the timeline interval of the first text.

[0084] Specifically, the length of the first timeline section of the first text in the edit track clip of the target video needs to be adjusted, and the target text displayed in the text area needs to be updated, and the specific content to be updated is determined by the specific operation of the first adjustment operation. Based on this, the length of the first timeline section is adjusted, and the text corresponding to the adjusted portion in the text area is updated in real time, so that the time corresponding to the timeline section of the invalid text can be accurately adjusted.

[0085] In one embodiment, the first adjustment operation includes a lengthening operation, and updating the target text for display in the text area includes: The method includes moving text content in the second text corresponding to an extended portion of the first timeline interval to the first text.

[0086] A length extension operation may be considered to be an operation that extends the length of the first timeline section. For example, a length extension operation may be an operation of dragging the start position of the first timeline section to the left or dragging the end position of the first timeline section to the right.

[0087] In one embodiment, when the length of the first timeline interval of the first text is extended in the edit track clip, text content in the second text in the text area corresponding to the extended portion of the first timeline interval may be synchronously moved to the first text. For example, when the first adjustment operation is an operation of dragging the start position of the first timeline interval to the left, the second text is the text corresponding to the timeline interval located to the left of the first timeline interval. In this case, the text content in the second text corresponding to the extended portion of the first timeline interval needs to be moved to the first text.

[0088] In one embodiment, the length extension operation may support moving a portion of the text of the adjacent non-invalid text into the first text, and may or may not support performing a length extension operation if the second text is invalid text.

[0089] 6 is a schematic diagram of updating a target text according to an embodiment of the present disclosure. As shown in FIG. 6, when the end position of the first timeline section corresponding to the first text 9 is dragged right from position A to position B, the length of the first timeline section of the first text 9 may be extended in the edit track clip of the target video, and the text content (e.g., "sun") corresponding to the extended part of the first timeline section in the second text 10 may be moved to the first text 9.

[0090] In one embodiment, the first adjustment operation comprises a length shortening operation. moving text content in the first text corresponding to the shortened portion of the first timeline interval to the second text; or adding a second text to the text area and moving text content in the first text corresponding to the shortened portion of the first timeline interval to the second text; Includes.

[0091] A length shortening operation may be considered to be an operation that shortens the length of the first timeline interval. For example, a length shortening operation may be an operation of dragging the start position of the first timeline interval to the right, or an operation of dragging the end position of the first time interval to the left.

[0092] In an embodiment, when shortening the length of the first timeline section of the first text in the edit track clip, an update manner of the target text displayed in the text area may be determined according to the difference of the first text content. For example, when the first text content is a pause, the text content corresponding to the shortened portion of the first timeline section in the first text may be moved to the second text, whereas when the first text content is a repeat text, an invalid word, or a normal text other than the pause, repeat, and invalid word, the second text may be added to the text area, and the text content corresponding to the shortened portion of the first timeline section in the first text may be moved to the added second text.

[0093] 7 is another schematic diagram of updating target text according to an embodiment of the present disclosure. As shown in FIG. 7, when the end position of the first timeline section corresponding to the first text 9 is dragged leftward from position A to position C, the length of the first timeline section of the first text 9 may be shortened in the edit track clip of the target video, and the text content (e.g., “pause for 0.02 s”) corresponding to the shortened part of the first timeline section in the first text 9 may be moved to the second text 10.

[0094] 8 is a further schematic diagram of updating a target text according to an embodiment of the present disclosure. As shown in FIG. 8, when the end position of the first timeline section corresponding to the first text 11 is dragged right from position D to position E, the length of the first timeline section of the first text 11 may be shortened in the edit track clip of the target video. Meanwhile, the second text 12 may be added to the text area, and the text content (i.e., "repeat") corresponding to the shortened part of the first timeline section in the first text 11 may be moved to the second text 12.

[0095] In one embodiment, the override text is selected text in the track area. Adjusting the number of timeline segments marked in the edit track clip of the target video in response to a second adjustment operation of the override text includes: in response to a selection operation on a third text in the text area, switching the third text from an unselected state to a selected state and adding an indication of a timeline interval of the third text in the editing track clip; In response to a deselection operation on a fourth text in the text area, the method includes at least one of switching the fourth text from a selected state to an unselected state and canceling the marking of a timeline section of the fourth text in the editing track clip.

[0096] The third text may be any one of the texts in the text area, for example, text in an unselected state in the text area. The selection operation may refer to an operation of switching the text from an unselected state to a selected state. For example, the selection operation may be an operation of clicking a control. The control may have a first display type and a second display type. Also, the selection state of the text may change as the display type of the control changes, and the details of the first display type and the second display type are not limited as long as different display types can be distinguished.

[0097] The fourth text may be any one of the texts in the text region, for example, the selected text in the text region. The fourth text may be the same as the third text or may be different. The deselection operation may refer to an operation of switching the text from a selected state to an unselected state. For example, the deselection operation may be an operation of clicking a control corresponding to the selection operation.

[0098] Specifically, this embodiment may switch the third text from an unselected state to a selected state in response to a selection operation on the third text in the text area, and add a mark of the timeline section of the third text to the edit track clip, so that the user can add the third text and delete the timeline section of the third text as needed.

[0099] In this embodiment, in response to a selection cancel operation on the fourth text in the text area, the fourth text may be switched from a selected state to an unselected state, and the marking of the timeline section of the fourth text in the edit track clip may be canceled. Based on this, the user may cancel the deletion of the fourth text and the marking of the timeline section of the fourth text as necessary.

[0100] For example, clicking a control to display the control in a first display type may be considered to trigger a selection operation on a third text corresponding to the control. In this case, the third text may be switched from an unselected state to a selected state, and a timeline section marking of the third text may be added in the editing track clip. Clicking a control to display the control in a second display type may be considered to trigger a selection cancellation operation on a fourth text corresponding to the control. In this case, the fourth text may be switched from a selected state to an unselected state, and a timeline section marking of the fourth text may be cancelled in the editing track clip.

[0101] 9 is a structural schematic diagram of a video editing device according to an embodiment of the present disclosure. The device may be applied to video editing. The device may be realized by software and / or hardware, and is usually integrated into electronic equipment.

[0102] As shown in FIG. 9, the device includes: a text determination module 310 for determining invalid text of a voice text of the target video and a timeline location of the invalid text by voice identifying audio in the target video, the timeline location of the invalid text being for indicating a time of appearance of the voice audio of the invalid text in the target video; a section marking module 320 for displaying an edit track clip of the target video on an edit screen of the target video and marking a timeline section of the invalid text in the edit track clip based on a timeline position of the invalid text, the timeline section of the invalid text being a time section during which the voice audio of the invalid text appears in the target video; a segment adjustment module 330 for adjusting a timeline segment of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; a clip deletion module 340 for deleting a video clip within a timeline section of the invalid text in the edit track clip from the target video in response to a video clip deletion operation of the invalid text; Equipped with.

[0103] In the video editing device according to the embodiment of the present disclosure, a text determination module determines invalid text in the audio text of the target video and a timeline position of the invalid text by voice-identifying audio in the target video, the timeline position of the invalid text is for indicating the appearance time of the voice audio of the invalid text in the target video, a section marking module displays an editing track clip of the target video on an editing screen of the target video, and marks a timeline section of the invalid text in the editing track clip according to the timeline position of the invalid text, the timeline section of the invalid text is a time section during which the voice audio of the invalid text appears in the target video, a section adjustment module adjusts the timeline section of the invalid text in the editing track clip of the target video in response to an adjustment operation of the invalid text, and a clip deletion module deletes a video clip within the timeline section of the invalid text in the editing track clip from the target video in response to a video clip deletion operation of the invalid text. According to the device, the invalid text and the position of the invalid text in the target video can be determined by voice-identifying audio in the target video. Furthermore, by marking the timeline section of the invalid text in the edit track clip of the target video, the invalid clip can be deleted in response to the user's operation of adjusting the invalid text and deleting the video clip, thereby completing the video editing and improving the efficiency of the video editing.

[0104] Optionally, the section adjustment module 330 may specifically: a first response unit for adjusting a length of a timeline interval of the invalid text in an edit track clip of the target video in response to a first adjustment operation of the invalid text; a second response unit for adjusting a number of timeline segments marked in an edit track clip of the target video in response to the second adjusting operation of the invalid text; The present invention is provided with at least one of the following:

[0105] Optionally, the section marking module 320 may be specifically configured to: The edit screen is used to display an edit track clip of the target video in a track area of ​​the edit screen, and to display invalid text and non-invalid text of the voice text in a text area of ​​the edit screen, where the invalid text is displayed in a selected state on the edit screen, and the non-invalid text is displayed in a non-selected state on the edit screen.

[0106] Optionally, the first response unit comprises a coordination sub-unit. The adjusting subunit is used for adjusting a length of a first timeline section of a first text in an edit track clip of the target video, and updating a target text to be displayed in the text area, where the target text includes the first text and / or a second text related to the first text.

[0107] Optionally, the first adjustment operation includes a length extension operation. The adjustment subunit specifically includes: The text content corresponding to the extended portion of the first timeline interval in the second text is moved to the first text.

[0108] Optionally, the first adjustment operation includes a length shortening operation. The adjustment subunit specifically includes: moving text content in the first text that corresponds to the shortened portion of the first timeline interval to the second text; or It is used to add a second text to the text area and move text content in the first text corresponding to the shortened portion of the first timeline interval to the second text.

[0109] Alternatively, the invalid text is a selected text in the track area. in response to a selection operation on a third text in the text area, switching the third text from an unselected state to a selected state and adding an indication of a timeline interval of the third text in the editing track clip; In response to a deselection operation on a fourth text in the text area, the fourth text is switched from a selected state to an unselected state, and the marking of the timeline section of the fourth text in the editing track clip is cancelled.

[0110] As an option, a video editing device according to an embodiment of the present disclosure may include: a playback module for playing the target video in a video area of ​​the editing screen; a first progress adjusting module for adjusting a playback progress of the target video to a first time node when the target video is played to a start position of a second timeline section, and continuing to play the target video from the first time node; It further comprises: The second timeline interval is a timeline interval of the invalid text, and the first time node is a time node corresponding to an end position of the second timeline interval in the target video.

[0111] Optionally, the video editing apparatus according to the embodiment of the present disclosure further comprises a first response module. After playing the target video in a video area of ​​the editing screen, the first response module adjusts a playback progress of the target video to a second time node in response to a trigger operation on a fifth text, and continues to play the target video from the second time node. The fifth text is displayed in a text area of ​​the editing screen. The fifth text is invalid text or non-invalid text. The second time node is a time node corresponding to a start position of the fifth text in the target video.

[0112] Optionally, a time identifier is further displayed in the edit track clip. a second response module for determining a third timeline section in which a target position of the target mark is located in response to a position adjustment operation on the target mark after playing the target video in a video area of ​​the editing screen, the target position being a position of the target mark in the editing track clip when execution of the position adjustment operation is completed; a second progress adjustment module for adjusting a playback progress of the target video to a third time node when the third timeline interval is a timeline interval of the invalid text, and continuing to play the target video from the third time node, the third time node being a time node corresponding to a start position or an end position of the third timeline interval in the target video, or a time node until the target video was last played before receiving the position adjustment operation; a third progress adjusting module for adjusting a playback progress of the target video to a fourth time node and continuing to play the target video from the fourth time node when the third timeline section is not a timeline section of the invalid text, the fourth time node being a time node corresponding to the target position in the target video; It further comprises:

[0113] Optionally, the clip deletion module 340 may specifically: Delete the video clips in the timeline section of the invalid text in the editing track clip from the target video, and divide the target video into a plurality of independent video clips using the deleted video clips as division points; or The video clips in the timeline section of the invalid text in the edit track clip are deleted from the target video, and the target video after the video clips are deleted is integrated into one video.

[0114] Optionally, the edit track clip is an audio track clip of the target video.

[0115] The video editing apparatus can execute the video editing method according to any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0116] Referring to FIG. 10 below, a structural schematic diagram of an electronic device 400 suitable for implementing the embodiment of the present disclosure is shown. The terminal device of the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG. 10 is merely an example and should not limit the functions and scope of use of the embodiment of the present disclosure.

[0117] 10, the electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 401. The processing unit can perform various appropriate operations and processes based on a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. The RAM 403 further stores various programs and data necessary for the operation of the electronic device 400. The processing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0118] Typically, input devices 406 including a touch panel, touch pad, keyboard, mouse, web camera, microphone, accelerometer, gyroscope, etc., output devices 407 including a liquid crystal display (LCD), loudspeaker, oscillator, etc., storage devices 408 including magnetic tape, hard disk, etc., and communication devices 409 may be connected to the I / O interface 405. The communication devices 409 may allow the electronic device 400 to exchange data with other devices by wireless or wired communication. Although FIG. 10 illustrates the electronic device 400 with various devices, it should be understood that it is not required to implement or include all of the illustrated devices. More or fewer devices may alternatively be implemented or included.

[0119] In particular, according to the embodiment of the present disclosure, the above flow described with reference to the flowchart may be implemented as a computer software program. For example, the embodiment of the present disclosure includes a computer program product, the computer program product includes a computer program included in a non-transitory computer readable medium, the computer program includes a program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed via a network by the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, it performs the above functions limited to the method of the embodiment of the present disclosure.

[0120] It should be noted that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of both. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program. The program may be used by or in combination with an instruction execution system, apparatus, or device. And, in the present disclosure, the computer-readable signal medium may be included in a data signal propagated as part of a baseband or carrier wave, in which the computer-readable program code is included. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium. The computer-readable signal medium may transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted via any suitable medium, including, but not limited to, electrical wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0121] In some embodiments, the client terminals and servers may communicate using any now known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), extranets (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any now known or later developed networks.

[0122] The computer-readable medium may be included in the electronic device, or may exist independently of the electronic device.

[0123] The computer-readable medium includes one or more programs that, when executed by an electronic device, determining null text of the voice text of the target video and a timeline location of the null text by voice identifying audio in the target video, the timeline location of the null text being for indicating a time of appearance of the voice audio of the null text in the target video; displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text in the editing track clip according to a timeline position of the invalid text, the timeline section of the invalid text being a time section during which the voice audio of the invalid text appears in the target video; adjusting a timeline interval of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; in response to a video clip deletion operation of the invalid text, deleting the video clip within the timeline section of the invalid text in the edit track clip from the target video; The electronic device is caused to execute the above.

[0124] Computer program code for carrying out the operations of the present disclosure may be written in one or more programming languages ​​or combinations thereof. Such programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and may further include conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected via the Internet using an Internet Service Provider).

[0125] The flowcharts and block diagrams in the drawings illustrate possible system architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or part of code. The module, program segment, or part of code includes one or more executable instructions for implementing a specified logical function. It should be noted that in some permutation implementations, the functions depicted in the blocks may occur in a different order than the order depicted in the drawings. For example, two blocks depicted may actually be executed essentially simultaneously, or they may be executed in the reverse order, as determined by the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that executes the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0126] The units described in the embodiments of the present disclosure may be implemented in a software or hardware manner, and the names of the modules may not constitute limitations on the units themselves.

[0127] The functionality described herein above may be performed at least in part by one or more hardware logic components. For example, and without limitation, model types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific general purpose products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.

[0128] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program used by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection by one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0129] According to one or more embodiments of the present disclosure, Example 1 provides a method for editing a video, the method comprising: determining invalid text of spoken text of the target video and a timeline location of the invalid text by speech identifying audio in the target video, the timeline location of the invalid text indicating a time of occurrence of speech audio of the invalid text in the target video; Displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text on the editing track clip according to a timeline position of the invalid text, the timeline section of the invalid text being a time section during which the voice audio of the invalid text appears in the target video; adjusting a timeline interval of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; in response to a video clip deletion operation of the invalid text, deleting the video clip within the timeline section of the invalid text in the edit track clip from the target video; Includes.

[0130] According to one or more embodiments of the present disclosure, in accordance with the method of Example 1, in Example 2, adjusting a timeline segment of the invalid text in an edit track clip of the target video in response to an adjustment operation of the invalid text includes: adjusting a length of a timeline segment of the invalid text in an edit track clip of the target video in response to the first adjustment operation of the invalid text; adjusting a number of timeline segments marked in the edit track clip of the target video in response to the second adjustment operation of the invalid text; Includes at least one of the following:

[0131] According to one or more embodiments of the present disclosure, in Example 3, according to the method of Example 2, displaying an editing track clip of the target video on an editing screen of the target video includes: displaying an edit track clip of the target video in a track area of ​​the editing screen, and displaying invalid text and non-invalid text of the audio text in a text area of ​​the editing screen, the invalid text being displayed in a selected state on the editing screen, and the non-invalid text being displayed in a non-selected state on the editing screen.

[0132] According to one or more embodiments of the present disclosure, in Example 4, according to the method described in Example 3, adjusting a length of a timeline segment of the invalid text in an editing track clip of the target video includes: and adjusting a length of a first timeline segment of a first text in an edit track clip of the target video to update a target text displayed in the text area, the target text including the first text and / or a second text related to the first text.

[0133] According to one or more embodiments of the present disclosure, in Example 5, based on the method of Example 4, the first adjustment operation includes a length extension operation. Updating the target text to be displayed in the text area includes: The method includes moving text content in the second text corresponding to an extended portion of the first timeline interval to the first text.

[0134] According to one or more embodiments of the present disclosure, Example 6 is based on the method of Example 4, and the first adjustment operation includes a length shortening operation. Updating the target text for displaying in the text area includes: moving text content in the first text corresponding to the shortened portion of the first timeline interval to the second text; or adding a second text to the text area and moving text content in the first text corresponding to the shortened portion of the first timeline interval to the second text; Includes.

[0135] According to one or more embodiments of the present disclosure, Example 7 may be based on the method of Example 3, wherein the invalid text is selected text in the track area, and adjusting a number of marked timeline segments in an editing track clip of the target video in response to a second adjustment operation of the invalid text includes: in response to a selection operation on a third text in the text area, switching the third text from an unselected state to a selected state and adding an indication of a timeline interval of the third text in the editing track clip; in response to a deselection operation on a fourth text in the text area, switching the fourth text from a selected state to an unselected state and canceling marking of a timeline section of the fourth text in the editing track clip; Includes at least one of the following:

[0136] According to one or more embodiments of the present disclosure, Example 8 is based on the method according to any one of Examples 1 to 7, the method further comprising: playing the target video in a video area of ​​the editing screen; The method further includes, when the target video is played back to a start position of a second timeline section, adjusting a playback progress of the target video to a first time node, and continuing to play the target video from the first time node. The second timeline interval is a timeline interval of the invalid text, and the first time node is a time node corresponding to an end position of the second timeline interval in the target video.

[0137] According to one or more embodiments of the present disclosure, in Example 9, according to the method of Example 8, after playing the target video in a video area of ​​the editing screen, The method further includes adjusting a playback progress of the target video to a second time node in response to a trigger operation on a fifth text, and continuing to play the target video from the second time node. The fifth text is displayed in a text area of ​​the editing screen. The fifth text is invalid text or non-invalid text. The second time node is a time node corresponding to a start position of the fifth text in the target video.

[0138] According to one or more embodiments of the present disclosure, Example 10 further comprises the method according to Example 8, further displaying a time identifier in the editing track clip. After playing the target video in a video area of ​​the editing screen, determining, in response to an justification operation on a target mark, a third timeline interval in which a target position of the target mark is located, the target position being a position of the target mark in the edit track clip upon completion of performance of the justification operation; if the third timeline interval is the timeline interval of the invalid text, adjust a playback progress of the target video to a third time node and continue playing the target video from the third time node, the third time node being a time node corresponding to a start position or an end position of the third timeline interval in the target video, or a time node until the target video was last played before receiving the position adjustment operation; If the third timeline interval is not the timeline interval of the invalid text, adjust the playback progress of the target video to a fourth time node, and continue to play the target video from the fourth time node, where the fourth time node is a time node corresponding to the target position in the target video; Further includes:

[0139] According to one or more embodiments of the present disclosure, in Example 11, according to the method according to any one of Examples 1 to 7, deleting a video clip within a timeline section of the invalid text in the editing track clip from the target video includes: Delete the video clips in the timeline section of the invalid text in the editing track clip from the target video, and divide the target video into a plurality of independent video clips using the deleted video clips as division points; or removing the video clips in the timeline section of the invalid text in the edit track clip from the target video, and integrating the target video after removing the video clips into one video; Includes.

[0140] According to one or more embodiments of the present disclosure, Example 12 is a method according to any one of Examples 1 to 7, wherein the editing track clip is an audio track clip of the target video.

[0141] The above description is merely a description of the preferred embodiments and applied technical principles of the present disclosure. As will be understood by those skilled in the art, the scope of the disclosure of the present disclosure is not limited to the technical solution obtained by specifically combining the above technical features, but should also include other technical guidance obtained by arbitrarily combining the above technical features or their equivalent features without departing from the concept disclosed above. For example, the above features may be substituted with technical features having similar functions (not limited to those) disclosed in the present disclosure.

[0142] Also, although operations have been described in a particular order, it should not be understood that these operations require the particular order shown or sequential order. In certain circumstances, multitasking and simultaneous processing may be advantageous. Similarly, although the above discussion includes some specific implementation details, these should not be construed as limiting the scope of the disclosure. Some features that are described in the context of separate embodiments may also be combined and implemented in a single embodiment. Conversely, various features that are described in the context of a single embodiment may be implemented in multiple embodiments, either independently or in any suitable subcombination manner.

[0143] Although the present subject matter has been described in language specifying structural features and / or methodological operations, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are merely example forms of implementing the claims.

Claims

1. determining invalid text of a speech text of the target video and a timeline location of the invalid text by speech identifying audio in the target video, the timeline location of the invalid text being for indicating a time of appearance of speech audio of the invalid text in the target video; Displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text in the editing track clip according to a timeline position of the invalid text, the timeline section of the invalid text being a time section during which the voice audio of the invalid text appears in the target video; adjusting a timeline interval of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; in response to a video clip deletion operation of the invalid text, deleting the video clip within the timeline section of the invalid text from the target video in the edit track clip; 16. A method for editing a video, comprising:

2. adjusting a timeline segment of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation, adjusting a length of a timeline segment of the invalid text in an edit track clip of the target video in response to the first adjustment operation of the invalid text; adjusting a number of timeline segments marked in the edit track clip of the target video in response to the second adjustment operation of the invalid text; 2. The method of claim 1, further comprising at least one of:

3. Displaying an edit track clip of the target video on an edit screen of the target video includes: displaying an edit track clip of the target video in a track area of ​​the edit screen, and displaying invalid text and non-invalid text of the voice text in a text area of ​​the edit screen; The invalid text is displayed in a selected state on the editing screen, and the non-invalid text is displayed in a non-selected state on the editing screen.

3. The method of claim 2 .

4. Adjusting a length of the timeline section of the invalid text in the edit track clip of the target video includes: adjusting a length of a first timeline section of a first text in an edit track clip of the target video to update the target text displayed in the text area; The target text includes the first text and / or a second text related to the first text.

4. The method according to claim 3 .

5. The first adjustment operation includes a length extension operation, and updating the target text displayed in the text area includes: and moving text content corresponding to an extended portion of the first timeline interval in the second text to the first text.

5. The method of claim 4.

6. updating the target text displayed in the text area, wherein the first adjustment operation includes a length shortening operation; moving text content in the first text corresponding to the shortened portion of the first timeline interval to the second text; or adding a second text to the text area and moving text content in the first text corresponding to the shortened portion of the first timeline interval to the second text; 5. The method of claim 4, comprising:

7. The invalid text is a text in a selected state in the track area, and adjusting a number of timeline segments marked in the edit track clip of the target video in response to a second adjustment operation of the invalid text is in response to a selection operation on a third text in the text area, switching the third text from an unselected state to a selected state, and adding an indication of a timeline interval of the third text in the editing track clip; in response to a deselection operation on a fourth text in the text area, switching the fourth text from a selected state to an unselected state and canceling marking of a timeline section of the fourth text in the editing track clip; 4. The method of claim 3, further comprising at least one of:

8. playing the target video in a video area of ​​the editing screen; When the target video is played to a start position of a second timeline section, adjusting a playback progress of the target video to a first time node, and continuing to play the target video from the first time node; Further comprising: The second timeline interval is a timeline interval of the invalid text, and the first time node is a time node corresponding to an end position of the second timeline interval in the target video. The method according to any one of claims 1 to 7.

9. After playing the target video in the video area of ​​the editing screen, and adjusting a playback progress of the target video to a second time node in response to a trigger operation on a fifth text, and continuing to play the target video from the second time node; The fifth text is displayed in a text area of ​​the editing screen, the fifth text is invalid text or non-invalid text, and the second time node is a time node corresponding to a start position of the fifth text in the target video.

9. The method of claim 8.

10. A time identifier is further displayed in the editing track clip, and after playing the target video in a video area of ​​the editing screen, determining a third timeline interval in which a target position of the landmark is located in response to a position adjustment operation on the landmark; if the third timeline interval is the timeline interval of the invalid text, adjust the playback progress of the target video to a third time node and continue playing the target video from the third time node; If the third timeline interval is not the invalid text timeline interval, adjusting the playback progress of the target video to a fourth time node and continuing to play the target video from the fourth time node; Further comprising: the target position is a position of the target mark in the edit track clip when execution of the positioning operation is completed; the third time node is a time node corresponding to a start position or an end position of the third timeline interval in the target video, or a time node prior to receiving the positioning operation and last playing the target video; The fourth time node is a time node corresponding to the target position in the target video.

9. The method of claim 8.

11. Deleting the video clips within the timeline section of the invalid text in the edit track clip from the target video includes: Delete the video clips in the timeline section of the invalid text in the editing track clip from the target video, and divide the target video into a plurality of independent video clips using the deleted video clips as division points; or deleting the video clips in the timeline section of the invalid text in the edit track clip from the target video, and integrating the target video from which the video clips have been deleted into one video; The method according to any one of claims 1 to 7, characterized in that it comprises:

12. Method according to any one of the preceding claims, characterized in that said edit track clip is an audio track clip of said target video.

13. a text determination module for determining invalid text of a voice text of the target video and a timeline location of the invalid text by voice identifying audio in the target video, the timeline location of the invalid text being for indicating a time of appearance of the voice audio of the invalid text in the target video; a section marking module for displaying an editing track clip of the target video on an editing screen of the target video, and marking a timeline section of the invalid text in the editing track clip according to a timeline position of the invalid text, the timeline section of the invalid text being a time section during which the voice audio of the invalid text appears in the target video; a segment adjustment module for adjusting a timeline segment of the invalid text in an edit track clip of the target video in response to the invalid text adjustment operation; a clip deletion module for deleting a video clip within a timeline section of the invalid text in the edit track clip from the target video in response to a video clip deletion operation of the invalid text; A video editing device comprising:

14. At least one processor; a memory communicatively coupled to the at least one processor; Equipped with a computer program executable by the at least one processor stored in the memory, the computer program being executed by the at least one processor to cause the at least one processor to execute the video editing method according to any one of claims 1 to 7; 1. An electronic device comprising:

15. Stored therein are computer instructions which, when executed by a processor, implement the method for editing video according to any one of claims 1 to 7. A computer-readable storage medium comprising:

16. When executed by a processor, the method for editing a video according to any one of claims 1 to 7 is realized. A program characterized by:

Citation Information

Patent Citations

  • Video processing method, device, electronic equipment and storage medium

    CN113613068A

  • Audio and video editing method and device

    CN114999530A

  • Multimedia processing method and device, electronic equipment and storage medium

    CN115623279A

  • Non-volatile programmable bistable multivibrator particularly for memory redundant circuit with reduced parasite in readout mode

    JP1996007595A

  • Reproduction device

    JP2010266778A