Video editing method and apparatus, electronic device, and storage medium
By introducing the editing interface of multi-track areas and the function of automatically recording voice clips in the video editing method, the complex and time-consuming operation in the prior art is solved, and a more efficient video editing process is achieved.
Patent Information
- Application Number
- PCT/CN2024/118468
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-09-12
- Publication Date
- 2025-06-12
AI Technical Summary
The existing video editing methods are complex in operation, requiring users to manually adjust the position of dubbing and subtitles in the video time progress, resulting in time-consuming and complexity.
Provides a video editing method, by displaying an editing interface of a multi-track area, including a video track and a dubbing subtitle track, in response to the user's start and end dubbing operations, automatically record voice clips and recognize text, and generate a fused video.
It reduces the operation complexity of the video editing method, reduces the steps of users to adjust the dubbing and subtitle positions in the video time progress, and improves editing efficiency.
Smart Images

Figure CN2024118468_12062025_PF_FP_ABST
Abstract
Description
Video editing method, device, electronic device and storage medium
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 202311685731X, filed on December 7, 2023, entitled “Video Mixing Method, Device, Electronic Device and Storage Medium,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of computers, and in particular to a video editing method, device, electronic device, and storage medium. Background Art
[0004] Video muxing refers to combining multimedia streams into a single video. For example, combining video, audio, subtitles, and dubbing into a single video. Current video muxing requires users to manually import these multimedia streams and adjust the position of each stream within the video's timeline. For example, when setting subtitles, users need to manually enter the subtitles and adjust the time at which they appear and disappear within the video. For example, when adding dubbing, users need to manually import the entire dubbing file and adjust the time at which the dubbing plays within the video.
[0005] In order to save users the time spent on manually adding subtitles, some existing editing software provides an automatic subtitle function, which uses voice recognition technology to identify pre-prepared voice files, obtain their text content as subtitles, and then adjust the playback time of the subtitles in the video in the editing software.
[0006] However, since both voiceovers and subtitles need to be played at the appropriate position in the video's timeline, users need to adjust the timing of their appearance and disappearance. Therefore, users need to manually segment the entire audio file or text and adjust the position of each segment in the video's timeline. These steps often need to be repeated multiple times to configure multiple voiceovers and subtitles. Therefore, current video editing methods are highly complex and require users to spend a lot of time adjusting them.
[0007] Summary of the Invention
[0008] Embodiments of the present application provide a video editing method, apparatus, electronic device, and storage medium, which can reduce the operational complexity of the video editing method.
[0009] The present invention provides a video editing method, which is executed by a computer device and includes:
[0010] Display the editing interface, which includes a multi-track area. The multi-track area includes a video track and a dubbing and subtitle track. The time progress of the video track and the dubbing and subtitle track are aligned, and the video track indicates the time progress of the video;
[0011] In response to a start dubbing operation triggered at the first time progress, starting to record a voice segment at the first time progress of the dubbing subtitle track;
[0012] In response to the end dubbing operation triggered at the second time progress, loading the recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track, and displaying the voice recognition text recognized from the voice segment;
[0013] In response to the video editing completion operation, a fused video including the video, the voice clip and the voice recognition text is generated; wherein, the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
[0014] The present invention also provides a video editing device, including:
[0015] An interface unit is used to display an editing interface, the editing interface includes a multi-track area, the multi-track area includes a video track and a dubbing and subtitle track, the time progress of the video track and the dubbing and subtitle track are aligned, and the video track indicates the time progress of the video;
[0016] A start dubbing unit, configured to start recording a voice segment at the first time progress of the dubbing subtitle track in response to a start dubbing operation triggered at the first time progress;
[0017] a dubbing end unit, configured to, in response to a dubbing end operation triggered at a second time progress, load a recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track, and display a voice recognition text recognized from the voice segment;
[0018] A stream mixing unit is used to generate a fused video including the video, the voice clip and the voice recognition text in response to the completion of the video editing operation; wherein the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
[0019] An embodiment of the present application also provides an electronic device, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps of any video editing method provided in the embodiment of the present application.
[0020] An embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps in any video editing method provided in the embodiment of the present application.
[0021] An embodiment of the present application also provides a computer program product, comprising a plurality of instructions, wherein the instructions are suitable for loading by a processor to execute the steps of any one of the video editing methods provided in the embodiment of the present application.
[0022] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0024] FIG1 is a schematic diagram of a scenario of a video editing method provided in an embodiment of the present application;
[0025] FIG2 is a flow chart of a video editing method according to an embodiment of the present application;
[0026] FIG3 is a schematic diagram of an editing interface of a video editing method provided in an embodiment of the present application;
[0027] FIG4 is a schematic diagram of starting recording of a video editing method provided in an embodiment of the present application;
[0028] FIG5 is a schematic diagram of ending recording of the video editing method provided in an embodiment of the present application;
[0029] FIG6 is a schematic diagram of a homepage of a video editing method provided in an embodiment of the present application;
[0030] FIG7 is a recording diagram of a video editing method provided in an embodiment of the present application;
[0031] FIG8 is a schematic diagram of subtitles of a video editing method provided by an embodiment of the present application;
[0032] FIG9 is a schematic diagram showing a video editing method according to an embodiment of the present application;
[0033] FIG10 is a text diagram of a video editing method provided in an embodiment of the present application;
[0034] FIG11 is a schematic diagram of text modification in a video editing method according to an embodiment of the present application;
[0035] FIG12 is a schematic structural diagram of a video editing device provided in an embodiment of the present application;
[0036] FIG13 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] Embodiments of the present application provide a video editing method, apparatus, electronic device, and storage medium.
[0039] The video editing device can be integrated into an electronic device, such as a terminal or a server. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.
[0040] In some embodiments, the video editing device can also be integrated into multiple electronic devices. For example, the video editing device can be integrated into multiple servers, and the video editing method of the present application can be implemented by multiple servers.
[0041] In some embodiments, the server may also be implemented in the form of a terminal.
[0042] For example, referring to FIG1 , the electronic device may be a smart terminal equipped with a client of video editing software. The smart terminal may communicate with a server, and the server may be equipped with a server of video editing software.
[0043] In some embodiments, the server involves cloud computing technology and speech-to-text recognition technology, that is, the server can be a cloud server, and the server can use speech-to-text recognition technology to convert the recorded audio obtained from the client into speech recognition text, and send the speech recognition text to the client.
[0044] For example, the client can display an editing interface, which includes a multi-track area, and the multi-track area includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track are aligned in time progress, and the video track indicates the time progress of the video; in response to the start dubbing operation triggered at the first time progress, the voice clip is recorded at the first time progress of the dubbing subtitle track; in some embodiments, when recording the voice clip, a real-time calculated audio waveform can be displayed in the dubbing subtitle track; in response to the end dubbing operation triggered at the second time progress, a voice clip is generated and sent to the server, and the server recognizes and returns the voice recognition text recognized from the voice clip; wherein, the end dubbing operation carries the second time progress, and the voice clip is loaded between the first time progress and the second time progress of the dubbing subtitle track, and the voice recognition text is displayed at the voice clip; in response to the video editing completion operation, a fused video including the video, the voice clip and the voice recognition text is generated; wherein the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
[0045] In some embodiments, the client may split the speech recognition text into multiple sentences and determine the time range in which each sentence appears and disappears in the video time progress, thereby adding each sentence paragraph to the corresponding video time progress according to the time range.
[0046] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0047] In this embodiment, a video editing method based on natural language processing technology involving artificial intelligence is provided. As shown in FIG2 , the specific process of the video editing method may be as follows:
[0048] S201. Display an editing interface, which includes a video track and a dubbing subtitle track. The time progress of the video track and the dubbing subtitle track are aligned, and the video track indicates the time progress of the video.
[0049] Among them, time progress alignment means that the time progress of the video track and the dubbing subtitle track are synchronized, facilitating unified editing of the video track and the dubbing subtitle track based on the same time progress. The editing interface can include multiple areas and controls, each area or control has a specific function to support users in editing multimedia streams. The editing interface can include a multi-track area, which includes the video track and the dubbing subtitle track. In some embodiments, in addition to the multi-track area, the editing interface can also include a preview playback area, a resource library area, and a special effects panel, etc.
[0050] In some embodiments, referring to Figure 3, the editing interface may also include a timeline and a pointer, wherein the timeline is a horizontal line representing the time span of the video, and its scale may be frames or seconds. The horizontal direction of the timeline represents the passage of time, and from left to right represents the beginning to the end of the video; each track in the multi-track area is parallel to the timeline and is used to accommodate different types of multimedia streams.
[0051] The user can locate the video screen at a specific time point of the video by dragging the pointer to move on the time progress axis or directly clicking a position on the time progress axis.
[0052] The first time progress refers to the starting position of the dubbing recording on the time progress axis, corresponding to a specific time point of the video. The first time progress can have a preset value or can be adjusted by the user.
[0053] In some embodiments, the position pointed to by the pointer on the time progress axis is the first time progress. In some embodiments, the user can control the pointer to move on the time progress axis, such as dragging the pointer to move on the time progress axis, using shortcut keys to control the pointer to move on the time progress axis, etc.
[0054] In some embodiments, the first time progress may be preset as a position corresponding to the first frame of the video on the time progress axis.
[0055] In some embodiments, the first time progress may be preset to 00:00 on the time progress axis.
[0056] In some embodiments, the preview playback area can be used to preview the video content being edited. The preview playback area may include a playback control bar, which may include playback control controls such as play, pause, fast forward, and fast rewind. Users can use these buttons to control the playback of the video in the preview playback area. The multi-track area can be used to edit, split, adjust the order of multimedia streams, and add elements such as audio and subtitles. The multi-track area may include video tracks, audio tracks, subtitle tracks, dubbing subtitle tracks, special effects tracks, etc. The resource library area can display multimedia files in the resource library, such as video, audio, images, text, etc. Users can drag and drop multimedia files from the resource library area to the multi-track area for editing. The special effects panel can display controls for various video special effects, and users can select and apply them to the multimedia stream in the multi-track area. The pointer can indicate the video time progress of the video. When the pointer or editing position, the movement of the playhead on the time progress axis indicates the current position of video playback or editing.
[0057] In some embodiments, the first time progress is the position pointed to by the pointer in the time progress axis, and displaying the editing interface further includes: updating the first time progress in response to a movement operation of the pointer in the time progress axis.
[0058] Among them, the operations for moving the pointer in the time progress axis can include dragging the pointer to move in the time progress axis, long pressing the space bar to control the pointer to move backward in the time progress axis, short pressing the recovery control to control the pointer to return to the initial position in the time progress axis, and so on.
[0059] In some embodiments, the preview playback area may display a video frame of the first time progress of the video.
[0060] In some embodiments, the initial position of the first time progress can be set at the 0 scale of the time progress axis; in some embodiments, the initial position of the first time progress can be set at the starting position or ending position corresponding to the video in the time progress axis.
[0061] S202: In response to a start dubbing operation triggered at a first time progress, start recording a voice segment at the first time progress of the dubbing subtitle track.
[0062] Among them, the start of dubbing operation can be set by technical personnel according to actual needs. For example, the start of dubbing operation can be triggered by triggering recording controls, physical buttons, somatosensory control, etc., or the editing software can automatically trigger the start of dubbing operation when entering the editing interface.
[0063] For example, referring to FIG3 , the editing interface may include a recording control “Press and hold to speak”, and when the user presses and holds the recording control, the dubbing operation is triggered to start.
[0064] For example, when the user presses the space bar on the keyboard, the dubbing operation is triggered.
[0065] For example, when the user shakes the mobile terminal, the dubbing operation is triggered to start.
[0066] Referring to FIG3 , when the user presses and holds the recording control, the dubbing operation is triggered. At this time, the first time progress is 00:00, so the voice clip is recorded at 00:00 corresponding to the dubbing subtitle track.
[0067] In some embodiments, referring to FIG. 4 , during the recording process, a waveform of the recorded audio may be displayed in real time in the dubbing subtitle track.
[0068] In some embodiments, in order to facilitate the abandonment of the recording, in the step of responding to the start dubbing operation triggered at the first time progress, after starting to record the voice clip at the first time progress of the dubbing subtitle track, a cancel recording control can also be displayed in the editing interface. In response to the cancel recording operation on the cancel recording control, the recording of the voice clip can be stopped and the recorded voice clip can be deleted.
[0069] S203: In response to the end dubbing operation triggered at the second time progress, the recorded voice segment is loaded between the first time progress and the second time progress of the dubbing subtitle track, and the voice recognition text recognized from the voice segment is displayed.
[0070] Similar to step S202, the end of the dubbing operation can be set by the technician according to actual needs. For example, the end of the dubbing operation can be triggered by ending the recording control, physical buttons, somatosensory control, etc., or the editing software can automatically trigger the end of the dubbing operation when entering the editing interface.
[0071] For example, referring to FIG5 , when the user releases the recording control, the dubbing operation is triggered to end.
[0072] For example, when the user releases the space bar on the keyboard, the end dubbing operation is triggered.
[0073] For example, when the user stops shaking the mobile terminal, the dubbing operation is triggered to end.
[0074] Referring to Figure 5, in some embodiments, when the user releases the recording control, the dubbing operation is triggered to end, and the moment when the dubbing operation ends in the time progress axis is determined as the second time progress, and a voice segment and a voice recognition text recognized from the voice segment are generated. In addition to displaying the waveform of the voice segment in the dubbing subtitle track, the voice recognition text of the voice segment can also be displayed on the waveform, that is, a dubbing-subtitle file is generated in the dubbing-subtitle track, and the dubbing-subtitle file includes the voice segment and its corresponding voice recognition text.
[0075] In some embodiments, the dubbing recording can be paused midway and then continued at the pause point, that is, a complete voice segment can be divided into multiple dubbing sub-segments and recorded in sequence.
[0076] In some embodiments, in order to further improve the real-time display of subtitle text, each time the recording is paused, the speech recognition text of the dubbing sub-segment recorded this time can be displayed in the dubbing subtitle track. Therefore, after step S202 and before step S203, the following is further included:
[0077] In response to a pause dubbing operation, a dubbing sub-segment and a speech recognition sub-text corresponding to the dubbing sub-segment are generated, wherein the pause dubbing operation carries a pause time progress, the dubbing sub-segment is loaded between the first time progress and the pause time progress of the dubbing subtitle track, and the speech recognition sub-text is displayed at the speech segment;
[0078] In response to the continuing dubbing operation, the voice segment is continued to be recorded at the paused time progress of the dubbing subtitle track.
[0079] For example, suppose there is a text "The weather is fine today, I am walking the dog in the park...", the user starts dubbing at 00:00 on the time progress axis, that is, 00:42 is the first time progress; at 00:12 on the time progress axis, the user pauses dubbing immediately after reading "The weather is fine today", that is, 00:12 is the pause time progress. At this time, the dubbing sub-segment recorded this time is generated in the dubbing subtitle track corresponding to 00:00 to 00:12 on the time progress axis, and the voice recognition sub-text "The weather is fine today" is displayed in the dubbing sub-segment. The user can continue to record the rest of the text at 00:12.
[0080] In some embodiments, after step S203 , the method further includes: in response to a deletion operation on the speech recognition text in the dubbing subtitle track, simultaneously deleting the speech segment in the dubbing subtitle track.
[0081] For example, if you delete either the voice clip or the voice recognition text in the dubbing-subtitle file, both will be deleted simultaneously.
[0082] In some embodiments, after step S203 , the method further includes: updating the speech recognition text in response to a text modification operation on the speech recognition text.
[0083] In some embodiments, text modification operations include content modification operations, style modification operations, position modification operations, and animation special effects editing operations. For example, content modification operations may include modifying text content; style modification operations may include modifying the font, bolding, underlining, italicization, background color, highlighting, and other font styles of text; position modification operations may include modifying the display position of text content in the video screen; and animation special effects editing operations may include adding or removing animation effects such as fade-in, fade-out, bouncing, rotating, and flashing for text.
[0084] Among them, the text modification operation for the speech recognition text can be implemented through shortcut keys, modification controls, etc. For example, the speech recognition text in the dubbing subtitle track can be double-clicked to modify the text of the speech recognition text.
[0085] In some embodiments, after step S203, it also includes: determining a target special effect in response to an operation of adding special effects to a voice segment in a dubbing subtitle track; performing voice-changing processing on the voice segment using the target special effect to obtain a voice-changed voice segment; and replacing the voice segment in the dubbing subtitle track with the voice-changed voice segment.
[0086] Among them, audio effects can include reverb, echo, vibrato, surround sound, etc.
[0087] In some embodiments, a server may use a target special effect to perform voice-changing processing on a voice segment to obtain a voice-changed voice segment. For example, a client may send the voice segment and the special effect number of the target special effect to the server, and the server may perform voice-changing processing on the voice segment based on an algorithm using the target special effect to obtain a voice-changed voice segment, and finally return the voice-changed voice segment to the client.
[0088] In some embodiments, it is convenient for users to preview special effects and use the voice-changed voice segment to replace the voice segment in the dubbing subtitle track, including: auditioning and playing the voice-changed voice segment; in response to confirming the adding operation, using the voice-changed voice segment to replace the voice segment in the dubbing subtitle track.
[0089] In some embodiments, in order to facilitate users to fine-tune the time progress of dubbing in the video, after step S203, it also includes: in response to the movement operation of the voice segment in the dubbing subtitle track, updating the first time progress and the second time progress based on the position of the moved voice segment in the dubbing subtitle track.
[0090] For example, the move operation may include dragging the voice clip to move it within the subtitle track.
[0091] In some embodiments, contrary to the design of the dubbing-subtitle file solution, the voice clip and speech recognition text in the dubbing-subtitle file can be separated so that the voice clip or speech recognition text can be adjusted separately. Therefore, the dubbing-subtitle track can be set to include a dubbing subtrack and a subtitle subtrack, with the voice clip loaded between the first time progress and the second time progress of the dubbing subtrack, and the speech recognition text loaded between the first time progress and the second time progress of the subtitle subtrack where the voice clip is loaded. In response to the movement operation of the voice clip in the dubbing-subtitle track, based on the position of the moved voice clip in the dubbing-subtitle track, the first time progress and the second time progress are updated, including:
[0092] In response to a move operation on the voice segment in the dubbing sub-track, updating a first time progress and a second time progress of the voice segment based on a position of the voice segment in the dubbing sub-track;
[0093] In response to a move operation on the speech recognition text in the subtitle sub-track, a position of the speech recognition text in the dubbing sub-track is updated, and a first time progress and a second time progress of the speech recognition text are updated.
[0094] In some embodiments, after step S203, the method further includes:
[0095] In response to a deletion operation on the voice segment in the dubbing subtitle track, the voice segment and the speech recognition text are deleted from the dubbing subtitle track at the same time.
[0096] In some embodiments, the editing interface further includes a preview playback area, the video track and the dubbing subtitle track have the same time progress axis, and the time progress axis includes a preview time progress. After step S203, the following steps are further included:
[0097] In response to the preview operation, when the preview time progress is between the first time progress and the second time progress, the video screen of the preview time progress is played in the preview playback area, and the voice recognition text is displayed on the video screen of the preview time progress.
[0098] For example, a user dubs "It's a nice day today" at t1-t2 of a video, and "I went to the park to walk my dog" at t3-t4. Therefore, when the preview area reaches t1-t2 of the video, the video screen will display the subtitle "It's a nice day today." When the preview area reaches t3-t4 of the video, the video screen will display the subtitle "It's a nice day today." T1-t4 refers to four different time points in the video.
[0099] In some embodiments, playing a video screen of a video at a preview time progress in a preview playback area, and displaying voice recognition text on the video screen of the preview time progress, further includes:
[0100] Play the video audio and voice clips of the video at the preview time progress.
[0101] S204. In response to the video editing completion operation, a fused video including the video, the voice clip, and the voice recognition text is generated; wherein the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
[0102] An embodiment of the present application provides a mixing method for simultaneously adding dubbing and dubbing subtitles. The user can record dubbing at the appropriate video progress through the editing interface. After the recording is completed, the dubbing and dubbing subtitles can be generated immediately at the video progress. That is, the user only needs to perform three steps: starting dubbing, ending dubbing and mixing to quickly complete the mixing of adding dubbing and dubbing subtitles at the same time.
[0103] This application eliminates the need to 1. prepare a dubbing file in advance; 2. import the dubbing into editing software; 3. adjust the appearance and end position of the dubbing in the video progress; 4. prepare a dubbing subtitle file corresponding to the dubbing in advance; 5. import the dubbing subtitle into editing software; 6. adjust the appearance and end position of the dubbing subtitle in the video progress; 7. repeat steps 1 to 6 multiple times to add multiple pre-prepared dubbing and dubbing subtitles. Therefore, the embodiments of this application provide a faster and simpler method for mixing streams, which can reduce the operational complexity of current video mixing methods.
[0104] In some embodiments, in addition to the mixing method of simultaneously mixing dubbing and subtitles corresponding to the dubbing by recording dubbing proposed in steps S201 to S204, a mixing method of simultaneously mixing subtitles and dubbing corresponding to the subtitles by configuring subtitles is also proposed, including: in response to a subtitle adding operation, adding a target subtitle to a dubbing subtitle track, the target subtitle being located between a third time progress and a fourth time progress of the dubbing subtitle track; generating a speech recitation segment corresponding to the target subtitle between the third time progress and the fourth time progress of the dubbing subtitle track;
[0105] In response to the completion of the video editing operation, a fused video including the video, the voice clip and the voice recognition text is generated, including: in response to the completion of the video editing operation, a fused video including the video, the voice clip, the voice recognition text, the target subtitles and the voice recitation clip is generated; wherein the voice recitation clip is used to dub the video clip between the third time progress and the fourth time progress in the fused video, and the target subtitles are used as subtitles for the voice recitation clip in the fused video.
[0106] For example, AI dubbing can be used to convert subtitles into dubbing, that is, voice recitation clips.
[0107] As can be seen from the above, an embodiment of the present application can display an editing interface, the editing interface including a multi-track area, the multi-track area including a video track and a dubbing and subtitle track, the video track and the dubbing and subtitle track time progress aligned, and the video track indicating the time progress of the video; in response to a start dubbing operation triggered at a first time progress, a voice clip is recorded at the first time progress of the dubbing and subtitle track; in response to an end dubbing operation triggered at a second time progress, a voice clip and a speech recognition text recognized from the voice clip are generated; wherein the end dubbing operation carries the second time progress, the dubbing and subtitle track has a voice clip loaded between the first time progress and the second time progress, and the speech recognition text is displayed at the voice clip; in response to a video editing completion operation, a fused video including the video, the voice clip, and the speech recognition text is generated; wherein the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the speech recognition text is used as a subtitle for the voice clip in the fused video. Thus, this solution simultaneously mixes the dubbing and the corresponding subtitles by recording the dubbing. The user only needs to perform the three steps of starting dubbing, ending dubbing, and mixing to quickly complete the mixing of adding the dubbing and subtitles simultaneously, thereby reducing the operational complexity of the video editing method.
[0108] The method described in the above embodiment will be further described below.
[0109] In this embodiment, the method of the embodiment of the present application will be described in detail by taking the video editing software installed on the smart terminal and combining it with the communication with the cloud server as an example.
[0110] In this embodiment, the user records the voice and can directly see the subtitles that match the voice content in the dubbing subtitle track when the recording is finished.
[0111] 6 , after the user opens a video through a video editor, he or she clicks the “narration” control on the homepage of the video editor to enter the editing interface.
[0112] Referring to Figure 7 (1), the editing interface includes a video track, a dubbing subtitle track, a preview playback area, a time progress axis, a pointer, a recording control "Press and Talk", etc. The video can be automatically imported into the video track, or the user can manually import other videos; referring to Figure 7 (2), when the user long presses the recording control "Press and Talk", the dubbing can be recorded in the dubbing subtitle track in real time, and the waveform of the dubbing can be displayed in real time. When the user releases the recording control "Press and Talk", the recorded voice clip and its corresponding subtitle text "The weather is really nice today" are displayed in the dubbing subtitle track; referring to Figure 7 (3), the user can continue to record the next dubbing. At each recording, the cancel recording control "Release to Cancel" can be displayed in the editing interface. When the user long presses the recording control "Press and Talk" and drags to the cancel recording control "Release to Cancel", the voice clip can be stopped and the audio data recorded this time can be deleted.
[0113] Referring to Figure 8, each time a recording is completed, the client can send the recorded dubbing to the server, and the server will convert it into corresponding subtitles and return it to the client, and the client will display the subtitles on the waveform diagram of its corresponding dubbing.
[0114] Referring to FIG9(1), when communication with the server is impossible, a communication failure prompt "The current network is unavailable, and the text cannot be recognized" can be displayed on the waveform of the dubbing recorded this time; referring to FIG9(2), when waiting for the server to return the subtitles, a communication failure prompt "Recognizing" can be displayed on the waveform of the dubbing recorded this time.
[0115] Referring to Figure 10(1), the user can modify the text of the subtitle by clicking on the subtitle displayed in the selected preview playback area, such as dragging the subtitle to perform a position modification operation, that is, modifying the position of the subtitle in the video screen; such as deleting the subtitle; such as referring to Figure 10(2), modifying the content of the subtitle; such as referring to Figure 10(3), modifying the style of the subtitle.
[0116] Referring to Figure 11, the user can add special effects operations to the voice clip in the dubbing subtitle track in the client and determine the target special effect; the client sends the voice clip and the special effect number of the target special effect to the server, so that the server uses the target special effect to change the voice of the voice clip, obtain the changed voice clip, and then return the changed voice clip to the client; the client uses the changed voice clip to replace the voice clip in the dubbing subtitle track.
[0117] In some embodiments, referring to Figure 11, the editing interface can also include an all-apply control "Apply to All". After the user selects the target special effect, he clicks the all-apply control "Apply to All" to apply the target special effect to all dubbings in the dubbing subtitle track with one click.
[0118] From the above, it can be seen that the embodiments of the present application can reduce the operational complexity of the video editing method.
[0119] To better implement the above method, embodiments of the present application further provide a video editing device, which can be integrated into an electronic device, such as a terminal, a server, or the like. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, or the like; the server can be a single server or a server cluster consisting of multiple servers.
[0120] For example, in this embodiment, the method of the embodiment of the present application will be described in detail by taking the video editing device specifically integrated into a smart phone as an example.
[0121] For example, as shown in FIG12 , the video editing apparatus may include an interface unit 1201 , a start dubbing unit 1202 , an end dubbing unit 1203 , and a stream mixing unit 1204 , as follows:
[0122] (1) Interface unit 1201.
[0123] The interface unit 1201 is used to display the editing interface, which includes a multi-track area. The multi-track area includes a video track and a dubbing subtitle track. The time progress of the video track and the dubbing subtitle track is aligned, and the video track indicates the time progress of the video.
[0124] In some embodiments, the editing interface further includes a pointer, the first time progress is the position pointed to by the pointer in the time progress axis, and the display of the editing interface further includes:
[0125] In response to a movement operation of the pointer in the time progress axis, the first time progress is updated.
[0126] (2) Start dubbing unit 1202.
[0127] The start dubbing unit 1202 is configured to start recording a voice segment at the first time progress of the dubbing subtitle track in response to a start dubbing operation triggered at the first time progress.
[0128] (3) End dubbing unit 1203.
[0129] The end dubbing unit 1203 is used to load the recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track in response to the end dubbing operation triggered at the second time progress, and display the voice recognition text recognized from the voice segment.
[0130] In some embodiments, in response to the end dubbing operation triggered at the second time progress, after loading the recorded voice clip and displaying the voice recognition text recognized from the voice clip between the first time progress and the second time progress of the dubbing subtitle track, the method further includes:
[0131] In response to a text modification operation on the speech recognition text, the speech recognition text is updated.
[0132] In some embodiments, the text modification operation includes a content modification operation, a style modification operation, a position modification operation, and an animation special effect editing operation.
[0133] In some embodiments, in response to the end dubbing operation triggered at the second time progress, after loading the recorded voice clip and displaying the voice recognition text recognized from the voice clip between the first time progress and the second time progress of the dubbing subtitle track, the method further includes:
[0134] In response to an operation of adding a special effect to a voice segment in a dubbing subtitle track, determining a target special effect;
[0135] Using target special effects to perform voice change processing on the voice clip to obtain the voice-changed voice clip;
[0136] Replace the voice clip in the dubbing subtitle track with the voice-changed voice clip.
[0137] In some embodiments, replacing a voice segment in a dubbing subtitle track with a voice-changed voice segment includes:
[0138] Audition and play the voice clips after voice change;
[0139] In response to confirming the adding operation, the voice segment in the dubbing subtitle track is replaced with the voice-changed voice segment.
[0140] In some embodiments, in response to the end dubbing operation triggered at the second time progress, after loading the recorded voice clip and displaying the voice recognition text recognized from the voice clip between the first time progress and the second time progress of the dubbing subtitle track, the method further includes:
[0141] In response to a move operation on the voice segment in the dubbing subtitle track, the first time progress and the second time progress are updated based on the position of the moved voice segment in the dubbing subtitle track.
[0142] In some embodiments, a dubbing subtitle track includes a dubbing subtrack and a subtitle subtrack, a voice segment is loaded between a first time progress and a second time progress of the dubbing subtrack, and a voice recognition text is loaded between the first time progress and the second time progress of the subtitle subtrack where the voice segment is loaded. In response to a move operation on the voice segment in the dubbing subtitle track, updating the first time progress and the second time progress based on the position of the moved voice segment in the dubbing subtitle track includes:
[0143] In response to a move operation on the voice segment in the dubbing sub-track, updating a first time progress and a second time progress of the voice segment based on a position of the voice segment in the dubbing sub-track;
[0144] In response to a move operation on the speech recognition text in the subtitle sub-track, a position of the speech recognition text in the dubbing sub-track is updated, and a first time progress and a second time progress of the speech recognition text are updated.
[0145] In some embodiments, in response to the end dubbing operation triggered at the second time progress, after loading the recorded voice clip and displaying the voice recognition text recognized from the voice clip between the first time progress and the second time progress of the dubbing subtitle track, the method further includes:
[0146] In response to a deletion operation on the voice segment in the dubbing subtitle track, the voice segment and the speech recognition text are deleted from the dubbing subtitle track at the same time.
[0147] In some embodiments, the editing interface further includes a preview playback area, the video track and the dubbing subtitle track have the same timeline, the timeline includes a preview timeline, and in response to the end dubbing operation triggered at the second timeline, the recorded voice clip is loaded between the first timeline and the second timeline of the dubbing subtitle track, and the voice recognition text recognized from the voice clip is displayed, further comprising:
[0148] In response to the preview operation, when the preview time progress is between the first time progress and the second time progress, the video screen of the preview time progress is played in the preview playback area, and the voice recognition text is displayed on the video screen of the preview time progress.
[0149] In some embodiments, playing a video screen of a video at a preview time progress in a preview playback area, and displaying voice recognition text on the video screen of the preview time progress, further includes:
[0150] Play the video audio and voice clips of the video at the preview time progress.
[0151] (4) Flow mixing unit 1204.
[0152] The mixing unit 1204 is used to generate a fused video including video, voice clips and voice recognition text in response to the completion operation of video editing; wherein the voice clips are used to dub the video clips between the first time progress and the second time progress in the fused video, and the voice recognition text is used as subtitles for the voice clips in the fused video.
[0153] In some embodiments, the video editing method further includes:
[0154] In response to the subtitle adding operation, adding a target subtitle to the dubbing subtitle track, where the target subtitle is located between the third time progress and the fourth time progress of the dubbing subtitle track;
[0155] Generating a speech recitation segment corresponding to the target subtitle between the third time progress and the fourth time progress of the dubbing subtitle track;
[0156] In response to the video editing completion operation, a fused video including the video, the voice clip, and the voice recognition text is generated, including:
[0157] In response to the completion of the video editing operation, a fused video is generated including a video, a voice clip, voice recognition text, a target subtitle and a voice recitation clip; wherein the voice recitation clip is used to dub the video clip between the third time progress and the fourth time progress in the fused video, and the target subtitle is used as a subtitle for the voice recitation clip in the fused video.
[0158] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0159] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0160] As can be seen from the above, the video editing device of this embodiment displays an editing interface by an interface unit, and the editing interface includes a multi-track area, which includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track are aligned in time progress, and the video track indicates the time progress of the video; the start dubbing unit responds to the start dubbing operation triggered at the first time progress, and starts recording a voice segment at the first time progress of the dubbing subtitle track; the end dubbing unit responds to the end dubbing operation triggered at the second time progress, and loads the recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track, and displays the voice recognition text recognized from the voice segment; the mixing unit generates a fused video including video, voice segment and voice recognition text in response to the video editing completion operation; wherein, the voice segment is used to dub the video segment between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice segment in the fused video.
[0161] Therefore, the embodiments of the present application can reduce the operational complexity of the video editing method.
[0162] The present application also provides an electronic device, which may be a terminal, a server, or the like. The terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or the like; the server may be a single server or a server cluster consisting of multiple servers, or the like.
[0163] In some embodiments, the video editing device can also be integrated into multiple electronic devices. For example, the video editing device can be integrated into multiple servers, and the video editing method of the present application can be implemented by multiple servers.
[0164] In this embodiment, the electronic device of this embodiment is a smartphone. For example, as shown in FIG13 , it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application. Specifically:
[0165] The electronic device may include components such as a processor 1301 with one or more processing cores, a memory 1302 with one or more computer-readable storage media, a power supply 1303, an input module 1304, and a communication module 1305. It will be understood by those skilled in the art that the electronic device structure shown in FIG13 does not limit the electronic device and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0166] Processor 1301 is the control center of the electronic device, connecting all parts of the electronic device using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 1302 and accessing data stored in memory 1302, it performs various functions of the electronic device and processes data, thereby performing overall testing of the electronic device. In some embodiments, processor 1301 may include one or more processing cores. In some embodiments, processor 1301 may integrate an application processor and a modem processor, wherein the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1301.
[0167] The memory 1302 can be used to store software programs and modules. The processor 1301 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302. The memory 1302 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1302 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1302 may also include a memory controller to provide the processor 1301 with access to the memory 1302.
[0168] The electronic device also includes a power supply 1303 for supplying power to various components. In some embodiments, the power supply 1303 can be logically connected to the processor 1301 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 1303 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0169] The electronic device may further include an input module 1304, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0170] The electronic device may also include a communication module 1305. In some embodiments, the communication module 1305 may include a wireless module. The electronic device may perform short-range wireless transmission via the wireless module of the communication module 1305, thereby providing the user with wireless broadband Internet access. For example, the communication module 1305 may be used to help the user send and receive emails, browse web pages, and access streaming media.
[0171] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 1301 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 1302 according to the following instructions, and the processor 1301 will run the application programs stored in the memory 1302 to implement various functions as follows:
[0172] Display the editing interface, which includes a multi-track area. The multi-track area includes a video track and a dubbing and subtitle track. The time progress of the video track and the dubbing and subtitle track are aligned, and the video track indicates the time progress of the video;
[0173] In response to a start dubbing operation triggered at the first time progress, starting to record a voice segment at the first time progress of the dubbing subtitle track;
[0174] In response to the end dubbing operation triggered at the second time progress, loading the recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track, and displaying the voice recognition text recognized from the voice segment;
[0175] In response to the video editing completion operation, a fused video including the video, the voice clip and the voice recognition text is generated; wherein, the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
[0176] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0177] From the above, it can be seen that the embodiments of the present application can reduce the operational complexity of the video editing method.
[0178] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0179] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the video editing methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0180] Display the editing interface, which includes a multi-track area. The multi-track area includes a video track and a dubbing and subtitle track. The time progress of the video track and the dubbing and subtitle track are aligned, and the video track indicates the time progress of the video;
[0181] In response to a start dubbing operation triggered at the first time progress, starting to record a voice segment at the first time progress of the dubbing subtitle track;
[0182] In response to the end dubbing operation triggered at the second time progress, loading the recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track, and displaying the voice recognition text recognized from the voice segment;
[0183] In response to the video editing completion operation, a fused video including the video, the voice clip and the voice recognition text is generated; wherein, the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
[0184] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0185] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the video mixing or multimedia editing aspects provided in the above-mentioned embodiments.
[0186] Since the instructions stored in the storage medium can execute the steps in any video editing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any video editing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0187] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0188] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A video editing method, performed by an electronic device, comprising: Display an editing interface, wherein the editing interface includes a video track and a dubbing subtitle track, wherein the time progress of the video track and the dubbing subtitle track are aligned, and the video track indicates the time progress of the video; In response to a start dubbing operation triggered at a first time progress, start recording a voice segment at the first time progress of the dubbing subtitle track; In response to the end dubbing operation triggered at the second time progress, loading the recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track, and displaying the voice recognition text recognized from the voice segment; In response to the video editing completion operation, a fused video including the video, the voice clip and the voice recognition text is generated; wherein the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
2. The video editing method according to claim 1, further comprising: In response to a text modification operation on the speech recognition text, the speech recognition text is updated.
3. The video editing method as described in claim 2, wherein the text modification operation includes a content modification operation, a style modification operation, a position modification operation, and an animation special effect editing operation.
4. The video editing method according to any one of claims 1 to 3, further comprising: In response to an operation of adding special effects to the voice segment in the dubbing subtitle track, determining a target special effect; Using the target special effect to perform voice change processing on the voice segment to obtain the voice-changed voice segment; The voice segment after the voice change is used to replace the voice segment in the dubbing subtitle track.
5. The video editing method according to claim 4, wherein the step of replacing the voice segment in the dubbing subtitle track with the voice segment after the voice change comprises: Auditioning and playing the voice segment after the voice change; In response to confirming the adding operation, the voice segment in the dubbing subtitle track is replaced with the voice segment after the voice change.
6. The video editing method according to any one of claims 1 to 5, further comprising: In response to a move operation on the voice segment in the dubbing subtitle track, the first time progress and the second time progress are updated based on a position of the moved voice segment in the dubbing subtitle track.
7. The video editing method according to claim 6, wherein the dubbing subtitle track comprises a dubbing subtrack and a subtitle subtrack, the voice segment is loaded between the first time progress and the second time progress of the dubbing subtrack, and the voice recognition text is loaded at the position where the voice segment is loaded between the first time progress and the second time progress of the subtitle subtrack; In response to the moving operation of the voice segment in the dubbing subtitle track, updating the first time progress and the second time progress based on the position of the moved voice segment in the dubbing subtitle track includes: In response to a move operation on the voice segment in the dubbing sub-track, updating a first time progress and a second time progress of the voice segment based on a position of the voice segment in the dubbing sub-track; In response to a move operation on the speech recognition text in the subtitle sub-track, the position of the speech recognition text in the dubbing sub-track is updated, and the first time progress and the second time progress of the speech recognition text are updated.
8. The video editing method according to any one of claims 1 to 7, further comprising: In response to a deletion operation on the voice segment in the dubbing subtitle track, the voice segment and the speech recognition text are deleted simultaneously in the dubbing subtitle track.
9. The video editing method according to any one of claims 1 to 8, wherein the editing interface further comprises a preview playback area, the video track and the dubbing subtitle track have the same time progress axis, the time progress axis comprises a preview time progress, and the method further comprises: In response to the preview operation, when the preview time progress is between the first time progress and the second time progress, the video is played in the preview playback area on the video screen of the preview time progress, and the voice recognition text is displayed on the video screen of the preview time progress.
10. The video editing method according to claim 9, further comprising: Play the video and audio of the video at the preview time progress, as well as the voice clip.
11. The video editing method according to any one of claims 1 to 10, wherein the editing interface further comprises a pointer, the first time progress is the position pointed to by the pointer in the time progress axis, and the display editing interface further comprises: In response to a movement operation of the pointer in the time progress axis, the first time progress is updated.
12. The video editing method according to any one of claims 1 to 11, further comprising: In response to a subtitle adding operation, adding a target subtitle in the dubbing subtitle track, the target subtitle being located between a third time progress and a fourth time progress of the dubbing subtitle track; Generating a voice recitation segment corresponding to the target subtitle between the third time progress and the fourth time progress of the dubbing subtitle track; In response to the video editing completion operation, generating a fused video including the video, the voice segment and the voice recognition text, comprises: In response to the video editing completion operation, a fused video including the video, the voice segment, the voice recognition text, the target subtitles and the voice recitation segment is generated; wherein the voice recitation segment is used for dubbing the video segment between the third time progress and the fourth time progress in the fused video, and the target subtitles are used as subtitles of the voice recitation segment in the fused video.
13. A video editing device, comprising: An interface unit, used for displaying an editing interface, wherein the editing interface includes a multi-track area, wherein the multi-track area includes a video track and a dubbing subtitle track, wherein the time progress of the video track is aligned with the time progress of the dubbing subtitle track, and the video track indicates the time progress of the video; A start dubbing unit, configured to start recording a voice segment at the first time progress of the dubbing subtitle track in response to a start dubbing operation triggered at the first time progress; A dubbing end unit, configured to load a recorded voice segment between the first time progress and the second time progress of the dubbing subtitle track in response to a dubbing end operation triggered at a second time progress, and to display a voice recognition text recognized from the voice segment; A stream mixing unit is used to generate a fused video including the video, the voice clip and the voice recognition text in response to a video editing completion operation; wherein the voice clip is used to dub the video clip between the first time progress and the second time progress in the fused video, and the voice recognition text is used as a subtitle for the voice clip in the fused video.
14. An electronic device comprising a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the video editing method according to any one of claims 1 to 12.
15. A computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor to execute the steps of the video editing method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Video dubbing method and device, equipment and storage medium
CN111741231A
Subtitle generation method and device for mobile terminal, equipment and storage medium
CN112653932A
Video recording method, device and equipment and readable storage medium
CN112752047A
Video editing method and device, electronic equipment and storage medium
CN116366917A
Data editing device and computer program for data editing
JP2007316322A
Cited By
Demonstration and understanding-oriented large-model multi-channel interactive joint output method
CN121071030A