Video stream mixing method and device, electronic equipment and storage medium

By providing an editing interface and automated dubbing and subtitle generation functions in the video mixing method, the complex and time-consuming operation in the prior art is solved, and a faster and simpler video mixing process is achieved.

CN120128754APending Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311685731.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing video mixing method is complex in operation, and users need to manually import and adjust multimedia streams, which consumes a lot of adjustment time.

Method used

A video mixing method is provided. By displaying an editing interface, users can record dubbing in a multi-track area, and automatically generate dubbing clips and audio recognition text, and generate mixed streaming videos in combination with target videos.

Benefits of technology

The operation complexity of the video mixing method is reduced. Users only need to perform three steps: starting dubbing, ending dubbing and mixing streaming to quickly complete the mixing streaming of adding dubbing and dubbing subtitles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128754A_ABST
    Figure CN120128754A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video stream mixing method and device, electronic equipment and a storage medium. According to the embodiment of the invention, an editing interface comprising a multi-track area can be displayed, the multi-track area comprises a video track and a dubbing subtitle track, and a target video is loaded in the video track; in response to the starting dubbing operation, starting to record audio at the first progress of the dubbing subtitle track; in response to the dubbing ending operation, generating a dubbing fragment and an audio recognition text corresponding to the dubbing fragment; a mixed stream video is generated in response to the stream mixing operation. According to the invention, a rapid flow mixing mode for adding dubbing and corresponding subtitles at the same time is provided, manual configuration is not needed, a user only needs to input the voice, the audio recognition text corresponding to the voice can be directly displayed in the dubbing subtitle track as the subtitles when recording is completed, and the voice and the audio recognition text are mixed into the video during flow mixing. Therefore, according to the scheme, the operation complexity of the video stream mixing method can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to a video mixing method, apparatus, electronic device, and storage medium. Background Art

[0002] Video mixing refers to combining multimedia streams into a single video. For example, combining video, audio, subtitles, dubbing, etc. into a single video. Currently, video mixing requires users to manually import these multimedia streams and adjust the positions where each multimedia stream appears in the video progress. For example, when setting subtitles, users need to manually input the subtitles and adjust the time when the subtitles appear and disappear in the video; for example, when dubbing, users need to manually import the entire dubbing file and adjust the time when the dubbing is played in the video.

[0003] To save the time consumed by users in manually adding subtitles, some existing editing software provides an automatic subtitle function, that is, using speech recognition technology to recognize a prepared speech file, obtaining its text content as subtitles, and then adjusting the time when the subtitles are played in the editing software.

[0004] However, since both dubbing and subtitles need to be played at appropriate positions in the video progress, that is, users need to adjust the time when dubbing and subtitles correctly appear and disappear in the video. Therefore, users need to manually segment the entire speech file or speech text and adjust the positions where each slice appears in the video progress. These steps often need to be repeated multiple times to configure multiple dubbings and multiple subtitles in sequence. Therefore, the operation complexity of the current video mixing method is high and it takes a lot of adjustment time for users. Summary of the Invention

[0005] Embodiments of this application provide a video mixing method, apparatus, electronic device, and storage medium, which can reduce the operation complexity of the video mixing method.

[0006] Embodiments of this application provide a video mixing method, including:

[0007] Display an editing interface, the editing interface includes a multi-track area, the multi-track area includes a video track and a dubbing and subtitle track, the video track and the dubbing and subtitle track have the same linked progress axis, and a target video is loaded in the video track;

[0008] In response to a start dubbing operation, start recording audio at the first progress of the dubbing and subtitle track;

[0009] In response to an end dubbing operation, generate a dubbing segment and an audio recognition text corresponding to the dubbing segment; where the end dubbing operation carries a second progress, a dubbing segment is loaded between the first progress and the second progress of the dubbing and subtitle track, and the audio recognition text is displayed at the dubbing segment.

[0010] Generate a mixed video in response to a mixed flow operation, where the mixed video includes a target video, a dubbed segment, and an audio recognition text; wherein, the dubbed segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbed segment and displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

[0011] The embodiment of the present application also provides a video mixing device, including:

[0012] An interface unit for displaying an editing interface, where the editing interface includes a multi-track area, the multi-track area includes a video track and a dubbed subtitle track, the video track and the dubbed subtitle track have the same linked progress axis, and the target video is loaded in the video track;

[0013] A start dubbing unit for starting to record audio at the first progress of the dubbed subtitle track in response to a start dubbing operation;

[0014] An end dubbing unit for generating a dubbed segment and an audio recognition text corresponding to the dubbed segment in response to an end dubbing operation; wherein, the end dubbing operation carries the second progress, the dubbed segment is loaded between the first progress and the second progress of the dubbed subtitle track, and the audio recognition text is displayed at the dubbed segment;

[0015] A mixing unit for generating a mixed video in response to a mixed flow operation, where the mixed video includes a target video, a dubbed segment, and an audio recognition text; wherein, the dubbed segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbed segment and displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

[0016] In some embodiments, after generating the dubbed segment and the audio recognition text corresponding to the dubbed segment in response to the end dubbing operation, it further includes:

[0017] Update the audio recognition text in response to a text modification operation on the audio recognition text.

[0018] In some embodiments, the text modification operation includes a content modification operation, a style modification operation, a position modification operation, and an animation effect editing operation.

[0019] In some embodiments, after generating the dubbed segment and the audio recognition text corresponding to the dubbed segment in response to the end dubbing operation, it further includes:

[0020] Determine a target special effect in response to an operation of adding a special effect to the dubbed segment in the dubbed subtitle track;

[0021] Perform voice transformation processing on the dubbed segment using the target special effect to obtain a voice-transformed dubbed segment;

[0022] Replace the voice clip in the voice subtitle track with the voice clip after voice conversion.

[0023] In some embodiments, replacing the voice clip in the voice subtitle track with the voice clip after voice conversion includes:

[0024] Play the voice clip after voice conversion for audition;

[0025] In response to confirming the addition operation, replace the voice clip in the voice subtitle track with the voice clip after voice conversion.

[0026] In some embodiments, after generating the voice clip and the audio recognition text corresponding to the voice clip in response to the end of the voice-over operation, it further includes:

[0027] In response to the movement operation on the voice clip in the voice subtitle track, update the first progress and the second progress based on the position of the moved voice clip in the voice subtitle track.

[0028] In some embodiments, the voice subtitle track includes a voice sub-track and a subtitle sub-track. A voice clip is loaded between the first progress and the second progress of the voice sub-track, and an audio recognition text is loaded at the position where the voice clip is loaded between the first progress and the second progress of the subtitle sub-track. In response to the movement operation on the voice clip in the voice subtitle track, updating the first progress and the second progress based on the position of the moved voice clip in the voice subtitle track includes:

[0029] In response to the movement operation on the voice clip in the voice sub-track, update the first progress and the second progress of the voice clip based on the position of the voice clip in the voice sub-track;

[0030] In response to the movement operation on the audio recognition text in the subtitle sub-track, update the position of the audio recognition text in the voice sub-track and update the first progress and the second progress of the audio recognition text.

[0031] In some embodiments, after generating the voice clip and the audio recognition text corresponding to the voice clip in response to the end of the voice-over operation, it further includes:

[0032] In response to the deletion operation on the voice clip in the voice subtitle track, delete the voice clip and the audio recognition text in the voice subtitle track simultaneously.

[0033] In some embodiments, the editing interface further includes a preview play area, and the progress bar includes a preview progress. After generating the voice clip and the audio recognition text corresponding to the voice clip in response to the end of the voice-over operation, it further includes:

[0034] In response to a preview operation, when the preview progress is between a first progress and a second progress, play the video frame of the target video at the preview progress in the preview playback area, and display the audio recognition text on the video frame at the preview progress.

[0035] In some embodiments, playing the video frame of the target video at the preview progress in the preview playback area, and displaying the audio recognition text on the video frame at the preview progress, further includes:

[0036] Play the video audio of the target video at the preview progress, and the dubbed segment.

[0037] In some embodiments, the editing interface further includes a pointer. The first progress is the position pointed to by the pointer on the progress axis. Displaying the editing interface further includes:

[0038] In response to a move operation of the pointer on the progress axis, update the first progress.

[0039] In some embodiments, the video mixing method further includes:

[0040] In response to a subtitle addition operation, add a target subtitle to the dubbed subtitle track. The target subtitle is between a third progress and a fourth progress of the dubbed subtitle track;

[0041] Generate a voice recitation segment corresponding to the target subtitle between the third progress and the fourth progress of the dubbed subtitle track;

[0042] In response to a mixing operation, generate a mixed video. The mixed video includes the target video, the target subtitle, and the voice recitation segment; wherein, the voice recitation segment is used to dub the target video segment, the target subtitle is used as the subtitle of the voice recitation segment and is displayed in the target video segment, and the target video segment is the target video between the third progress and the fourth progress.

[0043] An embodiment of the present application further provides an electronic device, including a memory storing multiple instructions; the processor loads the instructions from the memory to execute the steps in any one of the video mixing methods provided by the embodiments of the present application.

[0044] An embodiment of the present application further provides a computer-readable storage medium, which stores multiple instructions. The instructions are suitable for being loaded by a processor to execute the steps in any one of the video mixing methods provided by the embodiments of the present application.

[0045] Embodiments of the present application can display an editing interface. The editing interface includes a multi-track area. The multi-track area includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track have the same linked progress axis. A target video is loaded in the video track; in response to a start dubbing operation, audio is recorded at the first progress of the dubbing subtitle track; in response to an end dubbing operation, a dubbing segment and an audio recognition text corresponding to the dubbing segment are generated; wherein, the end dubbing operation carries a second progress, the dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment; in response to a mixing operation, a mixed video is generated. The mixed video includes the target video, the dubbing segment, and the audio recognition text; wherein, the dubbing segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbing segment and is displayed in the video segment. The video segment is the target video between the first progress and the second progress.

[0046] Embodiments of the present application provide a mixing method for adding dubbing and dubbing subtitles simultaneously. Through the editing interface, users can record dubbing at an appropriate video progress. After recording, dubbing and dubbing subtitles can be generated immediately at this video progress. That is, users only need to perform three steps: start dubbing, end dubbing, and mixing, to quickly complete the mixing of adding dubbing and dubbing subtitles simultaneously.

[0047] This application does not require: 1. Prepare the dubbing file in advance; 2. Import the dubbing into the editing software; 3. Adjust the position where the dubbing appears and ends in the video progress; 4. Prepare the dubbing subtitle file corresponding to the dubbing in advance; 5. Import the dubbing subtitles into the editing software; 6. Adjust the position where the dubbing subtitles appear and end in the video progress; 7. Repeat the above steps 1-6 multiple times to add multiple prepared dubbings and dubbing subtitles. Therefore, embodiments of the present application provide a faster and more concise mixing method, which can reduce the operation complexity of the current video mixing method. Description of the Drawings

[0048] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0049] Figure 1a It is a schematic diagram of the scenario of the video mixing method provided by the embodiments of the present application;

[0050] Figure 1b It is a schematic flowchart of the video mixing method provided by the embodiments of the present application;

[0051] Figure 1cIt is a schematic diagram of the editing interface of the video mixing method provided by the embodiment of the present application;

[0052] Figure 1d It is a schematic diagram of starting recording of the video mixing method provided by the embodiment of the present application;

[0053] Figure 1e It is a schematic diagram of ending recording of the video mixing method provided by the embodiment of the present application;

[0054] Figure 2a It is a schematic diagram of the home page of the video mixing method provided by the embodiment of the present application;

[0055] Figure 2b It is a schematic diagram of recording of the video mixing method provided by the embodiment of the present application;

[0056] Figure 2c It is a schematic diagram of subtitles of the video mixing method provided by the embodiment of the present application;

[0057] Figure 2d It is a schematic diagram of tips of the video mixing method provided by the embodiment of the present application;

[0058] Figure 2e It is a schematic diagram of text of the video mixing method provided by the embodiment of the present application;

[0059] Figure 2f It is a schematic diagram of text modification of the video mixing method provided by the embodiment of the present application;

[0060] Figure 3 It is a schematic diagram of the structure of the video mixing device provided by the embodiment of the present application;

[0061] Figure 4 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0063] The embodiment of the present application provides a video mixing method, device, electronic device and storage medium.

[0064] Among them, the video mixing device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC); the server can be a single server or a server cluster composed of multiple servers.

[0065] In some embodiments, the video mixing device can also be integrated in multiple electronic devices. For example, the video mixing device can be integrated in multiple servers, and multiple servers are used to implement the video mixing method of this application.

[0066] In some embodiments, the server can also be implemented in the form of a terminal.

[0067] For example, referring to Figure 1a , the electronic device can be a smart terminal, and a client of video editing software is installed on the smart terminal. The smart terminal can communicate with the server, and a server-side of video editing software can be installed on the server.

[0068] In some embodiments, the server involves cloud computing technology and speech and text recognition technology. That is, the server can be a cloud server. The server can use the speech and text recognition technology to convert the recorded audio obtained from the client into an audio recognition text and send the audio recognition text to the client.

[0069] For example, the client can display an editing interface. The editing interface includes a multi-track area. The multi-track area includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track have the same linked progress axis. The target video is loaded in the video track; in response to the start dubbing operation, audio is recorded at the first progress of the dubbing subtitle track; in some embodiments, a real-time calculated audio waveform diagram can be displayed in the dubbing subtitle track when recording audio; in response to the end dubbing operation, a dubbing segment is generated and the dubbing segment is sent to the server, and the server recognizes and returns the audio recognition text corresponding to the dubbing segment; among them, the end dubbing operation carries a second progress, and the dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment; in response to the mixing operation, a mixed video is generated. The mixed video includes the target video, the dubbing segment, and the audio recognition text; among them, the dubbing segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbing segment to be displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

[0070] In some embodiments, the client can split the audio recognition text into multiple sentences, and determine the time range when each sentence appears and disappears in the video progress, so as to add each sentence paragraph at the corresponding video progress according to the time range.

[0071] The following will be described in detail respectively. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.

[0072] Artificial Intelligence (AI) is a technology that uses digital computers to simulate human beings' perception of the environment, acquisition of knowledge, and use of knowledge. This technology can enable machines to have functions similar to human perception, reasoning, and decision-making. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, machine learning / deep learning, autonomous driving, and intelligent transportation.

[0073] Among them, the key technologies of Speech Technology are automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction. Among them, speech has become one of the most promising human-computer interaction methods in the future.

[0074] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0075] In this embodiment, a video mixing method based on natural language processing technology involving artificial intelligence is provided, as Figure 1b shown. The specific process of this video mixing method can be as follows:

[0076] 101. Display an editing interface. The editing interface includes a multi-track area. The multi-track area includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track have the same linked progress axis, and the target video is loaded in the video track.

[0077] The editing interface may include multiple regions and controls, each with a specific function to support users in editing, clipping, and processing multimedia streams. For example, in some embodiments, in addition to the multi-track region, the editing interface may further include a preview playback region, a resource library region, and a special effects panel, etc.

[0078] In some embodiments, referring to Figure 1c , the editing interface may further include a timeline and a pointer. Among them, the timeline is a horizontal line representing the time span of the video, and its scale can be in frames or seconds. The horizontal direction of the timeline represents the passage of time, from left to right indicating the start to the end of the video; each track in the multi-track region is parallel to the timeline and is used to accommodate different types of multimedia streams.

[0079] Among them, the user can locate the video frame at a specific time point of the video by dragging the pointer to move on the timeline or directly clicking on a certain position on the timeline.

[0080] Among them, the first progress refers to the starting position on the timeline where the voiceover recording starts, corresponding to a specific time point of the target video. The first progress can have a preset value or can be adjusted by the user himself.

[0081] In some embodiments, the position pointed to by the pointer on the timeline is the first progress. In some embodiments, the user can control the pointer to move on the timeline, such as dragging the pointer to move on the timeline, controlling the pointer to move on the timeline through shortcut keys, etc.

[0082] In some embodiments, the first progress can be preset to the position corresponding to the first frame of the target video on the timeline.

[0083] In some embodiments, the first progress can be preset to 00:00 on the timeline.

[0084] In some embodiments, the preview playback area can be used to preview the video content being edited. The preview playback area can include a playback control bar, which can contain playback control controls such as play, pause, fast forward, and rewind. Users can use these buttons to control the playback of the video in the preview playback area. The multi-track area can be used to clip, split, adjust the order of, and add elements such as audio and subtitles to the multimedia stream. Among them, the multi-track area can include a video track, an audio track, a subtitle track, a dubbed subtitle track, a special effect track, and so on. The resource library area can display the multimedia files in the resource library, such as videos, audios, images, texts, etc. Users can drag the multimedia files from the resource library area to the multi-track area for editing. Various video effect controls can be displayed in the special effect panel, and users can select and apply them to the multimedia stream in the multi-track area. The pointer can indicate the video progress, that is, the video progress or the editing position. The movement of this playback head on the progress axis represents the current position of video playback or editing.

[0085] In some embodiments, the first progress is the position pointed to by the pointer on the progress axis. The display editing interface further includes:

[0086] In response to a movement operation of the pointer on the progress axis, update the first progress.

[0087] Among them, the movement operation of the pointer on the progress axis can include dragging the pointer to move on the progress axis, long-pressing the space bar to control the pointer to move backward on the progress axis, short-pressing the restore control to control the pointer to fall back to the initial position on the progress axis, and so on.

[0088] In some embodiments, the video picture of the first progress of the target video can be displayed in the preview playback area.

[0089] In some embodiments, the initial position of the first progress can be set at the 0 scale of the progress axis; in some embodiments, the initial position of the first progress can be set at the starting position or the ending position corresponding to the target video on the progress axis.

[0090] 102. In response to a start dubbing operation, start recording audio at the first progress of the dubbed subtitle track.

[0091] Among them, the start dubbing operation can be set by technicians according to actual needs. For example, the start dubbing operation can be triggered by triggering a recording control, a physical button, a somatosensory control, etc., or the editing software can automatically trigger the start dubbing operation when entering the editing interface.

[0092] For example, referring to Figure 1c , the editing interface can include a recording control "Press and hold to speak". When the user presses and holds this recording control, the start dubbing operation is triggered.

[0093] For example, when the user holds down the space bar on the keyboard, the operation of starting voice dubbing is triggered.

[0094] For example, when the user shakes the mobile terminal, the operation of starting voice dubbing is triggered.

[0095] Reference Figure 1c , when the user holds down the recording control, the operation of starting voice dubbing is triggered. At this time, the first progress is 00:00. Therefore, the audio is recorded starting at 00:00 corresponding to the voice dubbing subtitle track.

[0096] In some embodiments, reference Figure 1d , during the recording process, the waveform diagram of the recorded audio can be displayed in real time in the voice dubbing subtitle track.

[0097] In some embodiments, to facilitate abandoning this recording, after starting to record audio at the first progress of the voice dubbing subtitle track in response to the operation of starting voice dubbing, a cancel recording control can also be displayed in the editing interface. In response to the cancel recording operation for this cancel recording control, the recording of the audio can be stopped and the audio data of this recording can be deleted.

[0098] 103. In response to the operation of ending voice dubbing, generate a voice dubbing segment and the corresponding audio recognition text; wherein, the operation of ending voice dubbing carries a second progress, a voice dubbing segment is loaded between the first progress and the second progress of the voice dubbing subtitle track, and the audio recognition text is displayed at the voice dubbing segment.

[0099] Similar to step 102, the operation of ending voice dubbing can be set by the technician according to actual needs. For example, the operation of ending voice dubbing can be triggered by ending the trigger recording control, physical button, body sensing control, etc., or the editing software can automatically trigger the operation of ending voice dubbing when entering the editing interface.

[0100] For example, reference Figure 1e , when the user releases the recording control, the operation of ending voice dubbing is triggered.

[0101] For example, when the user releases the space bar on the keyboard, the operation of ending voice dubbing is triggered.

[0102] For example, when the user stops shaking the mobile terminal, the operation of ending voice dubbing is triggered.

[0103] Reference Figure 1e, in some embodiments, when the user releases the recording control, the voice-over operation is triggered to end, and the moment of this end voice-over operation in the progress bar is determined as the second progress. A voice-over segment and the corresponding audio recognition text of the voice-over segment are generated. In addition to displaying the waveform diagram of the voice-over segment in the voice-over subtitle track, the audio recognition text of the voice-over segment can also be displayed on the waveform diagram. That is, a voice-over - subtitle file is generated in the voice-over subtitle track, and the voice-over - subtitle file includes the voice-over segment and its corresponding audio recognition text.

[0104] In some embodiments, the voice-over recording can be paused midway and then continued at the pause point. That is, a complete voice-over segment can be recorded in multiple sub - voice-over segments in sequence.

[0105] In some embodiments, in order to further improve the real - time display of subtitle text, each time the recording is paused, the audio recognition text of the sub - voice-over segment that has been recorded this time can be displayed in the voice-over subtitle track. Therefore, after step 102 and before step 103, it further includes:

[0106] In response to the voice-over pause operation, a sub - voice-over segment and the corresponding audio recognition sub - text of the sub - voice-over segment are generated. Among them, the voice-over pause operation carries the pause progress. The sub - voice-over segment is loaded between the first progress and the pause progress of the voice-over subtitle track, and the audio recognition sub - text is displayed at the voice-over segment;

[0107] In response to the continue voice-over operation, the audio recording continues at the pause progress of the voice-over subtitle track.

[0108] For example, assume a copywriting "It's sunny today. I'm walking my dog in the park...". The user starts the voice-over at 00:00 on the progress bar, that is, 00:42 is the first progress; when the user finishes reading "It's sunny today" and immediately pauses the voice-over at 00:12 on the progress bar, that is, 00:12 is the pause progress. At this time, the sub - voice-over segment recorded this time and the audio recognition sub - text "It's sunny today" are generated in the voice-over subtitle track corresponding to 00:00 - 00:12 on the progress bar. The user can continue to record the remaining copywriting at 00:12.

[0109] In some embodiments, after step 103, it further includes:

[0110] In response to the deletion operation on the audio recognition text in the voice-over subtitle track, the voice-over segment is simultaneously deleted in the voice-over subtitle track.

[0111] For example, for the voice-over segment or the audio recognition text in the voice-over - subtitle file, if either of them is deleted, both of them will be directly deleted simultaneously.

[0112] In some embodiments, after step 103, it further includes:

[0113] Update the audio recognition text in response to a text modification operation on the text for audio recognition.

[0114] In some embodiments, the text modification operation includes a content modification operation, a style modification operation, a position modification operation, and an animation effect editing operation. For example, the content modification operation may include modifying the text content; the style modification operation may include modifying font styles such as the font, bold, underline, italic, background color, and highlight of the text; the position modification operation may include modifying the display position of the text content in the video frame; and the animation effect editing operation may include adding or deleting animation effects such as fade-in, fade-out, bounce, rotation, and flicker to the text.

[0115] Among them, the text modification operation for the audio recognition text can be implemented through shortcut keys, modification controls, etc. For example, the audio recognition text in the dubbing subtitle track can be double-clicked to perform a text modification on the audio recognition text.

[0116] In some embodiments, after step 103, it further includes:

[0117] Determine the target special effect in response to an operation of adding a special effect to the dubbing segment in the dubbing subtitle track;

[0118] Perform voice transformation processing on the dubbing segment using the target special effect to obtain a voice-transformed dubbing segment;

[0119] Replace the dubbing segment in the dubbing subtitle track with the voice-transformed dubbing segment.

[0120] Among them, the audio special effects may include reverberation, echo, vibrato, surround sound, etc.

[0121] In some embodiments, the server can be used to perform voice transformation processing on the dubbing segment using the target special effect to obtain a voice-transformed dubbing segment. For example, the client sends the dubbing segment and the special effect number of the target special effect to the server, and the server performs voice transformation processing on the dubbing segment based on the algorithm using the target special effect to obtain a voice-transformed dubbing segment, and finally returns the voice-transformed dubbing segment to the client.

[0122] In some embodiments, to facilitate the user to pre-listen to the special effect and replace the dubbing segment in the dubbing subtitle track with the voice-transformed dubbing segment, it includes:

[0123] Play a trial listen to the voice-transformed dubbing segment;

[0124] In response to a confirmation addition operation, replace the dubbing segment in the dubbing subtitle track with the voice-transformed dubbing segment.

[0125] In some embodiments, to facilitate the user to finely adjust the progress of the dubbing appearing in the video, after step 103, it further includes:

[0126] In response to a moving operation on a dubbed segment in a dubbed subtitle track, update a first progress and a second progress based on the position of the moved dubbed segment in the dubbed subtitle track.

[0127] For example, the moving operation may include dragging the dubbed segment to move in the dubbed subtitle track.

[0128] In some embodiments, contrary to the design of the dubbed-subtitle file scheme, the dubbed segments and the audio recognition texts in the dubbed-subtitle file can be split so as to adjust the dubbed segments or the audio recognition texts separately. Therefore, the dubbed subtitle track can be set to include a dubbed sub-track and a subtitle sub-track. A dubbed segment is loaded between the first progress and the second progress of the dubbed sub-track, and the audio recognition text is loaded at the position where the dubbed segment is loaded between the first progress and the second progress of the subtitle sub-track. In response to a moving operation on a dubbed segment in the dubbed subtitle track, update the first progress and the second progress based on the position of the moved dubbed segment in the dubbed subtitle track, including:

[0129] In response to a moving operation on a dubbed segment in the dubbed sub-track, update the first progress and the second progress of the dubbed segment based on the position of the dubbed segment in the dubbed sub-track;

[0130] In response to a moving operation on the audio recognition text in the subtitle sub-track, update the position of the audio recognition text in the dubbed sub-track, and update the first progress and the second progress of the audio recognition text.

[0131] In some embodiments, after step 103, it further includes:

[0132] In response to a deletion operation on a dubbed segment in the dubbed subtitle track, delete the dubbed segment and the audio recognition text in the dubbed subtitle track simultaneously.

[0133] In some embodiments, the editing interface further includes a preview playing area, and the progress axis includes a preview progress. After step 103, it further includes:

[0134] In response to a preview operation, when the preview progress is between the first progress and the second progress, play the video frame of the target video at the preview progress in the preview playing area, and display the audio recognition text on the video frame at the preview progress.

[0135] For example, the user dubbed "The weather is really nice today" at the t 1 ~t 2 of the target video, and at the t 3 ~t 4The voiceover here is "I take my dog for a walk in the park". Therefore, when playing the target video from the t1-th to the t2-th position in the preview play area, in addition to showing the video image, the subtitle "The weather is really nice today" will also be shown. When playing the target video from the t3-th to the t4-th position in the preview play area, in addition to showing the video image, the subtitle "I take my dog for a walk in the park" will also be shown. Among them, t 1 ~t 4 refer to four different time points in the target video.

[0136] In some embodiments, playing the video image of the target video at the preview progress in the preview play area, and showing the audio recognition text on the video image at the preview progress further includes:

[0137] Playing the video audio of the target video at the preview progress, as well as the voiceover segment.

[0138] 104. Generating a mixed video in response to a mixing operation, the mixed video including the target video, the voiceover segment, and the audio recognition text; wherein, the voiceover segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the voiceover segment and shown in the video segment. The video segment is the target video between the first progress and the second progress.

[0139] In some embodiments, in addition to the mixing method of simultaneously mixing the voiceover and the subtitle corresponding to the voiceover by recording the voiceover proposed in steps 101 to 104, a mixing method of simultaneously mixing the subtitle and the voiceover corresponding to the subtitle by configuring the subtitle is also proposed, including:

[0140] Responding to the subtitle addition operation, adding the target subtitle to the voiceover subtitle track, and the target subtitle is located between the third progress and the fourth progress of the voiceover subtitle track;

[0141] Generating a voice recitation segment corresponding to the target subtitle between the third progress and the fourth progress of the voiceover subtitle track;

[0142] Responding to the mixing operation to generate a mixed video, the mixed video including the target video, the target subtitle, and the voice recitation segment; wherein, the voice recitation segment is used to dub the target video segment, and the target subtitle is used as the subtitle of the voice recitation segment and shown in the target video segment. The target video segment is the target video between the third progress and the fourth progress.

[0143] For example, converting the subtitle to a voiceover, that is, the voice recitation segment, through AI voiceover.

[0144] As can be seen from the above, the embodiment of the present application can display an editing interface. The editing interface includes a multi-track area. The multi-track area includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track have the same linked progress axis. The target video is loaded in the video track. In response to the start of the dubbing operation, audio recording starts at the first progress of the dubbing subtitle track. In response to the end of the dubbing operation, a dubbing segment and the corresponding audio recognition text of the dubbing segment are generated. Among them, the end of the dubbing operation carries the second progress. The dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track. The audio recognition text is displayed at the dubbing segment. In response to the mixing operation, a mixed video is generated. The mixed video includes the target video, the dubbing segment, and the audio recognition text. Among them, the dubbing segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbing segment and displayed in the video segment. The video segment is the target video between the first progress and the second progress. Thus, in this solution, by recording the dubbing, the dubbing and the corresponding subtitle of the dubbing are mixed at the same time. The user only needs to perform three steps: start dubbing, end dubbing, and mixing, to quickly complete the mixing of adding the dubbing and the dubbing subtitle at the same time, thereby reducing the operation complexity of the video mixing method.

[0145] The method described in the above embodiment will be further described in detail below.

[0146] In this embodiment, taking the video editing software installed on the smart terminal and combined with the communication with the cloud server as an example, the method of the embodiment of the present application will be described in detail.

[0147] In this embodiment, when the user inputs voice, the subtitle matching the voice content can be directly seen in the dubbing subtitle track when the recording ends.

[0148] Refer to Figure 2a , after the user opens the target video through the video editor, click the "Commentary" control on the home page of the video editor to enter the editing interface.

[0149] Refer to Figure 2b (1). The editing interface includes a video track, a dubbing subtitle track, a preview playback area, a progress axis, a pointer, a recording control "Press and hold to speak", etc. The target video can be automatically imported into the video track, or other videos can be manually imported by the user; Refer to Figure 2b (2) When the user long-presses the recording control "Press and hold to speak", the dubbing can be input in the dubbing subtitle track in real time, and the waveform diagram of the dubbing is displayed in real time. When the user releases the recording control "Press and hold to speak", the recorded dubbing segment and its corresponding subtitle text "The weather is really nice today" are displayed in the dubbing subtitle track; Refer to Figure 2b(3) The user can continue to record the next voiceover. Each time recording, the cancel recording control "Release to Cancel" can be displayed in the editing interface. When the user long - presses the recording control "Press to Speak" and drags it to the cancel recording control "Release to Cancel", the audio recording can be stopped and the audio data recorded this time can be deleted.

[0150] Reference Figure 2c After each recording is completed, the client can send the voiceover recorded this time to the server. The server converts it into the corresponding subtitles and then returns them to the client, and the client displays the subtitles on the waveform diagram corresponding to the voiceover.

[0151] Reference Figure 2d (1) When communication with the server fails, a communication failure prompt "The current network is unavailable and the text cannot be recognized" can be displayed on the waveform diagram of the voiceover recorded this time; Reference Figure 2d (2) When waiting for the server to return the subtitles, a communication failure prompt "Recognizing" can be displayed on the waveform diagram of the voiceover recorded this time.

[0152] Reference Figure 2e (1) The user can click on the subtitles displayed in the selected preview playback area to perform text modification operations on the subtitles. For example, the position can be modified by dragging the subtitles, that is, modifying the position of the subtitles in the video picture; for example, deleting the subtitles; for example, Reference Figure 2e (2) Perform content modification operations on the subtitles; for example, Reference Figure 2e (3) Perform style modification operations on the subtitles.

[0153] Reference Figure 2f The user can add special effect operations to the voiceover segments in the voiceover subtitle track in the client, and determine the target special effect; the client sends the voiceover segment and the special effect number of the target special effect to the server, so that the server can perform voice - changing processing on the voiceover segment using the target special effect to obtain the voice - changed voiceover segment, and then return the voice - changed voiceover segment to the client; the client replaces the voiceover segment in the voiceover subtitle track with the voice - changed voiceover segment.

[0154] In some embodiments, Reference Figure 2f The editing interface may also include an all - apply control "Apply to All". When the user selects the target special effect and clicks this all - apply control "Apply to All", the target special effect can be applied to all the voiceovers in the voiceover subtitle track with one click.

[0155] As can be seen from the above, the embodiments of the present application can reduce the operation complexity of the video mixing method.

[0156] To better implement the above method, an embodiment of the present application further provides a video mixing device, which can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer; the server can be a single server or a server cluster composed of multiple servers.

[0157] For example, in this embodiment, taking the video mixing device being specifically integrated in a smart phone as an example, the method of the embodiment of the present application will be described in detail.

[0158] For example, as Figure 3 shown, the video mixing device may include an interface unit 301, a start voice-over unit 302, an end voice-over unit 303, and a mixing unit 304, as follows:

[0159] (1) Interface unit 301.

[0160] The interface unit 301 is used to display an editing interface, and the editing interface includes a multi-track area. The multi-track area includes a video track and a voice-over subtitle track. The video track and the voice-over subtitle track have the same linked progress axis, and the target video is loaded in the video track.

[0161] In some embodiments, the editing interface further includes a pointer. The first progress is the position pointed to by the pointer on the progress axis. Displaying the editing interface further includes:

[0162] Responding to a moving operation of the pointer on the progress axis, updating the first progress.

[0163] (2) Start voice-over unit 302.

[0164] The start voice-over unit 302 is used to start recording audio at the first progress of the voice-over subtitle track in response to a start voice-over operation.

[0165] (3) End voice-over unit 303.

[0166] The end voice-over unit 303 is used to generate a voice-over segment and an audio recognition text corresponding to the voice-over segment in response to an end voice-over operation; wherein, the end voice-over operation carries a second progress, and the voice-over segment is loaded between the first progress and the second progress of the voice-over subtitle track, and the audio recognition text is displayed at the voice-over segment.

[0167] In some embodiments, after generating the voice-over segment and the audio recognition text corresponding to the voice-over segment in response to the end voice-over operation, it further includes:

[0168] Responding to a text modification operation on the audio recognition text, updating the audio recognition text.

[0169] In some embodiments, the text modification operation includes a content modification operation, a style modification operation, a position modification operation, and an animation effect editing operation.

[0170] In some embodiments, after generating a dubbing segment and the audio recognition text corresponding to the dubbing segment in response to the end of the dubbing operation, the following is further included:

[0171] In response to an operation of adding an effect to the dubbing segment in the dubbing subtitle track, determine the target effect;

[0172] Perform voice transformation processing on the dubbing segment using the target effect to obtain a voice-transformed dubbing segment;

[0173] Replace the dubbing segment in the dubbing subtitle track with the voice-transformed dubbing segment.

[0174] In some embodiments, replacing the dubbing segment in the dubbing subtitle track with the voice-transformed dubbing segment includes:

[0175] Play the voice-transformed dubbing segment for audition;

[0176] In response to a confirmation of the addition operation, replace the dubbing segment in the dubbing subtitle track with the voice-transformed dubbing segment.

[0177] In some embodiments, after generating a dubbing segment and the audio recognition text corresponding to the dubbing segment in response to the end of the dubbing operation, the following is further included:

[0178] In response to an operation of moving the dubbing segment in the dubbing subtitle track, update the first progress and the second progress based on the position of the moved dubbing segment in the dubbing subtitle track.

[0179] In some embodiments, the dubbing subtitle track includes a dubbing sub-track and a subtitle sub-track. A dubbing segment is loaded between the first progress and the second progress of the dubbing sub-track, and an audio recognition text is loaded at the position where the dubbing segment is loaded between the first progress and the second progress of the subtitle sub-track. In response to an operation of moving the dubbing segment in the dubbing subtitle track, updating the first progress and the second progress based on the position of the moved dubbing segment in the dubbing subtitle track includes:

[0180] In response to an operation of moving the dubbing segment in the dubbing sub-track, update the first progress and the second progress of the dubbing segment based on the position of the dubbing segment in the dubbing sub-track;

[0181] In response to an operation of moving the audio recognition text in the subtitle sub-track, update the position of the audio recognition text in the dubbing sub-track and update the first progress and the second progress of the audio recognition text.

[0182] In some embodiments, after generating a dubbed clip and the corresponding audio recognition text of the dubbed clip in response to the end of the dubbing operation, the following is further included:

[0183] In response to a deletion operation on the dubbed clip in the dubbed subtitle track, the dubbed clip and the audio recognition text are simultaneously deleted in the dubbed subtitle track.

[0184] In some embodiments, the editing interface further includes a preview playback area, and the progress bar includes a preview progress. After generating a dubbed clip and the corresponding audio recognition text of the dubbed clip in response to the end of the dubbing operation, the following is further included:

[0185] In response to a preview operation, when the preview progress is between a first progress and a second progress, the video frame of the target video at the preview progress is played in the preview playback area, and the audio recognition text is displayed on the video frame at the preview progress.

[0186] In some embodiments, playing the video frame of the target video at the preview progress in the preview playback area and displaying the audio recognition text on the video frame at the preview progress further includes:

[0187] Playing the video audio of the target video at the preview progress and the dubbed clip.

[0188] (4) Mixing unit 304.

[0189] The mixing unit 304 is configured to generate a mixed video in response to a mixing operation. The mixed video includes a target video, a dubbed clip, and an audio recognition text; wherein, the dubbed clip is used to dub the video clip, and the audio recognition text is used as the subtitle of the dubbed clip and displayed in the video clip. The video clip is the target video between the first progress and the second progress.

[0190] In some embodiments, the video mixing method further includes:

[0191] In response to a subtitle adding operation, a target subtitle is added to the dubbed subtitle track, and the target subtitle is located between a third progress and a fourth progress of the dubbed subtitle track;

[0192] A voice recitation segment corresponding to the target subtitle is generated between the third progress and the fourth progress of the dubbed subtitle track;

[0193] In response to a mixing operation, a mixed video is generated. The mixed video includes a target video, a target subtitle, and a voice recitation segment; wherein, the voice recitation segment is used to dub the target video clip, and the target subtitle is used as the subtitle of the voice recitation segment and displayed in the target video clip. The target video clip is the target video between the third progress and the fourth progress.

[0194] In specific implementation, each of the above units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the method embodiments described above, which will not be elaborated here.

[0195] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0196] As can be seen from the above, the video mixing device in this embodiment displays an editing interface by an interface unit. The editing interface includes a multi-track area. The multi-track area includes a video track and a dubbing subtitle track. The video track and the dubbing subtitle track have the same linked progress axis. A target video is loaded in the video track; the start dubbing unit responds to the start dubbing operation and starts recording audio at the first progress of the dubbing subtitle track; the end dubbing unit responds to the end dubbing operation and generates a dubbing segment and an audio recognition text corresponding to the dubbing segment; wherein, the end dubbing operation carries a second progress, and the dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment; the mixing unit responds to the mixing operation and generates a mixed video. The mixed video includes the target video, the dubbing segment and the audio recognition text; wherein, the dubbing segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbing segment and is displayed in the video segment. The video segment is the target video between the first progress and the second progress.

[0197] Thus, the embodiments of the present application can reduce the operation complexity of the video mixing method.

[0198] The embodiments of the present application also provide an electronic device, which can be a device such as a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0199] In some embodiments, the video mixing device can also be integrated in multiple electronic devices. For example, the video mixing device can be integrated in multiple servers, and multiple servers are used to implement the video mixing method of the present application.

[0200] In this embodiment, the electronic device in this embodiment is taken as a smart phone as an example for detailed description. For example, as Figure 4As shown, it shows a schematic structural diagram of an electronic device involved in an embodiment of the present application. Specifically:

[0201] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, a communication module 405, and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0202] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling the data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby performing an overall detection of the electronic device. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401.

[0203] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the electronic device. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0204] The electronic device also includes a power supply 403 that powers each component. In some embodiments, the power supply 403 may be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.

[0205] The electronic device may further include an input module 404, which may be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0206] The electronic device may further include a communication module 405. In some embodiments, the communication module 405 may include a wireless module. The electronic device may perform short-range wireless transmission through the wireless module of the communication module 405, thereby providing users with wireless broadband Internet access. For example, the communication module 405 may be used to help users send and receive emails, browse web pages, and access streaming media, etc.

[0207] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0208] Display an editing interface, the editing interface includes a multi-track area, the multi-track area includes a video track and a dubbing subtitle track, the video track and the dubbing subtitle track have the same linked progress axis, and a target video is loaded in the video track;

[0209] In response to a start dubbing operation, start recording audio at the first progress of the dubbing subtitle track;

[0210] In response to an end dubbing operation, generate a dubbing segment and an audio recognition text corresponding to the dubbing segment; wherein, the end dubbing operation carries a second progress, a dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment;

[0211] In response to a mixing operation, generate a mixed video, the mixed video includes a target video, a dubbing segment and an audio recognition text; wherein, the dubbing segment is used to dub the video segment, and the audio recognition text is used as the subtitle of the dubbing segment to be displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

[0212] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0213] As can be seen from the above, the embodiments of the present application can reduce the operation complexity of the video mixing method.

[0214] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0215] For this reason, an embodiment of the present application provides a computer-readable storage medium, in which multiple instructions are stored. The instructions can be loaded by a processor to execute the steps in any one of the video mixing methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0216] Display an editing interface, the editing interface includes a multi-track area, the multi-track area includes a video track and a dubbing subtitle track, the video track and the dubbing subtitle track have the same linked progress axis, and a target video is loaded in the video track;

[0217] In response to a start dubbing operation, start recording audio at the first progress of the dubbing subtitle track;

[0218] In response to an end dubbing operation, generate a dubbing segment and an audio recognition text corresponding to the dubbing segment; wherein, the end dubbing operation carries a second progress, the dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment;

[0219] In response to a mixing operation, generate a mixed video, the mixed video includes a target video, a dubbing segment and an audio recognition text; wherein, the dubbing segment is used to dub the video segment, the audio recognition text is used as the subtitle of the dubbing segment and is displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

[0220] Wherein, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc.

[0221] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations in the aspects of video mixing or multimedia editing provided in the above embodiments.

[0222] Since the instructions stored in the storage medium can execute the steps in any of the video mixing methods provided in the embodiments of the present application, the beneficial effects achievable by any of the video mixing methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0223] The above has introduced in detail a video mixing method, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A video editing method, characterized in that, comprising: displaying an editing interface, the editing interface including a multi-track area, the multi-track area including a video track and a dubbing subtitle track, the video track and the dubbing subtitle track having the same linked progress axis, and a target video being loaded in the video track; responding to a start dubbing operation, starting to record audio at a first progress of the dubbing subtitle track; responding to an end dubbing operation, generating a dubbing segment and an audio recognition text corresponding to the dubbing segment; wherein, the end dubbing operation carries a second progress, the dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment; responding to a mixing operation to generate a mixed video, the mixed video including the target video, the dubbing segment and the audio recognition text; wherein, the dubbing segment is used to dub a video segment, the audio recognition text is used as a subtitle of the dubbing segment to be displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

2. The video mixing method according to claim 1, characterized in that, after the step of responding to the end dubbing operation to generate a dubbing segment and an audio recognition text corresponding to the dubbing segment, further comprising: responding to a text modification operation on the audio recognition text, and updating the audio recognition text.

3. The video mixing method according to claim 2, characterized in that, the text modification operation includes a content modification operation, a style modification operation, a position modification operation, and an animation effect editing operation.

4. The video mixing method according to claim 1, characterized in that, after the step of responding to the end dubbing operation to generate a dubbing segment and an audio recognition text corresponding to the dubbing segment, further comprising: responding to an add special effect operation on the dubbing segment in the dubbing subtitle track, and determining a target special effect; performing a voice change process on the dubbing segment by using the target special effect to obtain the voice-changed dubbing segment; replacing the dubbing segment in the dubbing subtitle track with the voice-changed dubbing segment.

5. The video mixing method according to claim 4, characterized in that, the step of replacing the dubbing segment in the dubbing subtitle track with the voice-changed dubbing segment includes: performing a trial play on the voice-changed dubbing segment; responding to a confirmation addition operation, and replacing the dubbing segment in the dubbing subtitle track with the voice-changed dubbing segment.

6. The video mixing method according to claim 1, characterized in that, after the step of responding to the end dubbing operation to generate a dubbing segment and an audio recognition text corresponding to the dubbing segment, further comprising: responding to a moving operation on the dubbing segment in the dubbing subtitle track, and updating the first progress and the second progress based on the position of the moved dubbing segment in the dubbing subtitle track.

7. The video mixing method according to claim 6, characterized in that, The dubbed subtitle track includes a dubbed sub-track and a subtitle sub-track. The dubbed clip is loaded between the first progress and the second progress of the dubbed sub-track. The audio recognition text is loaded at the position where the dubbed clip is loaded between the first progress and the second progress of the subtitle sub-track. Responding to a moving operation on the dubbed clip in the dubbed subtitle track, and updating the first progress and the second progress based on the position of the moved dubbed clip in the dubbed subtitle track includes: Responding to a moving operation on the dubbed clip in the dubbed sub-track, and updating the first progress and the second progress of the dubbed clip based on the position of the dubbed clip in the dubbed sub-track; Responding to a moving operation on the audio recognition text in the subtitle sub-track, updating the position of the audio recognition text in the dubbed sub-track, and updating the first progress and the second progress of the audio recognition text.

8. The video mixing method according to claim 1, wherein, after generating the dubbed clip and the audio recognition text corresponding to the dubbed clip in response to the end of the dubbing operation, it further includes: Responding to a deletion operation on the dubbed clip in the dubbed subtitle track, and simultaneously deleting the dubbed clip and the audio recognition text in the dubbed subtitle track.

9. The video mixing method according to claim 1, wherein, the editing interface further includes a preview playback area, and the progress bar includes a preview progress. After generating the dubbed clip and the audio recognition text corresponding to the dubbed clip in response to the end of the dubbing operation, it further includes: Responding to a preview operation, when the preview progress is between the first progress and the second progress, playing the video frame of the target video at the preview progress in the preview playback area, and displaying the audio recognition text on the video frame at the preview progress.

10. The video mixing method according to claim 9, wherein, playing the video frame of the target video at the preview progress in the preview playback area, and displaying the audio recognition text on the video frame at the preview progress further includes: Playing the video audio of the target video at the preview progress, and the dubbed clip.

11. The video mixing method according to claim 1, wherein, the editing interface further includes a pointer. The first progress is the position pointed to by the pointer in the progress bar. Displaying the editing interface further includes: Responding to a moving operation on the pointer in the progress bar, and updating the first progress.

12. The video mixing method according to claim 1, wherein, the video mixing method further includes: Responding to a subtitle adding operation, adding a target subtitle in the dubbed subtitle track, and the target subtitle is between the third progress and the fourth progress of the dubbed subtitle track; Generating a voice recitation segment corresponding to the target subtitle between the third progress and the fourth progress of the dubbed subtitle track; Generate a mixed video in response to a mixed streaming operation, the mixed video including the target video, the target subtitle, and the voice recitation segment; wherein, the voice recitation segment is used to dub a target video segment, and the target subtitle is used as the subtitle of the voice recitation segment to be displayed in the target video segment, and the target video segment is the target video between the third progress and the fourth progress.

13. A video mixing device Characterized in that Comprising: An interface unit for displaying an editing interface, the editing interface including a multi-track area, the multi-track area including a video track and a dubbing subtitle track, the video track and the dubbing subtitle track having the same linked progress axis, and a target video being loaded in the video track; A start dubbing unit for starting to record audio at the first progress of the dubbing subtitle track in response to a start dubbing operation; An end dubbing unit for generating a dubbing segment and an audio recognition text corresponding to the dubbing segment in response to an end dubbing operation; wherein, the end dubbing operation carries a second progress, the dubbing segment is loaded between the first progress and the second progress of the dubbing subtitle track, and the audio recognition text is displayed at the dubbing segment; A mixing unit for generating a mixed video in response to a mixed streaming operation, the mixed video including the target video, the dubbing segment, and the audio recognition text; wherein, the dubbing segment is used to dub a video segment, and the audio recognition text is used as the subtitle of the dubbing segment to be displayed in the video segment, and the video segment is the target video between the first progress and the second progress.

14. An electronic device Characterized in that Comprising a processor and a memory, the memory storing multiple instructions; the processor loads the instructions from the memory to execute the steps in the video mixing method according to any one of claims 1 to 12.

15. A computer-readable storage medium Characterized in that The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the video mixing method according to any one of claims 1 to 12.

Citation Information

Cited By

  • Video editing method and device, storage medium and program product

    CN120455805A