Method and apparatus for editing a shot

By using text-to-speech commands in the subtitle track area when generating audio in video editing software, the frame rate can be automatically adjusted or the scenes can be trimmed, thus solving the problem of mismatch between audio duration and scene duration, achieving consistency in playback duration and optimizing user experience.

CN119767063BActive Publication Date: 2025-12-26SHANGHAI BILIBILI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411997312.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-12-26
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing video editing software cannot match the audio length with the storyboard length when generating audio, resulting in a poor user experience.

Method used

By placing subtitles in the subtitle track area to trigger text-to-speech commands, target audio is generated, and the audio is placed in the audio track area. The frame rate is automatically adjusted or the scenes are trimmed according to the duration of the audio and the scene to match their duration.

Benefits of technology

Ensure consistent playback duration for audio and storyboards to optimize user experience and improve editing efficiency and results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119767063B_ABST
    Figure CN119767063B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of split lens editing method device, computer equipment, computer readable storage medium, computer program product, it is related to video technical field, the method comprises: display split lens editing interface, split lens editing interface includes video track area, subtitle track area and audio track area;In response to the text-to-speech instruction triggered by the subtitle of target split lens based on subtitle track area placement, the target audio corresponding to target split lens is generated, and the target audio is placed in audio track area;In the case where the duration of target audio is greater than the duration of target split lens, the frame rate of target split lens is adjusted, in the case where the duration of target audio is less than the duration of target split lens, the target split lens is cropped to process, so that the playing duration of the target split lens is matched with the duration of the target audio.The technical scheme of the embodiment of the application can make the playing duration of target split lens match with the duration of target audio, ensure the consistency of editing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of bullet screen, and in particular to a split-screen editing method and device, computer equipment, computer readable storage medium, and computer program product. BACKGROUND

[0002] With the development of digital media technology, the creation and consumption of video content has become an important part of modern society. When users create video content, they will create it through video editing software.

[0003] However, the inventors have found that the existing video editing software cannot match the duration of the generated audio with the duration of the split screen when generating audio based on the subtitles corresponding to a certain split screen when creating a complete video from multiple split screens.

[0004] It should be noted that the above content is not necessarily prior art and does not limit the patent protection scope of the present application. SUMMARY

[0005] Embodiments of the present application provide a split-screen editing method, device, computer equipment, computer readable storage medium, and computer program product to solve or alleviate one or more technical problems.

[0006] One aspect of an embodiment of the present application provides a split-screen editing method, the method comprising:

[0007] displaying a split-screen editing interface, the split-screen editing interface comprising a video track area, a subtitle track area, and an audio track area;

[0008] in response to a text-to-speech instruction triggered by subtitles of a target split screen placed based on the subtitle track area, generating a target audio corresponding to the target split screen and placing the target audio in the audio track area;

[0009] in the case where the duration of the target audio is greater than the duration of the target split screen, adjusting the frame rate of the target split screen to match the playback duration of the target split screen with the duration of the target audio;

[0010] in the case where the duration of the target audio is less than the duration of the target split screen, performing a trimming process on the target split screen to match the playback duration of the target split screen with the duration of the target audio.

[0011] Optionally, the response to the text-to-speech instruction triggered by the subtitles of the target split screen placed based on the subtitle track area, the generation of the target audio corresponding to the target split screen, and the placement of the target audio in the audio track area comprises:

[0012] In response to a selection operation on any one of all subtitles corresponding to a target shot based on a subtitle track area placement, all subtitles corresponding to the target shot are obtained;

[0013] Based on all subtitles corresponding to the target shot, a text-to-speech instruction is generated, and the text-to-speech instruction is used to convert all subtitles into a target audio;

[0014] According to the text-to-speech instruction, the target audio corresponding to the target shot is generated, and the target audio is placed in the audio track area.

[0015] Optionally, the text-to-speech instruction triggered by the subtitle of the target shot based on the subtitle track area placement, the target audio corresponding to the target shot is generated, and the target audio is placed in the audio track area, comprising:

[0016] In response to a subtitle selection instruction triggered by a subtitle of a target shot based on a subtitle track area placement, a text editing interface is displayed on the shot editing interface, and the text editing interface includes a text-to-speech control;

[0017] In response to a trigger operation on the text-to-speech control, a text-to-speech instruction is generated, and the text-to-speech instruction is used to convert all subtitles corresponding to the target shot into a target audio;

[0018] Based on the text-to-speech instruction, a target audio corresponding to the target shot is generated, and the target audio is placed in the audio track area.

[0019] Optionally, the text editing interface further includes a tone selection sub-interface, and before the step of generating the text-to-speech instruction in response to the trigger operation on the text-to-speech control, the method further comprises:

[0020] In response to a selection instruction on any one of a plurality of tones displayed in the tone selection sub-interface, the selected tone is used as the tone of the target audio;

[0021] Correspondingly, in response to a trigger operation on the text-to-speech control, the text-to-speech instruction is generated, comprising:

[0022] In response to a trigger operation on the text-to-speech control, the text-to-speech instruction is generated, wherein the text-to-speech instruction carries information of the selected tone.

[0023] Optionally, the method further comprises:

[0024] In response to a trigger operation on the text-to-speech control, an audio determination pop-up window is popped up to determine whether to replace the original audio placed in the audio track area.

[0025] Correspondingly, placing the target audio in the audio track area includes:

[0026] In a case where the user determines, based on the audio, that the pop-up window selection does not select a replacement operation on the original audio, the original audio placed in the audio track area is retained, and a new audio track for placing the target audio is added in the audio track area.

[0027] In a case where the user determines, based on the audio, that the pop-up window selection does not select a replacement operation on the original audio, the original audio placed in the audio track area is retained, and a new audio track for placing the target audio is added in the audio track area.

[0028] Optionally, in a case where the duration of the target audio is greater than the duration of the target shot, adjusting a frame rate of the target shot to match the duration of the target shot with the duration of the target audio includes:

[0029] In a case where the duration of the target audio is greater than the duration of the target shot, determining a variable speed rate according to the duration of the target shot and the duration of the target audio.

[0030] Determining an adjusted frame rate according to the variable speed rate and an original frame rate corresponding to the target shot.

[0031] Taking the adjusted frame rate as a target frame rate corresponding to the target shot.

[0032] Optionally, the method further includes:

[0033] Obtaining a playing time interval of each piece of subtitle corresponding to the target shot.

[0034] Determining an adjustment multiple of the playing time interval according to the variable speed rate.

[0035] Adjusting the playing time interval of each piece of the subtitle based on the adjustment multiple to obtain an adjusted playing time interval.

[0036] Optionally, in a case where the duration of the target audio is less than the duration of the target shot, performing a clipping process on the target shot to match the duration of the target shot with the duration of the target audio includes:

[0037] In a case where the duration of the target audio is less than the duration of the target shot, performing a clipping process on a head segment and / or a tail segment of the target shot to match the duration of the target shot with the duration of the target audio.

[0038] Another aspect of the embodiment of the application provides a shot editing device, and the device includes:

[0039] a display module configured to display a split-screen editing interface, the split-screen editing interface comprising a video track area, a subtitle track area, and an audio track area;

[0040] a generation module configured to generate target audio corresponding to a target split-screen based on a text-to-speech instruction triggered by a subtitle of the target split-screen placed in the subtitle track area, and place the target audio in the audio track area;

[0041] an adjustment module configured to adjust a frame rate of the target split-screen to match a playing time of the target split-screen with a time length of the target audio, if the time length of the target audio is greater than the time length of the target split-screen;

[0042] a clipping module configured to clip the target split-screen to match a playing time of the target split-screen with a time length of the target audio, if the time length of the target audio is less than the time length of the target split-screen.

[0043] Another aspect of the embodiments of the present application provides a computer device, comprising:

[0044] at least one processor; and

[0045] a memory in communication connection with the at least one processor;

[0046] wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0047] Another aspect of the embodiments of the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions are executed by a processor to implement the method as described above.

[0048] Another aspect of the embodiments of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the method as described above.

[0049] The embodiments of the present application can have the following advantages by using the above technical solutions: the target audio corresponding to a target split-screen is generated based on a text-to-speech instruction triggered by a subtitle of the target split-screen placed in the subtitle track area. After the target video is generated, if the time length of the target audio is greater than the time length of the target split-screen, the frame rate of the target split-screen is automatically adjusted; if the time length of the target audio is less than the time length of the target split-screen, the target split-screen is also automatically clipped, so that the playing time of the target split-screen matches the time length of the target audio, ensuring the consistency of editing and optimizing the user experience. Attached Figure Description

[0050] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0051] Figure 1 This diagram schematically illustrates the operating environment of the storyboard editing method according to Embodiment 1 of this application;

[0052] Figure 2 A flowchart illustrating a storyboard editing method according to Embodiment 1 of this application is shown schematically;

[0053] Figure 3 A schematic diagram of the storyboard editing interface is shown.

[0054] Figure 4 This illustration schematically shows an embodiment according to the present application. Figure 2 Detailed flowchart of step S202 in the process;

[0055] Figure 5 The illustration shows a schematic diagram of the audio confirmation pop-up window;

[0056] Figure 6 This illustration schematically shows an embodiment according to the present application. Figure 2 Detailed flowchart of step S204 in the process;

[0057] Figure 7 The diagram illustrates a new flowchart of the storyboard editing method according to Embodiment 1 of this application;

[0058] Figure 8 The diagram illustrates a comparison before and after the adjustment of the subtitle playback time interval;

[0059] Figure 9 A block diagram of a storyboard editing apparatus according to Embodiment 2 of this application is schematically shown; and

[0060] Figure 10 A schematic diagram of the hardware architecture of a computer device according to Embodiment 3 of this application is shown. Detailed Implementation

[0061] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in details below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0062] It should be noted that the description of "first", "second" and the like in the embodiments of the present application is only for the purpose of description and should not be understood as indicating or implying the relative importance of the technical features indicated or implying the number of technical features indicated. Therefore, the features limited by "first", "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the fact that the technical solutions can be realized by those skilled in the art. When the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist and is not within the scope of protection claimed by the present application.

[0063] In the description of the present application, it should be understood that the numerical reference number before the step does not indicate the order of execution of the steps before and after, but is only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as limiting the present application.

[0064] Firstly, the explanation of the terms related to the present application is provided:

[0065] Video track area: an area in the split-screen editing interface for placing video tracks.

[0066] Caption track area: an area in the split-screen editing interface for placing caption tracks.

[0067] Audio track area: an area in the split-screen editing interface for placing audio tracks.

[0068] Video track (Video Track): a layer in the split-screen editing interface for placing split screens. On the timeline, video tracks can be stacked, allowing users to place multiple video clips at the same time point to create complex visual effects, such as picture overlay, transition effects, etc. Video tracks can also place static images, animations or graphic elements, etc.

[0069] Audio track (Audio Track): a layer in the split-screen editing interface for placing audio corresponding to the split screen. This audio track can place music, dialogue, sound effects or narration, etc. On the timeline, there can be multiple audio tracks, allowing users to place different sound elements, such as background music, environmental sound effects and dialogue. Audio tracks can also be stacked, making sound editing more flexible and rich.

[0070] Subtitle Track: A layer in the editing interface specifically designed to place text information corresponding to the shot. It is usually used to place subtitles or titles of the shot. Subtitle tracks can place static text or dynamic text animations.

[0071] Shot: An important technology in the production of movies, advertisements, and animations. Simply put, it refers to dividing continuous frames into several shots and filming or drawing video clips.

[0072] Secondly, in order to facilitate the understanding of the technical solution provided by the person skilled in the art, the related technology is explained as follows:

[0073] With the development of digital media technology, video content creation and consumption has become an important part of modern society. When users create video content, they will use video editing software to achieve the creation.

[0074] However, the inventor found that the existing video editing software cannot make the duration of the generated audio match the duration of the shot when generating audio based on the subtitles corresponding to a certain shot, and the user needs to manually edit the shot again, resulting in poor user experience.

[0075] Therefore, the embodiment of the present application provides a shot editing technical solution. The target audio corresponding to the target shot is generated based on the text-to-speech instruction triggered by the subtitles of the target shot placed in the subtitle track area. After generating the target video, if the duration of the target audio is greater than the duration of the target shot, the frame rate of the target shot will be automatically adjusted; if the duration of the target audio is less than the duration of the target shot, the target shot will also be automatically cropped, so that the playing duration of the target shot matches the duration of the target audio, ensuring the consistency of the editing and optimizing the user experience. See the following.

[0076] Finally, in order to facilitate understanding, an example running environment is provided as follows.

[0077] As shown in Figure 1 , the running environment diagram includes: a service platform 2, and clients (4A, 4B, …, 4N).

[0078] The service platform 2 can connect the clients (4A, 4B, …, 4N) through the network.

[0079] The service platform 2 can be a single server, a server cluster, or a cloud computing service center.

[0080] The service platform 2 can provide various creative materials to the clients, which can be videos, pictures, texts, etc.

[0081] The service platform 2 can be located in a data center at a single site or distributed among different geographic locations (e.g., at multiple sites). The service platform 2 can provide services via a network. The network includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links, such as coaxial cable links, twisted-pair cable links, fiber-optic links, combinations thereof, and / or the like, or wireless links, such as cellular links, satellite links, Wi-Fi links, and / or the like.

[0082] The clients (4A, 4B,..., 4N) can be configured to access content and services of the service platform 2. The clients (4A, 4B,..., 4N) can include electronic devices that carry or are peripheral to a display panel, such as mobile devices, tablet devices, laptop computers, workstations, virtual reality devices, gaming devices, digital streaming devices, vehicle terminals, smart televisions, set-top boxes, and / or the like, and can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, and / or the like. The virtual machines can be loaded by the computing devices based on virtual images and / or other data that define specific software (e.g., operating systems, specialized applications, servers) for emulation. Different virtual machines can be loaded and / or terminated on one or more computing devices as demand for different types of processing services changes.

[0083] The clients (4A, 4B,..., 4N) can be associated with one or more users. A single user can also use one or more of the clients (4A, 4B,..., 4N) to access the service platform 2. The clients (4A, 4B,..., 4N) can travel to various locations and use different networks to access the service platform 2.

[0084] The clients (4A, 4B,..., 4N) can include interfaces. The interfaces can include touchpads, touchscreens, mice, keyboards, or other sensory elements. For example, the input elements can be configured to receive user instructions that can cause the clients (4A, 4B,..., 4N) to perform various operations, such as XXXX, and / or the like.

[0085] Note that the above devices are exemplary, and the number and types of devices can be adjusted in different scenarios or according to different needs.

[0086] Note that the above devices are exemplary, and the number and types of devices can be adjusted in different scenarios or according to different needs.

[0087] The following describes the technical solutions of the present application by taking a client (e.g., 4A) as an execution subject through multiple embodiments. It should be noted that these embodiments can be implemented in various forms, and should not be interpreted as being limited to the embodiments described herein.

[0088] Embodiment One

[0089] Figure 2 An illustrative flowchart of a split-screen editing method according to Embodiment One of the present application is shown.

[0090] As shown in Figure 2 , the split-screen editing method can include steps S200-S206, in which:

[0091] Step S200: Display a split-screen editing interface, the split-screen editing interface including a video track area, a subtitle track area, and an audio track area.

[0092] Step S202: In response to a text-to-speech instruction triggered by a subtitle of a target split screen placed based on the subtitle track area, generate a target audio corresponding to the target split screen, and place the target audio in the audio track area.

[0093] Step S204: If the duration of the target audio is greater than the duration of the target split screen, adjust the frame rate of the target split screen so that the playing duration of the target split screen matches the duration of the target audio.

[0094] Step S206: If the duration of the target audio is less than the duration of the target split screen, perform a trimming process on the target split screen so that the playing duration of the target split screen matches the duration of the target audio.

[0095] The split-screen editing method of the present embodiment generates a target audio corresponding to a target split screen placed based on a subtitle trigger of a text-to-speech instruction of a subtitle of the target split screen. After generating the target video; if the duration of the target audio is greater than the duration of the target split screen, the frame rate of the target split screen is automatically adjusted; if the duration of the target audio is less than the duration of the target split screen, the target split screen is also automatically trimmed, so that the playing duration of the target split screen matches the duration of the target audio, ensuring consistency in editing and optimizing user experience.

[0096] The following describes steps S200-S206 and optional other steps in detail.

[0097] Step S200 Step S200: Display a split-screen editing interface, the split-screen editing interface including a video track area, a subtitle track area, and an audio track area.

[0098] The split-screen editing interface is an interface for editing operations on split screens.

[0099] As an example, referring to Figure 3 , the split shot editing interface can include a video track area A, a subtitle track area B, and an audio track area C. The video track area A can include one or more video tracks, each of which can be used to place a video clip. The subtitle track area B can include one or more subtitle tracks, each of which can be used to place a subtitle. The audio track area C can include one or more audio tracks, each of which can be used to place an audio.

[0100] In actual application, when editing each split shot, one or more subtitles can be configured in the subtitle track for the split shot, and one or more audios can be configured in the audio track for the split shot.

[0101] Step S202 In response to a text-to-speech instruction triggered by a subtitle of a target split shot placed based on the subtitle track area, the target audio corresponding to the target split shot is generated, and the target audio is placed in the audio track area.

[0102] The text-to-speech instruction is used to convert all subtitles corresponding to the target split shot into a corresponding audio. The audio obtained after the conversion is used as the target audio corresponding to the target split shot.

[0103] In some examples, the text-to-speech instruction can be triggered by directly clicking any one of all subtitles corresponding to the target split shot by the user, or can be triggered by directly clicking any one of all subtitles corresponding to the target split shot by the user, and then clicking one or more other controls. The specific generation triggering manner is not limited in the present application.

[0104] In an optional embodiment, step S202 can include: in response to a selection operation of any one of all subtitles corresponding to a target split shot placed based on the subtitle track area, obtaining all subtitles corresponding to the target split shot; generating a text-to-speech instruction based on all subtitles corresponding to the target split shot, the text-to-speech instruction being used to convert the all subtitles into a target audio; and generating the target audio corresponding to the target split shot according to the text-to-speech instruction, and placing the target audio in the audio track area.

[0105] In actual application, when the user selects any one of all subtitles corresponding to the target split shot, all subtitles corresponding to the target split shot are obtained. Then, a text-to-speech instruction is automatically generated based on all the obtained subtitles, so that all subtitles can be converted into a complete target audio based on the text-to-speech instruction in the future, without the need to convert each subtitle into an audio.

[0106] It should be noted that in other embodiments, all the subtitles of the target split shot can also be selected through multiple selection operations, and then the text-to-speech instruction is automatically generated based on all the subtitles.

[0107] In this embodiment, all the subtitles corresponding to the target split shot can be obtained through the selection of one subtitle by the user, so that the selection operation for each subtitle is not required, and the user operation is simplified.

[0108] In an optional embodiment, referring to Figure 4 , step S202 can further include:

[0109] In step S400, in response to a subtitle selection instruction triggered by a subtitle of a target split shot placed based on a subtitle track area, a text editing interface is displayed on the split shot editing interface, and the text editing interface includes a text-to-speech control.

[0110] In some examples, the subtitle selection instruction can be generated by triggering the subtitle selection instruction after the user clicks any one of all the subtitles of the target split shot placed based on the subtitle track area. In other embodiments, in order to avoid mis-triggering, the subtitle selection instruction can also be generated by triggering the subtitle selection instruction after the user clicks any one of all the subtitles of the target split shot placed based on the subtitle track area for a preset duration.

[0111] It should be noted that the user can also trigger the generation of the subtitle selection instruction in other ways, such as triggering the generation of the subtitle selection instruction by sliding the subtitle.

[0112] The text editing interface is an interface for editing the subtitle text, and the text editing interface at least includes a text-to-speech control. The text-to-speech control is used to trigger the generation of a text-to-speech instruction.

[0113] It should be noted that in order to facilitate the user to identify the text-to-speech control, a text description word can be displayed on the text-to-speech control, such as the text description word being “generate audio track”.

[0114] In step S402, in response to a trigger operation on the text-to-speech control, the text-to-speech instruction is generated, and the text-to-speech instruction is used to convert all the subtitles of the target split shot into a target audio.

[0115] In step S404, the target audio corresponding to the target split shot is generated based on the text-to-speech instruction, and the target audio is placed in the audio track area.

[0116] The trigger operation can be a single-click operation, a double-click operation, a sliding operation, etc.

[0117] After detecting the triggering operation on the text-to-speech control, a text-to-speech instruction is generated, so that the target audio corresponding to the target shot can be generated according to the text-to-speech instruction subsequently.

[0118] In this embodiment, after the user selects the subtitle, a text editing interface is displayed on the shot editing interface, so that the user can trigger the generation of a text-to-speech instruction based on the text-to-speech control displayed on the text editing interface, and then the target audio corresponding to the target shot can be generated according to the text-to-speech instruction. Through the above manner, the user can intuitively understand how to trigger the generation of a text-to-speech instruction.

[0119] In an optional embodiment, the text editing interface can further include a tone selection sub-interface. The tone selection sub-interface is used for the user to select a suitable tone to generate the target audio. The tone selection sub-interface can provide a plurality of tones for the user to select. In addition, in order to facilitate the user to select the tone, different categories of tones can be provided in the tone selection sub-interface for the user to select, and each category has one or more tones for the user to select.

[0120] In this embodiment, before the step S402, it further includes:

[0121] In response to a selection instruction of any tone of the plurality of tones displayed in the tone selection sub-interface, the selected tone is used as the tone of the target audio.

[0122] In some examples, the selection instruction of the tone can be triggered by the user clicking any tone of the plurality of tones displayed in the tone selection sub-interface. In other embodiments, in order to avoid mis-triggering, the selection instruction of the tone can also be triggered by the user when the time length of any tone of the plurality of tones displayed in the tone selection sub-interface reaches a preset time length.

[0123] In this embodiment, the tone selection sub-interface is displayed for the user to select the tone of the target audio generated subsequently, so that the user can select a suitable tone according to the actual demand, and improve the diversity of the generated target audio.

[0124] Correspondingly, the step S402 includes:

[0125] In response to the triggering operation on the text-to-speech control, the text-to-speech instruction is generated, wherein the text-to-speech instruction carries information of the selected tone.

[0126] In this embodiment, after the user selects the tone color, the text-to-speech control is triggered again, so that the generated text-to-speech instruction also carries the information of the selected tone color, which facilitates subsequent generation of target audio with the selected tone color based on the text-to-speech instruction.

[0127] In an optional embodiment, the method further comprises:

[0128] In response to the triggering operation of the text-to-speech control, an audio determination pop-up window is popped up to determine whether to replace the original audio placed in the audio track area.

[0129] As an example, the audio determination pop-up window includes Figure 5 As shown, the audio determination pop-up window includes text description information "whether to replace the generated text-to-speech audio" and two controls for the user to determine whether to replace the original audio placed in the audio track area, namely "yes" control and "no" control. When the user triggers the "yes" control, it means that the user selects to replace the original audio based on the audio determination pop-up window. When the user triggers the "no" control, it means that the user selects not to replace the original audio based on the audio determination pop-up window.

[0130] In this embodiment, in response to the triggering operation of the text-to-speech control, the audio determination pop-up window is popped up to determine whether to replace the original audio, which improves flexibility.

[0131] Correspondingly, placing the target audio in the audio track area includes: replacing the original audio placed in the audio track area with the target audio when the user selects to replace the original audio based on the audio determination pop-up window; and keeping the original audio placed in the audio track area and adding a new audio track for placing the target audio in the audio track area when the user selects not to replace the original audio based on the audio determination pop-up window.

[0132] The original audio is the audio originally generated based on the subtitle of the target shot.

[0133] The target audio is the audio newly generated based on the subtitle of the target shot.

[0134] It can be understood that when the original audio is replaced, it is first detected whether the original audio is placed in the audio track area. If the audio track area does not place the original audio, it indicates that the original audio is not generated based on the target shot subtitle in advance. At this time, the target audio can be directly placed in any position in the audio track area. When the audio track area places the original audio, the target audio is replaced and placed in the original audio in the audio track area, that is, the original audio is deleted from the audio track area, and the target audio is placed in the corresponding position.

[0135] In the embodiment, the original audio is directly replaced when the user selects the replacement operation on the original audio, so that the original audio can be avoided to interfere with the target audio. By simultaneously retaining two audios when the user selects not to replace the original audio, a more rich sound effect can be generated.

[0136] In an optional embodiment, after the target audio is generated, the system can provide a real-time preview function for the user to immediately listen to the alignment effect of the generated target audio and the subtitle.

[0137] In an optional embodiment, after the target audio is generated, the user can also use the "undo" function to return to the state before the target audio is generated, to restore or adjust the audio content.

[0138] Step S204 In the case where the duration of the target audio is greater than the duration of the target shot, the frame rate of the target shot is adjusted to match the playing duration of the target shot with the duration of the target audio.

[0139] In the embodiment, in the case where the duration of the target audio is greater than the duration of the target shot, the frame rate of the target shot is adjusted to increase the playing duration of the target shot, so that the playing duration of the target shot can match the duration of the target audio, the playing timeline of the target audio is aligned with the playing timeline of the target shot, and the situation that the video picture has been played but the audio has not been played is avoided, and the user experience is improved.

[0140] In an optional embodiment, referring to Figure 6 , step S204 comprises:

[0141] Step S600, in the case where the duration of the target audio is greater than the duration of the target shot, a variable speed rate is determined according to the duration of the target shot and the duration of the target audio.

[0142] Step S602, an adjusted frame rate is determined according to the variable speed rate and the original frame rate corresponding to the target shot.

[0143] Step S604, taking the adjusted frame rate as the target frame rate corresponding to the target split.

[0144] The variable speed rate = the time length of the target split / the time length of the target audio.

[0145] The original frame rate is the The frame rate originally set by the target split, for example, is 60 frames per second.

[0146] The adjusted frame rate = the variable speed rate * the original frame rate.

[0147] As an example, assuming that the time length of the target split is 4 seconds, the time length of the target audio is 5 seconds, The original frame rate is 60 frames per second, then the variable speed rate = 4 / 5 = 0.8. The adjusted frame rate = 60*0.8 = 48 frames per second.

[0148] In this embodiment, the variable speed rate is determined based on the time length of the target split and the time length of the target audio, and then the adjusted frame rate is determined according to the determined variable speed rate and the original frame rate, so that the playing time length of the target split can be perfectly matched with the time length of the target audio, improving the user experience.

[0149] In some embodiments, in order to avoid too much variable speed, which in turn causes the target split to be not smooth during playing. In this embodiment, a variable speed rate threshold can be set, for example, the variable speed rate threshold is 0.8. When the confirmed variable speed rate is lower than the variable speed rate threshold, only the variable speed rate threshold and the original frame rate corresponding to the target split are used to determine the adjusted frame rate. For the time that is not aligned after variable speed based on this way, a “+” placeholder can be left empty at the tail of the target split for the user to supplement the material, or part of the video segment corresponding to the original video segment of the target split can be taken out for supplementation, so as to finally realize the purpose of matching the playing time length of the target split with the time length of the target audio.

[0150] In an alternative embodiment, referring to Figure 7 , the method further comprises:

[0151] Step S700, obtaining the playing time interval of each caption corresponding to the target split.

[0152] Step S702, determining the adjustment rate of the playing time interval according to the variable speed rate.

[0153] Step S704, adjusting the playing time interval of each caption based on the adjustment rate to obtain an adjusted playing time interval.

[0154] The play time interval is preset by a user, and is used to indicate a time interval of a target shot in which a subtitle needs to be displayed. For example, the play time interval of the subtitle 1 is 40-80 ms, which indicates that if the target shot is played to the 40th ms, the subtitle 1 is displayed on the screen of the target shot, and the display of the subtitle 1 on the screen of the target shot ends when the target shot is played to the 80th ms.

[0155] The adjustment ratio is 1 / variable speed rate.

[0156] The adjusted play time interval is the play time interval multiplied by the adjustment ratio.

[0157] For example, referring to Figure 8 , the length of the target shot is 200 ms, the length of the target audio is 300 ms, the play time interval of the subtitle 1 is 40-80 ms, and the play time interval of the subtitle 2 is 100-140 ms. Based on the length of the target shot and the length of the target audio, the variable speed rate is determined to be 200 / 300. Based on the variable speed rate, the adjustment ratio is 300 / 200=1.5. Therefore, the adjusted play time interval of the subtitle 1 is 60-120 ms, and the adjusted play time interval of the subtitle 2 is 150-210 ms.

[0158] In this embodiment, the play time interval of each subtitle is adjusted, so that the subtitles can be more matched with the target audio, and the user experience is improved.

[0159] Step S206 In the case where the length of the target audio is less than the length of the target shot, the target shot is cropped to match the play time of the target shot with the length of the target audio.

[0160] In this embodiment, in the case where the length of the target audio is less than the length of the target shot, the target shot is cropped, so that the length of the target shot is reduced, the play time of the target shot is reduced, and the play time of the target shot can be matched with the length of the target audio. The play time line of the target audio is aligned with the play time line of the target shot, the situation that the video screen has been played but the audio has not been played is avoided, and the user experience is improved.

[0161] In an optional embodiment, the step S206 includes:

[0162] In the case where the length of the target audio is less than the length of the target shot, the first segment and / or the last segment of the target shot are cropped to match the play time of the target shot with the length of the target audio.

[0163] In this embodiment, the first segment and / or the last segment of the target shot are cropped, so that the picture of the target shot can still have good continuity after the target shot is cropped.

[0164] In other embodiments, segments at other positions of the target shot can also be cropped, such as segments at the middle position of the target shot.

[0165] It should be noted that when the target shot is cropped, the length of the specific segment to be cropped can be determined according to the difference between the playing time of the target shot and the time length of the target audio, for example, the playing time of the target shot is 3 seconds, and the time length of the target audio is 2 seconds, so the length of the segment to be cropped is 1 second.

[0166] Embodiment Two

[0167] Figure 9 A block diagram of a shot editing device according to Embodiment Two of the present application is schematically shown, which can be divided into one or more program modules, one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program module referred to in the embodiments of the present application refers to a series of computer program instruction segments capable of completing a specific function, and the functions of each program module in the embodiments will be specifically described below. As shown in the figure, the shot editing device 900 can include a display module 910, a generation module 920, an adjustment module 930 and a cropping module 940, wherein: Figure 9

[0168] The display module 910 is configured to display a shot editing interface, and the shot editing interface includes a video track area, a subtitle track area and an audio track area.

[0169] The generation module 920 is configured to generate a target audio corresponding to a target shot in response to a text-to-speech instruction triggered by a subtitle of the target shot placed based on the subtitle track area, and place the target audio in the audio track area.

[0170] The adjustment module 930 is configured to adjust the frame rate of the target shot to match the playing time of the target shot with the time length of the target audio in the case that the time length of the target audio is greater than the time length of the target shot.

[0171] The cropping module 940 is configured to crop the target shot to match the playing time of the target shot with the time length of the target audio in the case that the time length of the target audio is less than the time length of the target shot.

[0172] ​As an optional embodiment, in response to a text-to-speech instruction triggered by a subtitle trigger of a target shot based on a subtitle track area placement, generating a target audio corresponding to the target shot, and placing the target audio in the audio track area comprises:

[0173] In response to a selection operation on any one of all subtitles corresponding to the target shot based on the subtitle track area placement, obtaining all subtitles corresponding to the target shot;

[0174] Generating a text-to-speech instruction based on all subtitles corresponding to the target shot, the text-to-speech instruction being used to convert the all subtitles into a target audio;

[0175] Generating the target audio corresponding to the target shot according to the text-to-speech instruction, and placing the target audio in the audio track area.

[0176] As an optional embodiment, in response to a text-to-speech instruction triggered by a subtitle trigger of a target shot based on a subtitle track area placement, generating a target audio corresponding to the target shot, and placing the target audio in the audio track area comprises:

[0177] In response to a subtitle selection instruction triggered by a subtitle trigger of a target shot based on a subtitle track area placement, displaying a text editing interface on the shot editing interface, the text editing interface comprising a text-to-speech control;

[0178] In response to a trigger operation on the text-to-speech control, generating a text-to-speech instruction, the text-to-speech instruction being used to convert all subtitles corresponding to the target shot into a target audio;

[0179] Generating a target audio corresponding to the target shot based on the text-to-speech instruction, and placing the target audio in the audio track area.

[0180] As an optional embodiment, the text editing interface further comprises a tone selection sub-interface, and the apparatus 900 further comprises a selection module configured to:

[0181] In response to a selection instruction on any one of a plurality of tones displayed in the tone selection sub-interface, selecting the tone as a tone of the target audio.

[0182] Correspondingly, in response to a trigger operation on the text-to-speech control, generating the text-to-speech instruction comprises:

[0183] In response to a trigger operation on the text-to-speech control, generating the text-to-speech instruction, wherein the text-to-speech instruction carries information of the selected tone.

[0184] As an optional embodiment, the apparatus 900 further comprises a pop-up module, configured to:

[0185] In response to the triggering operation on the text-to-speech control, a pop-up audio determination pop-up window is popped up to determine whether to replace the original audio placed in the audio track area.

[0186] Correspondingly, placing the target audio in the audio track area comprises:

[0187] In the case that the user selects to replace the original audio based on the audio determination pop-up window, the original audio placed in the audio track area is replaced by the target audio;

[0188] In the case that the user selects not to replace the original audio based on the audio determination pop-up window, the original audio placed in the audio track area is retained, and a new audio track for placing the target audio is added in the audio track area.

[0189] As an optional embodiment, in the case that the duration of the target audio is greater than the duration of the target split shot, adjusting the frame rate of the target split shot to match the playing duration of the target split shot with the duration of the target audio comprises:

[0190] In the case that the duration of the target audio is greater than the duration of the target split shot, determining a variable speed rate according to the duration of the target split shot and the duration of the target audio;

[0191] Determining an adjusted frame rate according to the variable speed rate and the original frame rate corresponding to the target split shot;

[0192] Taking the adjusted frame rate as the target frame rate corresponding to the target split shot.

[0193] As an optional embodiment, the adjusting module 930 is further configured to:

[0194] Obtaining the playing time interval of each piece of subtitle corresponding to the target split shot;

[0195] Determining an adjustment rate of the playing time interval according to the variable speed rate;

[0196] Adjusting the playing time interval of each piece of the subtitle based on the adjustment rate to obtain an adjusted playing time interval.

[0197] As an optional embodiment, in the case that the duration of the target audio is less than the duration of the target split shot, performing a clipping process on the target split shot to match the playing duration of the target split shot with the duration of the target audio comprises:

[0198] In a case where the time length of the target audio is less than the time length of the target shot, the head segment and / or the tail segment of the target shot are cropped to match the time length of the target audio with the playing time length of the target shot.

[0199] Embodiment Three

[0200] Figure 10 A hardware architecture schematic diagram of a computer device 1000 suitable for implementing the shot editing method according to Embodiment Three of the present application is schematically shown. In some embodiments, the computer device 1000 can be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workstation, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 1000 can be a rack-mounted server, a blade server, a tower server, or a rack server (including a standalone server, or a server cluster composed of multiple servers), etc. As shown, the computer device 1000 includes, but is not limited to, a memory 1010, a processor 1020, and a network interface 1030 which are communicatively linked through a system bus. Figure 10

[0201] The memory 1010 includes at least one type of computer readable storage medium, which includes a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 1010 can be an internal storage module of the computer device 1000, such as a hard disk or a memory of the computer device 1000. In other embodiments, the memory 1010 can also be an external storage device of the computer device 1000, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 1000. Of course, the memory 1010 can also include both the internal storage module and the external storage device of the computer device 1000. In this embodiment, the memory 1010 is generally used to store an operating system and various application software installed on the computer device 1000, such as program codes of the shot editing method, etc. In addition, the memory 1010 can also be used to temporarily store various data that have been output or will be output.

[0202] ​The processor 1020 may, in some embodiments, be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chip. The processor 1020 is generally used to control the overall operation of the computer device 1000, such as performing control and processing related to data interaction or communication of the computer device 1000, etc. In the present embodiment, the processor 1020 is configured to execute program code or process data stored in the memory 1010.

[0203] The network interface 1030 may include a wireless network interface or a wired network interface, and is generally used to establish a communication link between the computer device 1000 and other computer devices. For example, the network interface 1030 is configured to connect the computer device 1000 with an external terminal through a network, establish a data transmission channel and a communication link between the computer device 1000 and the external terminal, etc. The network can be an intranet, the Internet, a Global System of Mobile communication (GSM), a Wideband Code Division Multiple Access (WCDMA), a 4G network, a 5G network, Bluetooth, Wi-Fi, or other wireless or wired networks.

[0204] It should be noted that, Figure 10 Only the computer device with components 1010-1030 is shown, but it should be understood that not all of the shown components are required to be implemented, and more or fewer components can be alternatively implemented.

[0205] In the present embodiment, the split-screen editing method stored in the memory 1010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 1020) to complete the embodiments of the present application.

[0206] Embodiment Four

[0207] The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium has a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the split-screen editing method in the embodiments.

[0208] In this embodiment, the computer readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a programmable read only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. In other embodiments, the computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Of course, the computer readable storage medium can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer readable storage medium is usually used to store an operating system and various application software installed on the computer device, for example, program codes of the sub-camera editing method in the embodiments, etc. In addition, the computer readable storage medium can also be used to temporarily store various data that have been output or will be output.

[0209] Embodiment five

[0210] The embodiment of the present application further provides a computer program product, comprising a computer program which is executed by a processor to realize the method in the above embodiment.

[0211] Obviously, those skilled in the art should understand that each module or each step of the above-mentioned embodiment of the present application can be realized by a general computer device, which can be concentrated on a single computer device or distributed on a network composed of multiple computer devices, and optionally, each module or each step can be realized by program codes executable by a computer device, so that each module or each step can be stored in a storage device and executed by a computer device, and in some cases, the steps shown or described can be executed in an order different from that shown here, or each module or each step can be manufactured into an individual integrated circuit module, or multiple modules or steps can be manufactured into a single integrated circuit module. Therefore, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0212] It should be noted that the above is only the preferred embodiment of the present application, and does not limit the patent protection scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings of the present application, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A split-screen editing method characterized by comprising: The method comprises: displaying a split-screen editing interface, the split-screen editing interface comprising a video track area, a subtitle track area and an audio track area; in response to a text-to-speech instruction triggered by a subtitle of a target split-screen placed based on the subtitle track area, generating target audio corresponding to the target split-screen and placing the target audio in the audio track area; in a case where a time length of the target audio is greater than a time length of the target split-screen, adjusting a frame rate of the target split-screen so that a playing time length of the target split-screen matches the time length of the target audio; in a case where the time length of the target audio is less than the time length of the target split-screen, performing a trimming process on the target split-screen so that the playing time length of the target split-screen matches the time length of the target audio; wherein, in the case where the time length of the target audio is greater than the time length of the target split-screen, adjusting the frame rate of the target split-screen so that the playing time length of the target split-screen matches the time length of the target audio comprises: in the case where the time length of the target audio is greater than the time length of the target split-screen, determining a variable speed rate according to the time length of the target split-screen and the time length of the target audio, the variable speed rate = the time length of the target split-screen / the time length of the target audio; determining an adjusted frame rate according to the variable speed rate and an original frame rate corresponding to the target split-screen, the adjusted frame rate = the variable speed rate × the original frame rate; taking the adjusted frame rate as a target frame rate corresponding to the target split-screen.

2. The method of claim 1, wherein, The response to the text-to-speech instruction triggered by the subtitle of the target split-screen placed based on the subtitle track area, the generation of the target audio corresponding to the target split-screen and the placing of the target audio in the audio track area comprises: in response to a selection operation of any one of all subtitles corresponding to the target split-screen placed based on the subtitle track area, obtaining all subtitles corresponding to the target split-screen; generating the text-to-speech instruction based on all subtitles corresponding to the target split-screen, the text-to-speech instruction being used to convert the all subtitles into a target audio; generating the target audio corresponding to the target split-screen according to the text-to-speech instruction and placing the target audio in the audio track area.

3. The method of claim 1, wherein, The response to the text-to-speech instruction triggered by the subtitle of the target split-screen placed based on the subtitle track area, the generation of the target audio corresponding to the target split-screen and the placing of the target audio in the audio track area comprises: in response to a subtitle selection instruction triggered by the subtitle of the target split-screen placed based on the subtitle track area, displaying a text editing interface on the split-screen editing interface, the text editing interface comprising a text-to-speech control; in response to a triggering operation of the text-to-speech control, generating the text-to-speech instruction, the text-to-speech instruction being used to convert all subtitles corresponding to the target split-screen into a target audio; generating the target audio corresponding to the target split-screen based on the text-to-speech instruction and placing the target audio in the audio track area.

4. The method of claim 3, wherein, The text editing interface further comprises a tone selection sub-interface, and before the step of generating the text-to-speech instruction in response to the triggering operation on the text-to-speech control, the method further comprises: in response to a selection instruction for any tone of the plurality of tones displayed in the tone selection sub-interface, selecting the tone as the tone of the target audio; correspondingly, in response to the triggering operation on the text-to-speech control, generating the text-to-speech instruction comprises: in response to the triggering operation on the text-to-speech control, generating the text-to-speech instruction, wherein the text-to-speech instruction carries information of the selected tone.

5. The method according to claim 3 or 4, characterized in that, The method further comprises: in response to the triggering operation on the text-to-speech control, popping up an audio determination pop-up window for the user to determine whether to replace the original audio placed in the audio track area; correspondingly, placing the target audio in the audio track area comprises: if the user selects to replace the original audio based on the audio determination pop-up window, replacing the original audio placed in the audio track area with the target audio; if the user selects not to replace the original audio based on the audio determination pop-up window, keeping the original audio placed in the audio track area, and adding a new audio track in the audio track area for placing the target audio.

6. The method of claim 1, wherein, The method further comprises: obtaining the playing time interval of each subtitle corresponding to the target shot; determining the adjustment ratio of the playing time interval according to the variable speed rate; adjusting the playing time interval of each subtitle based on the adjustment ratio to obtain an adjusted playing time interval.

7. The method according to any one of claims 1 to 4, characterized in that, In the case that the duration of the target audio is less than the duration of the target shot, the method further comprises: in the case that the duration of the target audio is less than the duration of the target shot, performing a trimming process on the head segment and / or tail segment of the target shot to match the playing duration of the target shot with the duration of the target audio.

8. A split edit apparatus, characterized by comprising: The device comprises: a display module configured to display a shot editing interface, wherein the shot editing interface comprises a video track area, a subtitle track area, and an audio track area; a generation module configured to generate a target audio corresponding to a target shot placed in the subtitle track area in response to a text-to-speech instruction triggered by a subtitle of the target shot; an adjustment module configured to, in the case that the duration of the target audio is greater than the duration of the target shot, adjust the frame rate of the target shot to match the playing duration of the target shot with the duration of the target audio; a trimming module configured to, in the case that the duration of the target audio is less than the duration of the target shot, perform a trimming process on the target shot to match the playing duration of the target shot with the duration of the target audio; wherein the adjustment module is further configured to: In a case where the time length of the target audio is greater than the time length of the target shot, a variable speed rate is determined according to the time length of the target shot and the time length of the target audio, the variable speed rate = the time length of the target shot / the time length of the target audio; An adjusted frame rate is determined according to the variable speed rate and an original frame rate corresponding to the target shot, the adjusted frame rate = the variable speed rate × the original frame rate; The adjusted frame rate is taken as a target frame rate corresponding to the target shot.

9. A computer device, characterized by Comprise: At least one processor; And The memory is in communication connection with the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are executed by the processor to implement the method of any one of claims 1 to 7.

11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video editing method and device, electronic equipment and storage medium

    CN116366917A

  • Video and audio synchronization method and device, electronic equipment and medium

    CN117939036A