Video editing method and device, and storage medium
By automatically acquiring and aligning text segments and audio segments from video drafts during video editing, the cumbersome text-to-audio alignment process in existing technologies is solved, improving the convenience and experience of video editing.
Patent Information
- Application Number
- PCT/CN2025/088203
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-04-10
- Publication Date
- 2025-12-04
AI Technical Summary
In existing technologies, aligning text and audio during video editing is cumbersome and negatively impacts the video editing experience.
By responding to the trigger command of the target text segment in the video draft, the corresponding audio segment is automatically obtained and added to the audio editing track according to the start point of the editing timeline. The length of the text segment is adjusted to align the timeline interval of the text with that of the audio segment.
It enables automatic alignment of text and audio segments during video editing, improving the convenience and experience of video editing.
Smart Images

Figure CN2025088203_04122025_PF_FP_ABST
Abstract
Description
Video editing methods, equipment and storage media
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410704507.9, filed on May 31, 2024, entitled "Video Editing Method, Apparatus and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of computer and network communication technology, and in particular to a video editing method, device and storage medium. Background Technology
[0004] With the development of communication technology and the rise of mobile short videos, the demand for video creation is becoming increasingly strong, and video creators are gradually spreading from professionals to the general public. Summary of the Invention
[0005] This disclosure provides a video editing method, device, and storage medium to bring more convenience to video editing and improve the video editing experience.
[0006] In a first aspect, embodiments of this disclosure provide a video editing method, including:
[0007] In response to a trigger command for a target text segment in a video draft, the first audio segment corresponding to the target text segment is obtained;
[0008] Based on the starting point of the target text fragment on the editing timeline of the video draft, the first reading audio fragment is added to the audio editing track of the video draft.
[0009] The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
[0010] Secondly, embodiments of this disclosure provide a video editing device, including:
[0011] The acquisition unit is used to acquire the first audio segment corresponding to the target text segment in response to a trigger command for a target text segment in a video draft.
[0012] The processing unit is configured to add the first reading audio segment to the audio editing track of the video draft based on the starting point of the target text segment on the editing timeline of the video draft.
[0013] The text adjustment unit is used to adjust the length of the target text segment on the text editing track of the video draft, so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
[0014] Thirdly, embodiments of this disclosure provide an electronic device, including: at least one processor and a memory;
[0015] The memory stores computer-executed instructions;
[0016] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video editing method as described in the first aspect and various possible designs of the first aspect.
[0017] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video editing method described in the first aspect and various possible designs of the first aspect.
[0018] Fifthly, embodiments of this disclosure provide a computer program product, including computer execution instructions, which, when executed by a processor, implement the video editing method described in the first aspect and various possible designs of the first aspect. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 is a scenario example diagram of a video editing method provided in an embodiment of this disclosure;
[0021] Figure 2 is a schematic flowchart of a video editing method provided in an embodiment of this disclosure;
[0022] Figure 3 is a schematic flowchart of a video editing method provided in another embodiment of this disclosure;
[0023] Figure 4 is a structural block diagram of a video editing device provided in an embodiment of this disclosure;
[0024] Figure 5 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0026] In one video editing scenario, it is necessary to add voice-over to text, with the text serving as the corresponding subtitle. Furthermore, the text and audio need to be aligned. However, existing technologies typically require users to perform these operations, making the video editing process cumbersome and impacting the video editing experience.
[0027] To address the aforementioned technical problems, this disclosure provides a video editing method. In response to a trigger command for a target text segment in a video draft, a first audio segment corresponding to the target text segment is obtained. Based on the starting point of the target text segment on the editing timeline of the video draft, the first audio segment is added to the audio editing track of the video draft. The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first audio segment on the audio editing track. In this embodiment, the first audio segment corresponding to the target text segment can be automatically obtained during video editing, and the alignment of the target text segment and the first audio segment can be achieved, bringing more convenience to video editing and improving the video editing experience.
[0028] The video editing method disclosed herein can be applied to any electronic device that can be used for video editing, such as a terminal device. Taking a terminal device as an example, a user can trigger a trigger command on a target text segment in a video draft. According to the trigger command, the user can obtain the first audio segment corresponding to the target text segment. The first audio segment can be generated by the terminal device itself or obtained from the server. For example, in the application scenario shown in Figure 1, the terminal device can send a request to the server to obtain the audio segment of the target text segment. The server can generate the corresponding first audio segment based on the target text segment included in the audio segment acquisition request and return it to the terminal device. After the terminal device obtains the first audio segment corresponding to the target text segment, it can add the first audio segment to the audio editing track (track 2) of the video draft and adjust the segment length of the target text segment on the text editing track (track 1) of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first audio segment on the audio editing track.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0030] The video editing method disclosed herein will be described in detail below with reference to specific embodiments.
[0031] Referring to Figure 2, which is a schematic flowchart of a video editing method according to an embodiment of this disclosure, the method of this embodiment can be applied to any electronic device such as a terminal device. The video editing method includes:
[0032] S201. In response to a trigger command for a target text segment in a video draft, obtain the first audio segment corresponding to the target text segment.
[0033] In this embodiment, during video editing on a terminal device, automatic reading can be triggered for target text segments such as subtitles in the video draft. For example, the target text segment can be automatically read aloud using a preset voice type or a user-selected target voice type, obtaining the first audio clip corresponding to the target text segment. Specific timbre, speech rate, etc., can be pre-set in any preset or target voice type. Any known method can be used to obtain the first audio clip corresponding to the target text segment; no limitation is made here. Specifically, the triggering command for the target text segment can be an automatic reading control provided in the video editing interface. The user can trigger the automatic reading control after selecting the target text segment, thus triggering the automatic reading of the target text segment. Of course, other methods can also be used for triggering; no limitation is made here.
[0034] Optionally, the terminal device can generate the first audio segment for the target text fragment, or other devices can generate the first audio segment for the target text fragment. The terminal device can then obtain the first audio segment from the other devices. For example, the server can generate the first audio segment for the target text fragment. Specifically, the terminal device can send a request to the server to obtain the audio segment for the target text fragment. The request includes the target text fragment. After the server generates the first audio segment, it can send the first audio segment to the terminal device. That is, the terminal device can receive the first audio segment returned by the server.
[0035] Optionally, different voice types can be provided to the user, such as the voice types of different cartoon characters. Furthermore, before obtaining the first audio segment corresponding to the target text fragment, the user's selection instruction on the voice type can be responded to, and the selected target voice type can be determined. Then, when obtaining the first audio segment corresponding to the target text fragment, the first audio segment reading the target text fragment using the target voice type can be obtained. If the first audio segment is generated by the server, the target voice type can be sent to the server after determining it. Optionally, the target voice type can be included in the above-mentioned request to obtain the audio segment reading of the target text fragment.
[0036] Optionally, the user can select one or more target speech types. Correspondingly, when acquiring the first audio segment, one or more of the first audio segments that read the target text segment using one or more target speech types can be acquired.
[0037] If a user selects only one target speech type, they can obtain one first audio segment read aloud using that target speech type. Of course, they can also obtain multiple first audio segments read aloud using that target speech type, and these multiple first audio segments are the same. If a user selects multiple target speech types, they can obtain one or more first audio segments read aloud using that target speech type for each target speech type.
[0038] S202. Based on the starting point of the target text segment on the editing timeline of the video draft, the first reading audio segment is added to the audio editing track of the video draft.
[0039] In this embodiment, the target text fragment is added to the text editing track of the video draft. The first reading audio can be added to the audio editing track of the video draft according to the starting point of the target text fragment on the editing timeline (that is, the starting point corresponding to the target text fragment on the editing timeline), so as to perform preliminary alignment between the first reading audio and the target text fragment. Specifically, the starting point of the first reading audio fragment on the editing timeline can be aligned with the starting point of the target text fragment on the editing timeline, that is, the starting time of the first reading audio fragment is aligned with the starting time of the target text fragment.
[0040] Optionally, if one or more first audio segments are obtained, the obtained one or more first audio segments can be added to different audio editing tracks of the video draft. If only one first audio segment is obtained, it is added to one audio editing track. If multiple first audio segments are obtained, each first audio segment is added to one audio editing track. The starting point of each first audio segment on the editing timeline is aligned with the starting point of the target text segment on the editing timeline.
[0041] S203. Adjust the length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
[0042] In this embodiment, after the first audio clip is added to the audio editing track of the video draft, the length of the target text clip in the video draft is automatically adjusted based on the timeline interval occupied by the first audio clip on the audio editing track, so that the timeline interval occupied by the target text on the text editing track is aligned with the timeline interval occupied by the first audio clip on the audio editing track. That is, the target text clip and the first audio clip are aligned not only at the start point of the editing timeline, but also at the end point of the editing timeline.
[0043] Optionally, when one or more first audio segments are acquired and added to different audio editing tracks of the video draft, since the timeline interval occupied by the target text segment on the text editing track is intelligently aligned with the timeline interval occupied by a first audio segment on the audio editing track, the latest acquired first audio segment can be selected. This means the length of the target text segment in the video draft can be automatically adjusted to align the timeline interval occupied by the target text on the text editing track with the timeline interval occupied by the latest acquired first audio segment on the audio editing track. Alternatively, the user can specify the first audio segment to be aligned. For example, after the user triggers a specified first audio segment on any audio editing track of the video draft, the length of the target text segment in the video draft can be automatically adjusted to align the timeline interval occupied by the target text on the text editing track with the timeline interval occupied by the specified first audio segment on the audio editing track.
[0044] In addition, if the latest obtained first audio segment is deleted, the second-to-last obtained first audio segment can be selected. That is, in response to the deletion command of the latest obtained first audio segment, the length of the target text segment in the video draft is adjusted so that the timeline interval occupied by the target text on the text editing track is aligned with the timeline interval occupied by the second-to-last obtained first audio segment on the audio editing track.
[0045] The video editing method provided in this embodiment obtains a first audio segment corresponding to the target text segment in response to a trigger command for a target text segment in a video draft; adds the first audio segment to the audio editing track of the video draft according to the starting point of the target text segment on the editing timeline of the video draft; and adjusts the segment length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first audio segment on the audio editing track. This embodiment can automatically obtain the first audio segment corresponding to the target text segment during video editing and achieve alignment between the target text segment and the first audio segment, bringing more convenience to video editing and improving the video editing experience.
[0046] Based on any of the above embodiments, the video editing method can be applied to any video editing scenario. Furthermore, the function of automatically aligning text and audio clips (i.e., whether to automatically execute the video editing method of the above embodiments) can be preset or set by the user within the video editing scenario. For example, in a video editing scenario that generates a video based on images and text, the function of automatically aligning text and audio clips can be preset to be enabled, and the user can freely choose whether to disable this function. In a normal video editing scenario, the function of automatically aligning text and audio clips can be preset to be disabled, and the user can freely choose whether to enable this function. Moreover, after the user selects whether to enable or disable the function of automatically aligning text and audio clips, the previous selection can be remembered by default. When text is added to the video draft again, the video editing method can be automatically executed based on the user's previous selection, or the video editing method can be not executed.
[0047] Based on any of the above embodiments, during the video editing process, the target text segment that has been aligned with the first audio segment can also be modified, as shown in Figure 3. The specific process is as follows:
[0048] S301. In response to the modification instruction for the target text segment, update the target text segment in the video draft, and reacquire the second reading audio segment corresponding to the updated target text segment;
[0049] S302. Replace the first reading audio segment in the audio editing track with the second reading audio segment;
[0050] S303. Adjust the length of the updated target text segment on the text editing track of the video draft so that the timeline interval occupied by the updated target text segment on the text editing track is aligned with the timeline interval occupied by the second reading audio segment on the audio editing track.
[0051] In this embodiment, the user can trigger a modification command for the target text segment, such as adding, deleting, or modifying the content of the target text segment on the text editing track, thereby updating the target text segment on the text editing track. Since the content of the target text segment has changed, it is necessary to re-acquire the second audio segment corresponding to the updated target text segment. The process of acquiring the second audio segment can be similar to the process of acquiring the first audio segment in the above embodiment, and will not be repeated here. The modification command for the target text segment can be any feasible operation command such as the user's confirmation operation on the modified target text segment, or the mouse defocusing from the target text segment (e.g., moving away from the target text segment). Furthermore, after obtaining the second audio segment, the first audio segment in the audio editing track can be replaced with the second audio segment. The specific replacement process can be to first delete the first audio segment and then add the second audio segment, or it can be any feasible replacement process. After the replacement is completed, the segment length of the updated target text segment on the text editing track of the video draft can be adjusted so that the timeline interval occupied by the updated target text segment on the text editing track is aligned with the timeline interval occupied by the second audio segment on the audio editing track. The process is similar to the process of aligning the target text segment with the first audio segment before the update in the above embodiment, and will not be described again here.
[0052] It should be noted that S301-S303 above can be executed automatically when the function of automatically aligning text with audio clips is enabled. Of course, if the function of automatically aligning text with audio clips is disabled, it can be executed after the function is enabled, or it can be executed after being triggered by the user.
[0053] Furthermore, in the above embodiments, if one or more first audio segments are acquired and added to different audio editing tracks of the video draft, and the target text segment is updated, one or more second audio segments of the updated target text segment can be acquired again and used to replace one or more first audio segments in the audio editing track, so that each first audio segment is updated.
[0054] Based on the above embodiments, the modification of the target text segment may further include the segmentation of the target text segment, that is, the segmentation of the target text segment into two or more sub-text segments. In other words, the above-mentioned modification instruction for the target text segment can be a segmentation instruction for the target text segment. Further, in response to the segmentation instruction for the target text segment, the target text segment in the video draft is segmented into two or more sub-text segments according to the segmentation instruction, and the second reading audio segment corresponding to each sub-text segment is re-obtained. Then, as in the above embodiments, the first reading audio segment in the audio editing track is replaced with each of the second reading audio segments, and the segment length of each sub-text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by each sub-text segment on the text editing track is aligned with the timeline interval occupied by the corresponding second reading audio segment on the audio editing track.
[0055] Based on any of the above embodiments, a data structure can also be provided to record the mapping relationship between target text segments and audio segments. For example, the mapping relationship between the identifier of the target text segment (text id) and the identifier of the audio segment (audio id) can be recorded so that when the video draft is loaded again in a later time, the target text segment and the audio segment can be queried and loaded according to the mapping relationship. Optionally, the target speech type (timbre id) used by each audio text can also be recorded to identify the timbre and other information of the audio text.
[0056] When the target text segment is updated, the above mapping relationship can be updated. Since the first audio segment is replaced by the second audio segment, the identifier (audio ID) of the audio segment changes in the above mapping relationship. Furthermore, if multiple first audio segments exist and any first audio segment is deleted, the corresponding mapping relationship needs to be deleted; or, if the user selects a new target speech type, a corresponding mapping relationship can be added.
[0057] Based on any of the above embodiments, undo and redo functions for modifying the target text segment can also be provided. When undoing the modification of the target text segment, the target text segment in the video draft can be restored to the target text segment before the update, and the second reading audio segment in the audio editing track can be replaced with the first reading audio segment. The first reading audio segment can be cached in memory. If it is not cached in memory, the first reading audio segment can be retrieved again, and the segment length of the target text segment on the text editing track of the video draft can be readjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track. When redoing the modification of the target text segment, S301-S303 can be re-executed. The second reading audio segment can be cached in memory. If it is not cached in memory, the second reading audio segment can be retrieved again.
[0058] In addition, it can provide deletion, undo, and redo functions for the first (or second) audio segment in the audio editing track. If only one first (or second) audio segment exists, it can be deleted directly. If multiple first (or second) audio segments exist, and the deleted first (or second) audio segment is not currently aligned with the target text segment (e.g., it is not the latest obtained first audio segment in the above embodiment), it can be deleted directly. If the target text segment is currently aligned with the first audio segment (e.g., the latest acquired first audio segment in the above embodiment), then after deleting the first audio segment (or the second audio segment), the segment length of the target text segment on the text editing track of the video draft is readjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by other first audio segments (or the second audio segment) on the audio editing track. These other first audio segments (or the second audio segment) can be the second-to-last acquired first audio segment (or the second audio segment). If an undo function is executed after the first audio segment (or the second audio segment) is deleted, the deleted first audio segment (or the second audio segment) can be added back to the audio editing track, and the display duration of the target text segment remains consistent with that before the first audio segment (or the second audio segment) was deleted. The redo function re-executes the deletion of the first audio segment (or the second audio segment), which will not be elaborated further here.
[0059] Based on any of the above embodiments, the speed of the first audio segment (or the second audio segment) can also be adjusted. That is, in response to the speed adjustment command for the first audio segment (or the second audio segment), the speed of the first audio segment (or the second audio segment) is adjusted in the audio editing track according to the speed adjustment command, including acceleration or deceleration. The playback duration of the corresponding first audio segment (or the second audio segment) changes. Therefore, it is also necessary to adjust the segment length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the speed-adjusted first audio segment (or the second audio segment) on the audio editing track.
[0060] Based on any of the above embodiments, a preview function can also be provided during video editing. In response to a preview command for a target text segment in the video draft, the audio stream data corresponding to the target text segment is obtained and played. The audio stream data does not need to be added to the audio editing track and can be used to allow users to quickly preview the reading effect of the target text segment, enabling them to adjust the target speech type, playback speed, and content of the target text segment. The audio stream data differs from the first (or second) audio segment mentioned above. While the first (or second) audio segment can be in the form of an audio file, the audio stream data can be in the form of streaming data. Furthermore, the audio stream data can also be obtained from a server. For example, the terminal device sends a preview request to the server, the server generates the audio stream data, and returns it to the terminal device for streaming playback.
[0061] Based on any of the above embodiments, during the process of acquiring the first audio segment (or the second audio segment) or listening to it, abnormal situations may occur, such as the text content of the target text segment being too long, or the target speech type being unable to read the target text segment (e.g., language type mismatch), or the target text segment containing some blocked words that cannot be read. In such cases, an error message can be displayed, and / or a retry can be performed.
[0062] Based on any of the above embodiments, the terminal device may provide an Application Programming Interface (API). This API can implement the functions of acquiring the first (or second) audio segment, listening to the audio, and handling exceptions. The functions of acquiring the first (or second) audio segment and listening to the audio can be achieved by sending a request to the server and receiving the first (or second) audio segment or audio stream data returned by the server. The functions of acquiring the first (or second) audio segment, listening to the audio, and handling exceptions can be implemented by calling the API through a TTS (Text-to-Speech) handle. When calling the API, the required parameters, including the target text segment and the target speech type, can be determined first. Furthermore, the TTS handle can be initialized; if no TTS handle exists, it can be created first to call the API. Alternatively, a SAMI (Synchronous Accessible Media Exchange) handle can be created, and the SAMI SDK (Software Development Kit) can be called. The software development kit (SDK) provides the functionality to acquire the first (or second) audio clip, perform a preview, and handle the aforementioned exceptions. It can also add the first (or second) audio clip to the audio editing track and adjust the length of the target text clip on the text editing track of the video draft to align the timeline segment occupied by the target text clip with that of the first (or second) audio clip on the audio editing track. Optionally, the API interfaces can be called using delegates or callbacks. Alternatively, lazy loading can be used to initialize and destroy handles as needed, ensuring efficient resource utilization of the terminal device and optimizing video editing performance.
[0063] Corresponding to the video editing method in the above embodiments, Figure 4 is a structural block diagram of the video editing device provided in this disclosure embodiment. For ease of explanation, only the parts related to this disclosure embodiment are shown. Referring to Figure 4, the video editing device 400 includes: an acquisition unit 401, a processing unit 402, and a text adjustment unit 403.
[0064] The acquisition unit 401 is used to acquire the first reading audio segment corresponding to the target text segment in response to a trigger command for the target text segment in the video draft.
[0065] Processing unit 402 is used to add the first reading audio segment to the audio editing track of the video draft according to the starting point of the target text segment on the editing timeline of the video draft.
[0066] The text adjustment unit 403 is used to adjust the length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
[0067] In one or more embodiments of this disclosure, the acquisition unit 401 is further configured to, in response to a modification instruction for the target text segment, update the target text segment in the video draft, and reacquire the second reading audio segment corresponding to the updated target text segment;
[0068] The processing unit 402 is further configured to replace the first reading audio segment in the audio editing track with the second reading audio segment;
[0069] The text adjustment unit 403 is also used to adjust the length of the updated target text segment on the text editing track of the video draft so that the timeline interval occupied by the updated target text segment on the text editing track is aligned with the timeline interval occupied by the second reading audio segment on the audio editing track.
[0070] In one or more embodiments of this disclosure, when the acquisition unit 401 updates the target text segment in the video draft in response to a modification instruction for the target text segment, and re-acquires the second reading audio segment corresponding to the updated target text segment, it is configured to:
[0071] If the modification instruction for the target text segment is a segmentation instruction for the target text segment, then the target text segment is segmented into two or more sub-text segments according to the segmentation instruction, and the second reading audio segment corresponding to each sub-text segment is re-acquired.
[0072] In one or more embodiments of this disclosure, when acquiring the first audio segment corresponding to the target text segment, the acquisition unit 401 is used to:
[0073] Send a request to the server to obtain the audio of the target text segment, wherein the audio request includes the target text segment;
[0074] Receive the first audio segment returned by the server.
[0075] In one or more embodiments of this disclosure, before acquiring the first audio segment corresponding to the target text segment, the acquisition unit 401 is further configured to:
[0076] In response to the instruction to select a speech type, the target speech type is determined;
[0077] Accordingly, when acquiring the first audio segment corresponding to the target text segment, the acquisition unit 401 is used to:
[0078] Obtain the first audio segment of the target text segment read aloud using the target speech type.
[0079] In one or more embodiments of this disclosure, the target speech type includes one or more target speech types; correspondingly, when acquiring the first audio segment of the target text segment read aloud using the target speech type, the acquisition unit 401 is used to:
[0080] Obtain one or more of the first audio segments of the reading aloud, which are read aloud using the target speech type.
[0081] In one or more embodiments of this disclosure, when the processing unit 402 adds the first read-aloud audio segment to the audio editing track of the video draft, it is configured to:
[0082] Add one or more of the first audio segments to different audio editing tracks of the video draft;
[0083] Accordingly, when the text adjustment unit 403 adjusts the length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track, it is used to:
[0084] The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the latest acquired first reading audio segment on the audio editing track.
[0085] In one or more embodiments of this disclosure, the text adjustment unit 403 is used to:
[0086] In response to the deletion command for the latest acquired first audio segment, the length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the second-to-last acquired first audio segment on the audio editing track.
[0087] In one or more embodiments of this disclosure, the processing unit 402 is further configured to, in response to a speed change instruction for the first audio segment, change the speed of the first audio segment in the audio editing track according to the speed change instruction;
[0088] The text adjustment unit 403 is also used to adjust the length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment after speed change on the audio editing track.
[0089] In one or more embodiments of this disclosure, the acquisition unit 401 is further configured to, in response to a listening instruction for a target text segment in a video draft, acquire audio stream data corresponding to the target text segment, and play the audio stream data.
[0090] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0091] Referring to Figure 5, a schematic diagram of the structure of an electronic device 500 suitable for implementing embodiments of the present disclosure is shown. The electronic device 500 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 5 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.
[0092] As shown in Figure 5, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0093] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 shows an electronic device 500 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0094] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0095] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0096] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0097] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0098] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0100] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0101] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0102] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] In a first aspect, according to one or more embodiments of this disclosure, a video editing method is provided, comprising:
[0104] In response to a trigger command for a target text segment in a video draft, the first audio segment corresponding to the target text segment is obtained;
[0105] Based on the starting point of the target text fragment on the editing timeline of the video draft, the first reading audio fragment is added to the audio editing track of the video draft.
[0106] The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
[0107] According to one or more embodiments of this disclosure, the method further includes:
[0108] In response to the modification instruction for the target text segment, the target text segment in the video draft is updated, and the second reading audio segment corresponding to the updated target text segment is re-acquired;
[0109] Replace the first audio segment in the audio editing track with the second audio segment.
[0110] The length of the updated target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the updated target text segment on the text editing track is aligned with the timeline interval occupied by the second reading audio segment on the audio editing track.
[0111] According to one or more embodiments of this disclosure, the step of updating the target text segment in the video draft in response to a modification instruction for the target text segment, and re-acquiring the second reading audio segment corresponding to the updated target text segment, includes:
[0112] If the modification instruction for the target text segment is a segmentation instruction for the target text segment, then the target text segment is segmented into two or more sub-text segments according to the segmentation instruction, and the second reading audio segment corresponding to each sub-text segment is re-acquired.
[0113] According to one or more embodiments of this disclosure, obtaining the first audio segment corresponding to the target text segment includes:
[0114] Send a request to the server to obtain the audio of the target text segment, wherein the audio request includes the target text segment;
[0115] Receive the first audio segment returned by the server.
[0116] According to one or more embodiments of this disclosure, before obtaining the first audio segment corresponding to the target text segment, the method further includes:
[0117] In response to the instruction to select a speech type, the target speech type is determined;
[0118] Accordingly, obtaining the first audio segment corresponding to the target text segment includes:
[0119] Obtain the first audio segment of the target text segment read aloud using the target speech type.
[0120] According to one or more embodiments of this disclosure, the target speech type includes one or more target speech types; correspondingly, obtaining the first audio segment of the target text segment read aloud using the target speech type includes:
[0121] Obtain one or more of the first audio segments of the reading aloud, which are read aloud using the target speech type.
[0122] According to one or more embodiments of this disclosure, adding the first read-aloud audio segment to the audio editing track of the video draft includes:
[0123] Add one or more of the first audio segments to different audio editing tracks of the video draft;
[0124] Accordingly, adjusting the length of the target text segment on the text editing track of the video draft, so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track, includes:
[0125] The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the latest acquired first reading audio segment on the audio editing track.
[0126] According to one or more embodiments of this disclosure, the method further includes:
[0127] In response to the deletion command for the latest acquired first audio segment, the length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the second-to-last acquired first audio segment on the audio editing track.
[0128] According to one or more embodiments of this disclosure, the method further includes:
[0129] In response to a speed change command for the first audio segment, the speed of the first audio segment is changed in the audio editing track according to the speed change command;
[0130] The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first audio segment after speed adjustment on the audio editing track.
[0131] According to one or more embodiments of this disclosure, the method further includes:
[0132] In response to a listening instruction for a target text segment in a video draft, the system acquires the audio stream data corresponding to the target text segment and plays the audio stream data.
[0133] Secondly, according to one or more embodiments of this disclosure, a video editing device is provided, comprising:
[0134] The acquisition unit is used to acquire the first audio segment corresponding to the target text segment in response to a trigger command for a target text segment in a video draft.
[0135] The processing unit is configured to add the first reading audio segment to the audio editing track of the video draft based on the starting point of the target text segment on the editing timeline of the video draft.
[0136] The text adjustment unit is used to adjust the length of the target text segment on the text editing track of the video draft, so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
[0137] According to one or more embodiments of this disclosure, the acquisition unit is further configured to, in response to a modification instruction for the target text segment, update the target text segment in the video draft, and reacquire the second reading audio segment corresponding to the updated target text segment;
[0138] The processing unit is also configured to replace the first reading audio segment in the audio editing track with the second reading audio segment;
[0139] The text adjustment unit is also used to adjust the length of the updated target text segment on the text editing track of the video draft so that the timeline interval occupied by the updated target text segment on the text editing track is aligned with the timeline interval occupied by the second reading audio segment on the audio editing track.
[0140] According to one or more embodiments of this disclosure, when the acquisition unit updates the target text segment in the video draft in response to a modification instruction for the target text segment, and re-acquires the second reading audio segment corresponding to the updated target text segment, it is configured to:
[0141] If the modification instruction for the target text segment is a segmentation instruction for the target text segment, then the target text segment is segmented into two or more sub-text segments according to the segmentation instruction, and the second reading audio segment corresponding to each sub-text segment is re-acquired.
[0142] According to one or more embodiments of this disclosure, when the acquisition unit acquires the first audio segment corresponding to the target text segment, it is configured to:
[0143] Send a request to the server to obtain the audio of the target text segment, wherein the audio request includes the target text segment;
[0144] Receive the first audio segment returned by the server.
[0145] According to one or more embodiments of this disclosure, before acquiring the first audio segment corresponding to the target text segment, the acquisition unit is further configured to:
[0146] In response to the instruction to select a speech type, the target speech type is determined;
[0147] Accordingly, when acquiring the first audio segment corresponding to the target text segment, the acquisition unit is used to:
[0148] Obtain the first audio segment of the target text segment read aloud using the target speech type.
[0149] According to one or more embodiments of this disclosure, the target speech type includes one or more target speech types; correspondingly, when the acquisition unit acquires the first audio segment of the target text segment read aloud using the target speech type, it is used to:
[0150] Obtain one or more of the first audio segments of the reading aloud, which are read aloud using the target speech type.
[0151] According to one or more embodiments of this disclosure, when the processing unit adds the first read-aloud audio segment to the audio editing track of the video draft, it is configured to:
[0152] Add one or more of the first audio segments to different audio editing tracks of the video draft;
[0153] Accordingly, when the text adjustment unit adjusts the length of the target text segment on the text editing track of the video draft so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track, it is used to:
[0154] The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the latest acquired first reading audio segment on the audio editing track.
[0155] According to one or more embodiments of this disclosure, the text adjustment unit is used for:
[0156] In response to the deletion command for the latest acquired first audio segment, the length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the second-to-last acquired first audio segment on the audio editing track.
[0157] According to one or more embodiments of this disclosure, the processing unit is further configured to, in response to a speed change instruction for the first audio segment, change the speed of the first audio segment in the audio editing track according to the speed change instruction;
[0158] The text adjustment unit is also used to adjust the length of the target text segment on the text editing track of the video draft, so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment after speed change on the audio editing track.
[0159] According to one or more embodiments of this disclosure, the acquisition unit is further configured to, in response to a listening instruction for a target text segment in a video draft, acquire audio stream data corresponding to the target text segment, and play the audio stream data.
[0160] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0161] The memory stores computer-executed instructions;
[0162] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video editing method as described in the first aspect and various possible designs of the first aspect.
[0163] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, which, when executed by a processor, implement the video editing method described in the first aspect and various possible designs of the first aspect.
[0164] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, including computer execution instructions that, when executed by a processor, implement the video editing method described in the first aspect and various possible designs of the first aspect.
[0165] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0166] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0167] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video editing method, comprising: In response to a trigger command for a target text segment in a video draft, the first audio segment corresponding to the target text segment is obtained; Based on the starting point of the target text fragment on the editing timeline of the video draft, the first reading audio fragment is added to the audio editing track of the video draft. The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
2. The method according to claim 1, further comprising: In response to the modification instruction for the target text segment, the target text segment in the video draft is updated, and the second reading audio segment corresponding to the updated target text segment is re-acquired; Replace the first audio segment in the audio editing track with the second audio segment. The length of the updated target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the updated target text segment on the text editing track is aligned with the timeline interval occupied by the second reading audio segment on the audio editing track.
3. The method according to claim 2, wherein updating the target text segment in the video draft and re-acquiring the second audio segment corresponding to the updated target text segment in response to the modification instruction for the target text segment comprises: If the modification instruction for the target text segment is a segmentation instruction for the target text segment, then the target text segment is segmented into two or more sub-text segments according to the segmentation instruction, and the second reading audio segment corresponding to each sub-text segment is re-acquired.
4. The method according to any one of claims 1-3, wherein obtaining the first audio segment corresponding to the target text segment includes: Send a request to the server to obtain the audio of the target text segment, wherein the audio request includes the target text segment; Receive the first audio segment returned by the server.
5. The method according to any one of claims 1-3, wherein before obtaining the first audio segment corresponding to the target text segment, the method further comprises: In response to the instruction to select a speech type, the target speech type is determined; Accordingly, obtaining the first audio segment corresponding to the target text segment includes: Obtain the first audio segment of the target text segment read aloud using the target speech type.
6. The method according to claim 5, wherein the target speech type includes one or more target speech types; correspondingly, obtaining the first audio segment of the target text segment read aloud using the target speech type includes: Obtain one or more of the first audio segments of the reading aloud, which are read aloud using the target speech type.
7. The method of claim 6, wherein adding the first read-aloud audio segment to the audio editing track of the video draft comprises: Add one or more of the first audio segments to different audio editing tracks of the video draft; Accordingly, adjusting the length of the target text segment on the text editing track of the video draft, so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track, includes: The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the latest acquired first reading audio segment on the audio editing track.
8. The method according to claim 7, further comprising: In response to the deletion command for the latest acquired first audio segment, the length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the second-to-last acquired first audio segment on the audio editing track.
9. The method according to claim 1, further comprising: In response to a speed change command for the first audio segment, the speed of the first audio segment is changed in the audio editing track according to the speed change command; The length of the target text segment on the text editing track of the video draft is adjusted so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first audio segment after speed adjustment on the audio editing track.
10. The method according to claim 1, further comprising: In response to a listening instruction for a target text segment in a video draft, the system acquires the audio stream data corresponding to the target text segment and plays the audio stream data.
11. A video editing device, comprising: The acquisition unit is used to acquire the first audio segment corresponding to the target text segment in response to a trigger command for a target text segment in a video draft. The processing unit is configured to add the first reading audio segment to the audio editing track of the video draft based on the starting point of the target text segment on the editing timeline of the video draft. The text adjustment unit is used to adjust the length of the target text segment on the text editing track of the video draft, so that the timeline interval occupied by the target text segment on the text editing track is aligned with the timeline interval occupied by the first reading audio segment on the audio editing track.
12. An electronic device, comprising: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1-10.
13. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method as described in any one of claims 1-10.
14. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Method and device for automatically adding subtitle fragments and computer equipment
CN112738563A
Multimedia data generation method and device, electronic equipment, medium and program product
CN116049452A
Video editing method and device, electronic equipment and storage medium
CN116366917A
Multimedia data processing method and device, equipment and medium
CN117956100A
Personalized audio and / or video shows
US20160064033A1