Video editing method and apparatus, and terminal device
By performing voice recognition and automatic editing of video materials, the problem of low efficiency of manual editing of videos is solved, and the terminal device can quickly generate voice-coherent videos.
Patent Information
- Application Number
- PCT/CN2025/078454
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-21
- Filing Date
- 2025-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
In the prior art, users are less efficient when they manually edit video materials to generate new videos, resulting in less efficient when the terminal device generates new videos.
By acquiring the video material and the first text, the video material is subject to speech recognition processing, and the text segments and periods corresponding to the voice are determined, the video material is automatically edited to generate the video segments corresponding to the target text segments, and splicing them into the target video according to the order of the text segments.
It realizes the terminal equipment to automatically edit video materials, improves the efficiency and accuracy of video generation, and reduces the complexity of manual processing.
Smart Images

Figure CN2025078454_28082025_PF_FP_ABST
Abstract
Description
Video editing method, device and terminal equipment
[0001] This application claims priority to Chinese Patent Application No. 202410195053.7 filed on February 21, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to a video editing method, apparatus, and terminal device. Background Art
[0003] The user can edit multiple video materials according to their voices to obtain multiple video segments, and then re-join the multiple video segments to obtain a new video, wherein the new video has better voice coherence.
[0004] Currently, users can manually extract audio frames from multiple video clips, edit them based on the audio frames, and merge the resulting video segments to create a new video with coherent audio. However, manual processing is inefficient, resulting in low efficiency in generating new videos on terminal devices. Summary of the Invention
[0005] The present disclosure provides a video editing method, apparatus, and terminal device, which are used to solve the technical problem of low efficiency in generating new videos by terminal devices.
[0006] In a first aspect, the present disclosure provides a video editing method, the video editing method comprising:
[0007] Obtain video material and first text;
[0008] Performing speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material;
[0009] Determining a target text segment constituting the first text from the text segments corresponding to the speech;
[0010] According to the target time period of the target text segment, the video material is edited to obtain a video segment corresponding to the target text segment, and a target video is determined based on the video segment corresponding to the target text segment.
[0011] In a second aspect, the present disclosure provides a video editing device, comprising an acquisition module, a processing module, a first determination module, and a second determination module, wherein:
[0012] The acquisition module is used to acquire video material and the first text;
[0013] The processing module is used to perform speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material;
[0014] The first determining module is used to determine a target text segment constituting the first text in the text segment corresponding to the speech;
[0015] The processing module is further configured to edit the video material according to the target time period of the target text segment to obtain a video segment corresponding to the target text segment;
[0016] The second determination module is configured to determine a target video according to a video segment corresponding to the target text segment.
[0017] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;
[0018] The memory stores computer-executable instructions;
[0019] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video editing method as described in the first aspect and various possible aspects of the first aspect.
[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible aspects of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0023] FIG2 is a flow chart of a video editing method provided by an embodiment of the present disclosure;
[0024] FIG3 is a schematic diagram of obtaining a first text according to an embodiment of the present disclosure;
[0025] FIG4 is a schematic diagram of a process for determining a text segment provided by an embodiment of the present disclosure;
[0026] FIG5 is a schematic diagram of determining a target text segment provided by an embodiment of the present disclosure;
[0027] FIG6 is a schematic diagram of a method for determining a video segment provided by an embodiment of the present disclosure;
[0028] FIG7 is a schematic diagram of a process for generating a target video according to an embodiment of the present disclosure;
[0029] FIG8 is a schematic diagram of a method for obtaining a first text provided by an embodiment of the present disclosure;
[0030] FIG9 is a schematic diagram of adjusting a video segment according to an embodiment of the present disclosure;
[0031] FIG10 is a schematic structural diagram of a video editing device provided by an embodiment of the present disclosure; and
[0032] FIG11 is a schematic structural diagram of a terminal device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0034] To facilitate understanding, the concepts involved in the embodiments of the present disclosure are explained below.
[0035] Terminal device: is a device with wireless transceiver function. The terminal device can be deployed on land, including indoors or outdoors, handheld, wearable or vehicle-mounted. The terminal device can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a vehicle-mounted terminal device, a wireless terminal in self-driving, a wireless terminal device in remote medical, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, a wearable terminal device, etc. The terminal device involved in the embodiments of the present disclosure can also be called a terminal, user equipment (UE), an access terminal device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile station, a remote station, a remote terminal device, a mobile device, a UE terminal device, a wireless communication device, a UE agent or a UE device, etc. Terminal devices can also be fixed or mobile.
[0036] In related technologies, users can edit multiple video materials based on the voice of multiple video materials to obtain multiple video segments, and then re-splice the multiple video segments to obtain a new video with better voice coherence. For example, if the user needs to edit a video with a voice text of word a-word b-word c, the terminal device can edit video segment 1 including the voice of word a in video material A, edit video segment 2 including the voice of word b in video material B, and edit video segment 3 including the voice of word c in video material C, and splice video segments 1, 2, and 3 to obtain the video, in which the voice of word a, word b, and word c are the voice. However, currently users can only edit new videos manually. For example, if the text of the speech that the user needs to edit is a video of word a-word b-word c, the user needs to manually extract the speech of words related to multiple video materials, and obtain the video segments related to word a, word b and word c through manual editing, and then splice the multiple video segments to obtain the video. In this way, the efficiency of manual processing is low, which leads to low efficiency in generating videos.
[0037] In order to solve the technical problems in the related art, the embodiment of the present disclosure provides a video editing method, which obtains video material and a first text, performs speech recognition processing on the video material, obtains the text segment corresponding to the speech of the video material and the target time period when the text segment appears in the video material, the terminal device can determine the target text segment that constitutes the first text in the text segment corresponding to the speech, the terminal device can determine the video material corresponding to the target text segment, and according to the target time period of the target text segment, determine the start timestamp and end timestamp in the video material corresponding to the target text segment, the terminal device can perform editing processing on the video material corresponding to the target text segment according to the start timestamp and the end timestamp, and then obtain the video segment corresponding to the target text segment, and the terminal device can determine the target video according to the video segment corresponding to the target text segment. In this way, since the terminal device can automatically perform editing processing on multiple video materials to obtain multiple video segments associated with the first text, the terminal device can quickly synthesize the target video associated with the first text, thereby improving the efficiency of video generation.
[0038] The application scenario of the embodiment of the present disclosure is described below with reference to FIG1 .
[0039] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Referring to Figure 1 , it includes: Video A and Video B. Video A may include video frames 1, 2, ..., and 6, and Video B may include video frames a, b, ..., and f. The text segment corresponding to the speech in Video Frames 1 and 2 is Text A, the text segment corresponding to the speech in Video Frame 3 is Text B, and the text segment corresponding to the speech in Video Frames 4, 5, and 6 is Text C.
[0040] Referring to Figure 1, the text segment corresponding to the speech in video frames a, b, and c in video B is text D, the text segment corresponding to the speech in video frames d and e is text E, and the text segment corresponding to the speech in video frame f is text F. The terminal device can automatically edit video A to obtain video frames 1 and 2, and edit video B to obtain video frames a, b, and c. The terminal device can concatenate video frames 1, 2, a, b, and c to obtain video C, where the text segment corresponding to the speech in video C is text A-text D. In this way, the terminal device can automatically edit videos with good speech coherence from multiple video sources, thereby improving the efficiency of video synthesis.
[0041] It should be noted that FIG1 is only an example of an application scenario of the embodiment of the present disclosure, and does not limit the application scenario of the embodiment of the present disclosure.
[0042] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0043] FIG2 is a flow chart of a video editing method provided by an embodiment of the present disclosure. Referring to FIG2 , the method may include:
[0044] S201: Obtain video material and first text.
[0045] The execution subject of the embodiment of the present disclosure may be a terminal device, or a video editing device provided in the terminal device. The video editing device may be implemented based on software, or based on a combination of software and hardware, which is not limited in the embodiment of the present disclosure.
[0046] The video material can be any video. For example, the video material can be any type of video, or a video of any scene, which is not limited in the present embodiment. For example, the video material can be a video of any user giving a speech, or a video of any user singing.
[0047] The video material may include voices. For example, the video material may include multiple voices, wherein the video material may include one object emitting voices, or may include multiple objects emitting voices, which is not limited in the present embodiment.
[0048] Optionally, the terminal device may determine the video material based on the target timbre. For example, the target timbre may be a timbre associated with the demand. If the terminal device needs to generate a video with timbre A (target timbre), the terminal device may obtain a video with the voice timbre of timbre A. If the terminal device needs to generate a video with timbre B (target timbre), the terminal device may obtain a video material with the voice timbre of timbre B.
[0049] It should be noted that the terminal device can also obtain video materials according to any feasible implementation method (for example, the database may include multiple videos, and the terminal device can obtain multiple videos in the database to obtain multiple video materials), and the embodiments of the present disclosure are not limited to this.
[0050] The first text may be any piece of text. The terminal device may obtain the first text from a database or obtain the first text input by a user in real time. This is not limited in the embodiment of the present disclosure.
[0051] Optionally, the terminal device may obtain the first text according to the following feasible implementation method: displaying a video editing page, wherein the video editing page includes a text input area, and in response to an operation of inputting text in the text input area, determining the text in the text input area as the first text.
[0052] For example, the terminal device can display a video editing page, which may include a text input area. If the user enters text 1 in the text input area, the terminal device can determine text 1 as the first text. If the user enters text 2 in the text input area, the terminal device can determine text 2 as the first text.
[0053] The following describes the process of the terminal device acquiring the first text in conjunction with FIG3 .
[0054] FIG3 is a schematic diagram of obtaining a first text provided by an embodiment of the present disclosure. Referring to FIG3 , it includes: a terminal device. The display page of the terminal device may be a video editing page, which may include a text input area and a video display area. The video display area is used to display the video generated by the terminal device. If the user enters "The weather is great today" in the text input area, the terminal device may determine that the first text is: "The weather is great today." In this way, the terminal device can flexibly obtain the first text based on the user's operation, thereby improving the user experience.
[0055] S202: Perform speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period in which the text segment appears in the video material.
[0056] The target time period can be the time period when the text segment corresponding to the speech appears in the video material. For example, if the text segments corresponding to the speech in the video material are Text 1 and Text 2, and if Text 1 appears between the 1st and 2nd seconds of the video material, and Text 2 appears between the 3rd and 5th seconds of the video material, then the target time period for Text 1 can be between the 1st and 2nd seconds of the video material, and the target time period for Text 2 can be between the 3rd and 5th seconds of the video material.
[0057] It should be noted that the text fragments may be words, short sentences, long texts, etc., which are not limited in the embodiments of the present disclosure.
[0058] It should be noted that after the terminal device performs voice recognition processing on the video material, it can obtain the text segment corresponding to the voice and the target time period corresponding to the text segment. The terminal device can also determine the target time period corresponding to the text segment according to any feasible implementation method. The embodiment of the present disclosure does not limit this.
[0059] Optionally, the terminal device can perform voice recognition processing on the video material according to the following feasible implementation method: obtain the voice of the video material, perform text recognition on the voice, obtain the text associated with the voice, and the time period of the text associated with the voice in the video material, perform word segmentation processing on the text associated with the voice, obtain the text segment corresponding to the voice and the target time period when the text segment appears in the video material.
[0060] Optionally, the terminal device may obtain the voice of the video material according to any feasible implementation method, which is not limited in the embodiments of the present disclosure.
[0061] Optionally, the terminal device can process the speech according to the Automatic Speech Recognition (ASR) technology, and then obtain the text segment associated with the speech and the time period corresponding to the text segment in the video material. For example, for any video material, after the terminal device obtains the speech of the video material, it can process the speech according to the ASR technology, and then obtain multiple text segments corresponding to the speech. Moreover, the terminal device can accurately determine the time period of the text segment corresponding to the speech in the video material based on the relationship between the speech and the video material. For example, the video material includes speech B of speech A, speech A is located in time period 1 of the video material, and speech B is located in time period 2 of the video material. Therefore, after the terminal device performs text recognition on speech A and speech B, the time period of the text segment corresponding to speech A in the video material is time period 1, and the time period of the text segment corresponding to speech B in the video material is time period 2.
[0062] 4 , the following describes a process in which a terminal device performs speech recognition processing on a video material to obtain a text segment corresponding to the speech of the video material and a target time period in which the text segment appears in the video material.
[0063] Figure 4 is a schematic diagram of a process for determining text fragments provided by an embodiment of the present disclosure. Please refer to Figure 4, which includes: video material and a speech recognition module. Among them, the terminal device (not shown in Figure 4) can input video material into the speech recognition module, and the speech recognition module can recognize the text corresponding to the speech of the video material, and perform word segmentation on the text to obtain multiple words (text fragments). Among them, the words corresponding to the speech of the video material can include word 1, word 2,..., word n, and the target time period of word 1 is 2 seconds to 3.59 seconds, the target time period of word 2 is 3.6 seconds to 6.79 seconds,..., and the target time period of word n is 70.8 seconds to 73.99 seconds. In this way, the terminal device can accurately obtain multiple words corresponding to the speech, and the target time period corresponding to each word, thereby improving the efficiency of video editing and the efficiency of video synthesis.
[0064] Optionally, the video editing page may also include a text segment display area. After the terminal device determines the text segment corresponding to the speech, the method further includes: displaying the text segment corresponding to the speech in the text segment display area. For example, the terminal device performs speech recognition processing on video material 1 and video material 2 to obtain text segment 1, text segment 2, ..., text segment n. The terminal device may display text segment 1, text segment 2, ..., text segment n in the text segment display area. In this way, the user can accurately obtain usable text segments, reduce the complexity of the synthesized video, and improve the efficiency of video synthesis.
[0065] S203: Determine a target text segment constituting the first text in the text segment corresponding to the speech.
[0066] Optionally, the target text segment may be a text segment that constitutes the first text among multiple text segments corresponding to the speech. For example, if the words corresponding to the speech include word 1, word 2, word 3, and word 4, and if the first text includes word 1, word 2, and word 3, the terminal device may determine word 1, word 2, and word 3 as target words. If the first text includes word 4, the terminal device may determine word 4 as the target word.
[0067] Among them, the terminal device can determine the target text segment constituting the first text in the text segment corresponding to the voice according to the following feasible implementation method: determine the text similarity between the text segment corresponding to the voice and the first text, and if the text similarity is greater than a preset threshold, determine the text segment corresponding to the voice as the target text segment.
[0068] Optionally, the terminal device can calculate the cosine similarity between the text segment corresponding to the speech and the first text, and then determine the cosine similarity as the text similarity between the text segment and the first text. It should be noted that the terminal device can also determine the text similarity between the text segment corresponding to the speech and the first text according to any feasible implementation method, and the embodiments of the present disclosure are not limited to this.
[0069] It should be noted that since the first text contains a large number of characters, in actual application, the preset threshold can be 0. If the text similarity between the text segment corresponding to the speech and the first text is greater than 0, it means that the first text includes the text segment corresponding to the speech, and the terminal device can determine the text segment as the target text segment.
[0070] The following describes the process of determining the target text segment by the terminal device in conjunction with FIG5 .
[0071] FIG5 is a schematic diagram of a method for determining a target text segment provided by an embodiment of the present disclosure. Please refer to FIG5 , which includes: a first text and a word set (a set of multiple text segments corresponding to the voice of the video material). Among them, the first text includes word 1, word 11, word 9 and word 17. The word set may include word 1 (text segment), word 2, ..., word 12. Since the text similarity between word 1, word 11 and word 9 in the word set and the first text is greater than a preset threshold, the terminal device can determine that the target text segment includes word 1, word 11 and word 9.
[0072] It should be noted that in the embodiment shown in Figure 5, the first text includes word 17. Since word 17 is not included in the word set, the terminal device can determine the meaning of word 17, and determine a word in the word set that has the same meaning or an opposite meaning to word 17, and determine the word as the target text segment corresponding to word 17. The terminal device can also determine the target text segment corresponding to word 17 in the word set according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.
[0073] S204: Edit the video material according to the target time period of the target text segment to obtain a video segment corresponding to the target text segment, and determine the target video according to the video segment corresponding to the target text segment.
[0074] The video segment may be the video segment corresponding to the target text segment in the video material. For example, if the target time period of the target text segment is from the 2nd second to the 3rd second of the video material, then after the terminal device edits the video material, it can obtain the video segment corresponding to the target text segment, which is from the 2nd second to the 3rd second of the video material.
[0075] Among them, the terminal device can edit the video material according to the following feasible implementation method: determine the video material corresponding to the target text segment, determine the start timestamp and end timestamp in the video material corresponding to the target text segment according to the target time period of the target text segment, and edit the video material corresponding to the target text segment according to the start timestamp and end timestamp.
[0076] Since the text segment corresponding to the speech may include multiple target text segments, and the multiple target text segments are associated with different video materials, the terminal device can determine the video material corresponding to the target text segment. For example, if target text segment 1 is the text segment corresponding to the speech of video material A, and target text segment 2 is the text segment corresponding to the speech of video material B, the terminal device can determine that the video material corresponding to target text segment 1 is video material 1, and the video material corresponding to target text segment 2 is video material B.
[0077] Optionally, since the terminal device can store text segments corresponding to speech based on a set, the terminal device can configure a corresponding text segment set for each video material. For example, the multiple text segments obtained by the terminal device based on the speech of video material 1 can be stored in text segment set A, and the multiple text segments obtained by the terminal device based on the speech of video material 2 can be stored in text segment set B. In this way, the video material corresponding to the target text segment can be accurately determined.
[0078] Optionally, the start timestamp may indicate the moment when the target text segment appears in the video material, and the end timestamp may indicate the moment when the target text segment disappears from the video material. The terminal device may determine the start timestamp and the end timestamp based on the target time period of the target text segment. For example, if the target text segment appears between the 2nd and 3rd seconds of the video material, the start timestamp may be the timestamp corresponding to the 2nd second, and the end timestamp may be the timestamp corresponding to the 3rd second.
[0079] It should be noted that the terminal device can determine the start timestamp and the end timestamp associated with the target text segment according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.
[0080] Optionally, after the terminal device determines the start timestamp and the end timestamp, it can edit the video material according to the start timestamp and the end timestamp to obtain the video segment corresponding to the target text segment. For example, the video material corresponding to the target text segment 1 is video material A, and the video material corresponding to the target text segment 2 is video material B. If the start timestamp of the target text segment 1 corresponds to the 2nd second and the end timestamp corresponds to the 3rd second, and the start timestamp of the target text segment 2 corresponds to the 7th second and the end timestamp corresponds to the 9th second, then the terminal device can edit the video material A from the 2nd second to the 3rd second to obtain video segment a, and edit the video material B from the 7th second to the 9th second to obtain video segment b, wherein video segment a is the video segment corresponding to the target text segment 1, and video segment b is the video segment corresponding to the target text segment 2.
[0081] 6 , the method for determining a video segment by a terminal device will be described below.
[0082] Figure 6 is a schematic diagram of a method for determining a video segment provided by an embodiment of the present disclosure. Please refer to Figure 6, which includes: a start timestamp and an end timestamp corresponding to the video material and the target text segment. Among them, the video material includes video frame 1, video frame 2, ..., video frame 26, and the timestamps corresponding to each video frame are timestamp A, timestamp B, ..., timestamp Z, the start timestamp is timestamp B, and the end timestamp is timestamp H. The terminal device (not shown in Figure 6) can edit the video material according to the start timestamp and end timestamp corresponding to the target text segment to obtain the video segment corresponding to the target text segment. Among them, the video segment can include video frame 2, video frame 3, ..., video frame 8. In this way, the terminal device can quickly edit the video material according to the start timestamp and end timestamp of the target text segment, thereby improving the efficiency of video editing.
[0083] The target video's speech may include multiple target text segments. For example, if the target text segments are word 1, word 2, and word 3, the target video's speech may include the speech of word 1, word 2, and word 3.
[0084] Optionally, the terminal device can determine the target video according to the following feasible implementation method: if the number of target text segments is 1, the video segment corresponding to the target text segment is determined as the target video; if the number of target text segments is greater than 1, the video segments corresponding to multiple target text segments are spliced to obtain the target video.
[0085] For example, if the first text includes only one word, the terminal device can determine a target text segment in the text segment corresponding to the voice of the video material, and then edit it to obtain a video segment. The terminal device can determine the video segment as the target video.
[0086] For example, if the number of target text segments is greater than one, the terminal device may edit multiple video materials to obtain multiple video segments, and the terminal device may splice the multiple video segments to obtain the target video. For example, after the terminal device edits multiple video materials based on the multiple target text segments, video segments 1, 2, and 3 may be obtained. The terminal device may splice the video segments 1, 2, and 3 to obtain the target video.
[0087] Among them, the terminal device splices the video segments corresponding to multiple target text fragments to obtain the target video. Specifically, it can be: determining the text order of multiple target text fragments in the first text, and splicing the multiple video segments corresponding to the multiple target text fragments according to the text order to obtain the target video.
[0088] The text order may be the order in which the target text segments appear in the first text. For example, if the first text is: Today's weather is really good, the target text segment may include: today, of, weather, really good, then the terminal device may determine the text order as: today - of - weather - really good.
[0089] Optionally, the terminal device can perform text recognition on the first text, and then obtain the text order of multiple target text fragments in the first text. The terminal device can also determine the text order of multiple target text fragments in the first text according to any feasible implementation method. The embodiments of the present disclosure are not limited to this.
[0090] Optionally, the terminal device can determine the splicing order of multiple video segments based on the text order, and splice the multiple video segments according to the splicing order to obtain the target video. For example, if the first text is: "The weather is great today", and the text order is: today - of - weather - great, the video segment corresponding to "today" is video segment 1, the video segment corresponding to "of" is video segment 2, the video segment corresponding to "weather" is video segment 3, and the video segment corresponding to "great" is video segment 4, the terminal device can determine the splicing order of the multiple video segments as: video segment 1 - video segment 2 - video segment 3 - video segment 4, and splice the multiple video segments according to this splicing order to obtain the target video.
[0091] The process of generating a target video is described below with reference to FIG7 .
[0092] FIG7 is a schematic diagram of a process for generating a target video provided by an embodiment of the present disclosure. Please refer to FIG7 , which includes: a video segment 1 corresponding to a target text segment A, a video segment 2 corresponding to a target text segment B, and a video segment 3 corresponding to a target text segment C. The terminal device (not shown in FIG7 ) can determine that the text order of the multiple target text segments is: target text segment C-target text segment A-target text segment B. Therefore, the terminal device can splice the multiple video segments according to the text order to obtain a target video, wherein the target video can include video segment 1, video segment 2, and video segment 3, and the playback order in the target video is: video segment 3-video segment 1-video segment 2. In this way, the terminal device can accurately and quickly generate the target video according to the order of appearance of the target words, thereby improving the generation efficiency and generation accuracy of the target video.
[0093] The disclosed embodiment provides a video editing method, which obtains video material and a first text, performs speech recognition processing on the video material, obtains a text segment corresponding to the speech of the video material and a target time period in which the text segment appears in the video material, and the terminal device can determine the target text segment that constitutes the first text in the text segment corresponding to the speech, and the terminal device can determine the video material corresponding to the target text segment, and according to the target time period of the text segment, determine the start timestamp and end timestamp in the video material corresponding to the target text segment, and the terminal device can perform editing processing on the video material corresponding to the target text segment according to the start timestamp and the end timestamp, and then obtain a video segment corresponding to the target text segment, and the terminal device can determine the target video according to the video segment corresponding to the target text segment. In this way, since the terminal device can automatically perform editing processing on multiple video materials to obtain multiple video segments associated with the first text, the terminal device can quickly synthesize the target video associated with the first text, thereby improving the efficiency of video generation.
[0094] Based on the embodiment shown in Figure 2, after the terminal device obtains the text segment corresponding to the voice of the video material, the above-mentioned video editing method also includes another method for obtaining the first text. Below, in combination with Figure 8, another method for obtaining the first text is described in detail.
[0095] FIG8 is a schematic diagram of a method for obtaining a first text provided by an embodiment of the present disclosure. Referring to FIG8 , the method includes:
[0096] S801: Obtain a second text describing the target intention.
[0097] The second text may be a text describing the target intent. For example, the target intent may be any intent such as making a funny video, making a video about wanting to eat, or making a video about reciting a poem, and the like, which is not limited in the present embodiment.
[0098] Optionally, the terminal device may determine the target intent based on the demand information and generate a second text based on the target intent. For example, if the demand information is to create a video of a poetry recitation, the terminal device may determine that the target intent is to create a video of a poetry recitation. Therefore, the terminal device may generate the second text: Create a video of a poetry recitation, or the terminal device may generate the second text: Create a video of a recitation of ancient poetry, etc., which is not limited in the present embodiment.
[0099] It should be noted that the terminal device can obtain the demand information according to any feasible implementation method, and the embodiments of the present disclosure are not limited to this.
[0100] Optionally, the terminal device may also determine the second text based on the text input by the user. For example, if the text input by the user is: make a speech video, the terminal device may determine the second text to be: make a speech video.
[0101] S802: Input the second text and the text segments corresponding to the speech into the neural network model to obtain a first text composed of the text segments corresponding to the speech.
[0102] The textual intent of the first text is the target intent. For example, if the target intent is to create a speech video, the textual intent of the first text can also be to create a speech video; if the target intent is to create a poetry recitation video, the textual intent of the first text can also be to create a poetry recitation video.
[0103] Optionally, the neural network model in the embodiment of the present disclosure may be any language model, and the embodiment of the present disclosure is not limited to this.
[0104] The terminal device can input the second text and the text segments corresponding to the speech into the neural network model, and the neural network model can generate the first text based on the target intent corresponding to the second text and the multiple text segments corresponding to the speech. For example, if the text segments corresponding to the speech include word 1, word 2, ... word n, and the second text can be: Make a speech video, then after the terminal device inputs the above words and the second text into the neural network model, the neural network model can output a speech text, which can be the first text. In this way, the terminal device can determine the multiple text segments that generate the first text as target text segments, and edit multiple video materials according to the target text segments to obtain multiple video segments. After the terminal device splices the multiple video segments, it can obtain a target video, and the text corresponding to the speech of the target video is the above speech text.
[0105] The disclosed embodiment provides a method for obtaining a first text. A terminal device can obtain a second text describing a target intent and input the second text and a text segment corresponding to the speech into a neural network model. The neural network model can compose a first text based on the text segment corresponding to the speech. The intent associated with the first text is the target intent. In this way, the terminal device can quickly generate the first text without the user having to create the first text, thereby reducing the complexity of video synthesis. Moreover, the first text is composed of text segments corresponding to the speech. Therefore, it can be ensured that the terminal device obtains the video segment corresponding to each target text segment in the video material, thereby improving the accuracy of video synthesis.
[0106] Based on any of the above embodiments, in the above video editing method, after the terminal device obtains the video segment, it can also adjust the playback speed of the video segment. The method for adjusting the playback speed of the above video segment is described below in conjunction with Figure 9.
[0107] FIG9 is a schematic diagram of adjusting a video segment according to an embodiment of the present disclosure. Referring to FIG9 , the method flow includes:
[0108] S901. Obtain background music.
[0109] The background music may be the background music of the target video. For example, after the terminal device generates the target video, the audio in the target video is mainly voice. Therefore, the terminal device may pre-set a piece of background music for the target video. After the terminal device generates the target video, the background music may be added to the target video, thereby improving the playback effect of the target video.
[0110] It should be noted that the terminal device can obtain background music according to any feasible implementation method, and the embodiments of the present disclosure are not limited to this.
[0111] S902: Determine the rhythm of the background music.
[0112] The rhythm of the background music may be the length and strength of the notes in the background music. For example, the rhythm of the background music may be accents or bass, the rhythm of the background music may be 4-4 beats or 8-8 beats, the rhythm of the background music may be single rhyme or double rhyme, etc., which is not limited in the present embodiment.
[0113] Optionally, the terminal device may determine the rhythm of the background music according to any feasible implementation method, which is not limited in the embodiment of the present disclosure.
[0114] S903: Adjust the playback speed of the video segment according to the rhythm of the background music.
[0115] Optionally, the terminal device can adjust the playback speed of the video segment according to the rhythm of the background music, and then adjust the audio playback speed corresponding to the target text segment so that the audio playback speed corresponding to the target text segment matches the rhythm of the background music. For example, if the rhythm of the background music is a 4-4 beat rhythm, and the 6 target text segments include 12 characters and 6 video segments, the terminal device can adjust the playback speed of the video segment corresponding to the 4th character, the video segment corresponding to the 8th character, and the video segment corresponding to the 12th character so that the 4th character, the 8th character, and the 12th character in the 6 text segments are played longer, thereby forming a beat that matches the background music.
[0116] For example, the target text segment consists of 4 characters, character 1 is located from 0 seconds to 0.5 seconds in the video segment corresponding to the target text segment, character 2 is located from 0.5 seconds to 1 second in the video segment corresponding to the target text segment, from 1 second to 1.5 seconds in the video segment corresponding to the target text segment, and from 1.5 seconds to 2 seconds in the video segment corresponding to the target text segment. If the terminal device reduces the playback speed of character 4, the terminal device can reduce the playback speed of the video segment from 1.5 seconds to 2 seconds in the video segment.
[0117] The disclosed embodiments provide a method for adjusting the playback speed of video segments. Background music is obtained and its rhythm is determined. Based on the rhythm of the background music, the playback speed of the video segments is adjusted so that the playback speed of the speech in the video segments matches the rhythm of the background music. A terminal device can then splice the adjusted video segments together to obtain a target video and add background music to the target video. This allows the rhythm of the speech corresponding to the target video to match the rhythm of the background music, thereby improving the playback quality of the target video.
[0118] FIG10 is a schematic diagram of the structure of a video editing device provided by an embodiment of the present disclosure. Referring to FIG10 , the video editing device 100 includes an acquisition module 101, a processing module 102, a first determination module 103, and a second determination module 104, wherein:
[0119] The acquisition module 101 is used to acquire video material and a first text;
[0120] The processing module 102 is configured to perform speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material;
[0121] The first determining module 103 is configured to determine a target text segment constituting the first text from the text segment corresponding to the speech;
[0122] The processing module 102 is further configured to edit the video material according to the target time period of the target text segment to obtain a video segment corresponding to the target text segment;
[0123] The second determining module 104 is configured to determine a target video according to the video segment corresponding to the target text segment.
[0124] According to one or more embodiments of the present disclosure, the processing module 102 is specifically configured to:
[0125] Determining the video material corresponding to the target text segment;
[0126] Determining a start timestamp and an end timestamp in the video material corresponding to the target text segment according to the target time period of the target text segment;
[0127] The video material corresponding to the target text segment is edited according to the start timestamp and the end timestamp.
[0128] According to one or more embodiments of the present disclosure, the second determining module 104 is specifically configured to:
[0129] If the number of the target text segments is 1, the video segment corresponding to the target text segment is determined as the target video;
[0130] If the number of the target text segments is greater than 1, the video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
[0131] According to one or more embodiments of the present disclosure, the second determining module 104 is specifically configured to:
[0132] determining a text order of the plurality of target text segments in the first text;
[0133] According to the text sequence, multiple video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
[0134] According to one or more embodiments of the present disclosure, the acquisition module 101 is specifically configured to:
[0135] Displaying a video editing page, wherein the video editing page includes a text input area;
[0136] In response to an operation of inputting text in the text input area, the text in the text input area is determined as the first text.
[0137] According to one or more embodiments of the present disclosure, the acquisition module 101 is specifically configured to:
[0138] Obtaining a second text describing the target intention;
[0139] The second text and the text segment corresponding to the speech are input into the neural network model to obtain a first text composed of the text segments corresponding to the speech, wherein the text intent of the first text is the target intent.
[0140] According to one or more embodiments of the present disclosure, the first determining module 103 is specifically configured to:
[0141] Determining text similarity between the text segment corresponding to the speech and the first text;
[0142] If the text similarity is greater than a preset threshold, the text segment corresponding to the speech is determined as the target text segment.
[0143] According to one or more embodiments of the present disclosure, the acquisition module 101 is further configured to:
[0144] Get background music;
[0145] determining the rhythm of the background music;
[0146] The playback speed of the video segment is adjusted according to the rhythm of the background music.
[0147] The video editing device provided in the embodiment of the present disclosure can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0148] FIG11 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present disclosure. Please refer to FIG11, which shows a schematic diagram of the structure of a terminal device 1100 suitable for implementing an embodiment of the present disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. The terminal device shown in FIG11 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0149] As shown in FIG11 , terminal device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1102 or programs loaded from a storage device 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of terminal device 1100 are also stored in RAM 1103. Processing device 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.
[0150] Typically, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the terminal device 1100 to communicate with other devices wirelessly or by wire to exchange data. Although FIG11 illustrates a terminal device 1100 having various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0151] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0152] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0153] The computer-readable medium may be included in the terminal device, or may exist independently without being incorporated into the terminal device.
[0154] The computer-readable medium carries one or more programs. When the one or more programs are executed by the terminal device, the terminal device executes the method shown in the above embodiment.
[0155] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, various methods that may be involved in the above embodiments are implemented.
[0156] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements various possible methods involved in the above embodiments when executed by a processor.
[0157] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0159] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0160] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0161] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0162] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0163] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0164] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0165] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thereby, the user can independently choose whether to provide personal information to the software or hardware such as the terminal device, application, server or storage medium that performs the operation of the technical solution of the present disclosure based on the prompt message. As an optional but non-limiting implementation method, in response to receiving an active request from the user, the method of sending the prompt message to the user can be, for example, a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the terminal device.
[0166] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0167] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws and regulations. Data may include information, parameters and messages, such as flow switching indication information.
[0168] In a first aspect, an embodiment of the present disclosure provides a video editing method, the video editing method comprising:
[0169] Obtain video material and first text;
[0170] Performing speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material;
[0171] Determining a target text segment constituting the first text from the text segments corresponding to the speech;
[0172] According to the target time period of the target text segment, the video material is edited to obtain a video segment corresponding to the target text segment, and a target video is determined based on the video segment corresponding to the target text segment.
[0173] According to one or more embodiments of the present disclosure, the editing of the video material according to the target time period of the target text segment includes:
[0174] Determining the video material corresponding to the target text segment;
[0175] Determining a start timestamp and an end timestamp in the video material corresponding to the target text segment according to the target time period of the target text segment;
[0176] The video material corresponding to the target text segment is edited according to the start timestamp and the end timestamp.
[0177] According to one or more embodiments of the present disclosure, determining a target video based on a video segment corresponding to the target text segment includes:
[0178] If the number of the target text segments is 1, the video segment corresponding to the target text segment is determined as the target video;
[0179] If the number of the target text segments is greater than 1, the video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
[0180] According to one or more embodiments of the present disclosure, the step of splicing the video segments corresponding to the multiple target text segments to obtain the target video includes:
[0181] determining a text order of the plurality of target text segments in the first text;
[0182] According to the text sequence, multiple video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
[0183] According to one or more embodiments of the present disclosure, obtaining a first text includes:
[0184] Displaying a video editing page, wherein the video editing page includes a text input area;
[0185] In response to an operation of inputting text in the text input area, the text in the text input area is determined as the first text.
[0186] According to one or more embodiments of the present disclosure, obtaining a first text includes:
[0187] Obtaining a second text describing the target intention;
[0188] The second text and the text segment corresponding to the speech are input into the neural network model to obtain a first text composed of the text segments corresponding to the speech, wherein the text intent of the first text is the target intent.
[0189] According to one or more embodiments of the present disclosure, determining a target text segment constituting the first text from the text segment corresponding to the speech includes:
[0190] Determining text similarity between the text segment corresponding to the speech and the first text;
[0191] If the text similarity is greater than a preset threshold, the text segment corresponding to the speech is determined as the target text segment.
[0192] According to one or more embodiments of the present disclosure, after obtaining the video segment, the method further includes:
[0193] Get background music;
[0194] determining the rhythm of the background music;
[0195] The playback speed of the video segment is adjusted according to the rhythm of the background music.
[0196] In a second aspect, an embodiment of the present disclosure provides a video editing device, comprising an acquisition module, a processing module, a first determination module, and a second determination module, wherein:
[0197] The acquisition module is used to acquire video material and the first text;
[0198] The processing module is used to perform speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material;
[0199] The first determining module is used to determine a target text segment constituting the first text in the text segment corresponding to the speech;
[0200] The processing module is further configured to edit the video material according to the target time period of the target text segment to obtain a video segment corresponding to the target text segment;
[0201] The second determination module is configured to determine a target video according to a video segment corresponding to the target text segment.
[0202] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:
[0203] Determining the video material corresponding to the target text segment;
[0204] Determining a start timestamp and an end timestamp in the video material corresponding to the target text segment according to the target time period of the target text segment;
[0205] The video material corresponding to the target text segment is edited according to the start timestamp and the end timestamp.
[0206] According to one or more embodiments of the present disclosure, the second determining module is specifically configured to:
[0207] If the number of the target text segments is 1, the video segment corresponding to the target text segment is determined as the target video;
[0208] If the number of the target text segments is greater than 1, the video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
[0209] According to one or more embodiments of the present disclosure, the second determining module is specifically configured to:
[0210] determining a text order of the plurality of target text segments in the first text;
[0211] According to the text sequence, multiple video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
[0212] According to one or more embodiments of the present disclosure, the acquisition module is specifically configured to:
[0213] Displaying a video editing page, wherein the video editing page includes a text input area;
[0214] In response to an operation of inputting text in the text input area, the text in the text input area is determined as the first text.
[0215] According to one or more embodiments of the present disclosure, the acquisition module is specifically configured to:
[0216] Obtaining a second text describing the target intention;
[0217] The second text and the text segment corresponding to the speech are input into the neural network model to obtain a first text composed of the text segments corresponding to the speech, wherein the text intent of the first text is the target intent.
[0218] According to one or more embodiments of the present disclosure, the first determining module is specifically configured to:
[0219] Determining text similarity between the text segment corresponding to the speech and the first text;
[0220] If the text similarity is greater than a preset threshold, the text segment corresponding to the speech is determined as the target text segment.
[0221] According to one or more embodiments of the present disclosure, the acquisition module is further configured to:
[0222] Get background music;
[0223] determining the rhythm of the background music;
[0224] The playback speed of the video segment is adjusted according to the rhythm of the background music.
[0225] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;
[0226] The memory stores computer-executable instructions;
[0227] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video editing method as described in the first aspect and various possible aspects of the first aspect.
[0228] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible aspects of the first aspect is implemented.
[0229] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0230] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0231] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A video editing method, comprising: Obtain video material and first text; Performing speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material; Determining a target text segment constituting the first text from the text segments corresponding to the speech; According to the target time period of the target text segment, the video material is edited to obtain a video segment corresponding to the target text segment, and a target video is determined based on the video segment corresponding to the target text segment.
2. The method according to claim 1, wherein The editing of the video material according to the target time period of the target text segment includes: Determining the video material corresponding to the target text segment; Determining a start timestamp and an end timestamp in the video material corresponding to the target text segment according to the target time period of the target text segment; The video material corresponding to the target text segment is edited according to the start timestamp and the end timestamp.
3. The method according to claim 1 or 2, wherein: The determining of the target video according to the video segment corresponding to the target text segment includes: If the number of the target text segments is 1, the video segment corresponding to the target text segment is determined as the target video; If the number of the target text segments is greater than 1, the video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
4. The method according to claim 3, wherein: The step of splicing the video segments corresponding to the plurality of target text segments to obtain the target video includes: determining a text order of the plurality of target text segments in the first text; According to the text sequence, multiple video segments corresponding to the multiple target text segments are spliced together to obtain the target video.
5. The method according to any one of claims 1 to 4, wherein: Get the first text, including: Displaying a video editing page, wherein the video editing page includes a text input area; In response to an operation of inputting text in the text input area, the text in the text input area is determined as the first text.
6. The method according to any one of claims 1 to 4, wherein: Get the first text, including: Obtaining a second text describing the target intention; The second text and the text segment corresponding to the speech are input into the neural network model to obtain a first text composed of the text segments corresponding to the speech, wherein the text intent of the first text is the target intent.
7. The method according to any one of claims 1 to 6, wherein: The determining of a target text segment constituting the first text from the text segment corresponding to the speech includes: Determining text similarity between the text segment corresponding to the speech and the first text; If the text similarity is greater than a preset threshold, the text segment corresponding to the speech is determined as the target text segment.
8. The method according to any one of claims 1 to 7, wherein: After obtaining the video segment, the method further includes: Get background music; determining the rhythm of the background music; The playback speed of the video segment is adjusted according to the rhythm of the background music.
9. A video editing device, comprising an acquisition module, a processing module, a first determination module and a second determination module, wherein: The acquisition module is configured to acquire video material and a first text; The processing module is configured to perform speech recognition processing on the video material to obtain a text segment corresponding to the speech of the video material and a target time period during which the text segment appears in the video material; The first determining module is configured to determine a target text segment constituting the first text in the text segment corresponding to the speech; The processing module is further configured to perform editing processing on the video material according to the target time period of the target text segment to obtain a video segment corresponding to the target text segment; The second determination module is configured to determine a target video according to a video segment corresponding to the target text segment.
10. A terminal device comprising: processor and memory, wherein The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the video editing method according to any one of claims 1 to 8.
11. A computer-readable storage medium storing computer-executable instructions, wherein: When the processor executes the computer-executable instruction, the video editing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
A method and apparatus for video retrieval
CN109101558A
Intelligent video automatic editing method
CN112423023A
Audio and video data processing method and device, electronic equipment and storage medium
CN116614669A
Automatic video mixing and cutting method based on text-video retrieval
CN116614672A
Video editing method and device and computer readable storage medium
CN117097944A