Short video generation method and device and electronic equipment
Through the combination of speech and text conversion and large language model, high-quality short videos are fully automated edited from long videos, solving the problems of low editing efficiency and high technical threshold in the existing technology, and improving the efficiency and quality of short video generation.
Patent Information
- Application Number
- CN202411958081.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
It is difficult for the existing technology to edit high-quality short videos from long videos, and the production process of traditional short videos is cumbersome and time-consuming, and there is a high technical threshold.
By obtaining prompt words and long videos with audio, voice text conversion is performed, combining prompt words, audio text, speaker information and voice language, a large language model is used to perform fully automated video editing to generate high-quality short videos.
It realizes fully automated video editing, improves the work efficiency of short video generation, reduces the technical threshold, and makes it easier for users to create high-quality short videos.
Smart Images

Figure CN119996789A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of video processing technology, and more specifically, to a short video generation method, device, and electronic device. Background Art
[0002] With the development of mobile Internet, short videos have become a form of content that quickly transmits information and is easy to consume, and are widely loved by users. However, despite the increasing demand for social media content creation, especially short video creation, the traditional short video production process is cumbersome, time-consuming, and has a high technical threshold, which to some extent inhibits users' enthusiasm for creation.
[0003] Thanks to breakthroughs in artificial intelligence technology in automatic speech recognition, natural language processing, and video content understanding, short video intelligent editing technology has emerged. This technology uses advanced artificial intelligence algorithms to automatically perform multiple tasks in video editing, significantly improving creative efficiency and reducing the difficulty of entry. Although short video intelligent editing technology has shown great potential, it is difficult to edit high-quality short videos from long videos due to the diversity of video types and content and limited computing resources. Summary of the invention
[0004] In response to the problems existing in the above-mentioned prior art, the embodiments of the present application provide a short video generation method, device, and electronic device, which implement short video editing operations based on voice information, thereby realizing fully automated video editing and improving the work efficiency of short video editing.
[0005] In a first aspect, an embodiment of the present application provides a short video generation method, comprising the following steps:
[0006] Get the cue words and long video with audio;
[0007] Converting the audio of the long video into speech-to-text to obtain text corresponding to the audio of the long video; and
[0008] The long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video.
[0009] Furthermore, before converting the audio of the long video into speech-to-text and obtaining the text corresponding to the audio of the long video, the method further includes:
[0010] Extracting audio from the long video; and
[0011] The audio is cut into speech segments for speech-to-text conversion.
[0012] Furthermore, after extracting the audio from the long video, the method further includes:
[0013] Separating human voice from background sound in the audio; and
[0014] Identify speaker information of a human voice in the audio.
[0015] Furthermore, after extracting the audio from the long video, the method further includes:
[0016] Identify the speech language of the audio.
[0017] Furthermore, the step of editing the long video into a short video according to the prompt word and the text corresponding to the audio of the long video includes:
[0018] The long video is edited into short videos through a large language model according to the prompt words, the text corresponding to the audio of the long video, the identified speaker information and the voice language of the audio.
[0019] Further, after the long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video, the method further includes:
[0020] Perform a deduplication operation on the short video.
[0021] Furthermore, after performing the deduplication operation on the short video, the method further includes:
[0022] The short video is rated.
[0023] In a second aspect, the embodiment of the present application further provides a short video generation device, including:
[0024] Long video acquisition module, used to acquire prompt words and long videos with audio;
[0025] a text conversion module, used to perform speech-to-text conversion on the audio of the long video to obtain text corresponding to the audio of the long video; and
[0026] The short video generation module is used to edit the long video into a short video according to the prompt word and the text corresponding to the audio of the long video.
[0027] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the short video generation method according to the first aspect described above when executing the program.
[0028] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein the computer program is used to implement the short video generation method according to the first aspect above.
[0029] The embodiments of the present application bring the following beneficial effects:
[0030] In the short video generation method provided in the embodiment of the present application, a prompt word and a long video with audio are first obtained, and the audio of the long video is converted from speech to text to obtain the text corresponding to the audio of the long video. Finally, according to the prompt word and the text corresponding to the audio of the long video, the long video is edited into a short video. Therefore, the short video generation method provided in the embodiment of the present application realizes a short video editing operation based on voice information, thereby realizing fully automated video editing and improving the work efficiency of short video editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0032] Figure 1 A schematic diagram of a process for generating a short video according to an embodiment of the present application;
[0033] Figure 2 A structural block diagram of a short video generation device provided in an embodiment of the present application;
[0034] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0035] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments described in the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0037] In the specification and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "multiple" means two or more. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.
[0038] Figure 1 is a flow chart of a short video generation method according to an embodiment of the present application. Figure 1 As shown, the short video generation method of the embodiment of the present application includes the following steps:
[0039] S101: Obtain prompt words and a long video with audio;
[0040] Currently, prompts, or prompts, play a key role in the process of interacting with AI. They are not just commands or queries, but also a communication mechanism to guide AI to understand the needs and intentions of users. The long videos mentioned here usually include human voices and are usually not less than 5 minutes long. Prompts convey the user's instructions to edit the long videos with human voice audio.
[0041] S102: Convert the audio of the long video into speech-to-text to obtain text corresponding to the audio of the long video; and
[0042] Specifically, after obtaining a long video with audio, it is necessary to use speech-to-text technology to obtain text and timestamp information corresponding to the speech content. The speech-to-text technology can be selected from any one of the following specific technical solutions: open source speech-to-text technologies such as Whisper and DeepSpeech, closed source commercial solutions such as Microsoft, Google, and Volcano, or self-developed speech-to-text algorithms.
[0043] S103: Editing the long video into short videos according to the prompt words and the text corresponding to the audio of the long video.
[0044] As described above, the input, output and specific tasks are specified through prompt words, and the text and speaker information corresponding to different speech time periods are obtained through the results of the speech-to-text output, and are edited together. For example, they are input into a large language model for editing, and the editing dimensions, included text, speaker, timestamp, etc. of each short video are output, so as to edit the input long video into short videos.
[0045] Therefore, in the short video generation method provided in the embodiment of the present application, the prompt word and the long video with audio are first obtained, and the audio of the long video is converted into speech text to obtain the text corresponding to the audio of the long video. Finally, according to the prompt word and the text corresponding to the audio of the long video, the long video is edited into a short video. Therefore, the short video generation method provided in the embodiment of the present application realizes a short video editing operation based on voice information, thereby realizing fully automated video editing and improving the work efficiency of short video editing.
[0046] Furthermore, in some embodiments of the present application, before converting the audio of the long video into speech-to-text and obtaining the text corresponding to the audio of the long video, the method further includes:
[0047] Extracting audio from the long video; and
[0048] The audio is cut into speech segments for speech-to-text conversion.
[0049] Specifically, before converting the audio of the long video to text and obtaining the text corresponding to the audio of the long video, it is necessary to extract the complete audio data in mp3 / wav format from the input long video. In order to reduce unnecessary resource consumption, some audio clips (for example, 30 seconds) are cut out for subsequent speech language recognition, that is, speech text conversion.
[0050] Furthermore, in some embodiments of the present application, after extracting the audio from the long video, the method further includes:
[0051] Separating human voice from background sound in the audio; and
[0052] Identify speaker information of a human voice in the audio.
[0053] Specifically, in order to obtain better quality speech text information, it is necessary to perform voice separation operation on the input audio, input the audio data separated from the long video, use, for example, a deep neural network to separate the human voice, distinguish the human voice from the background sound, and retain the human voice audio data for subsequent operations.
[0054] Furthermore, in some embodiments of the present application, after extracting the audio from the long video, the method further includes:
[0055] Identify the speech language of the audio.
[0056] That is, before using speech-to-text technology, it is necessary to obtain the language type information of the input audio. This information can be obtained using speech language recognition technology based on, for example, deep learning, in order to perform speech-to-text conversion and subsequent editing operations.
[0057] Further, in some embodiments of the present application, the step of editing the long video into a short video according to the prompt word and the text corresponding to the audio of the long video includes:
[0058] The long video is edited into short videos through a large language model according to the prompt words, the text corresponding to the audio of the long video, the identified speaker information and the voice language of the audio.
[0059] Specifically, the input, output and specific tasks of the language model are specified through prompt words. The results of speech-to-text conversion and speaker recognition output are integrated to obtain the text and speaker information corresponding to different speech time periods, which are input into the large language model for intelligent editing. The editing dimensions, text, speaker and timestamp of each short video are output. The specific editing strategy includes the following dimensions:
[0060] (1) Theme. Summarize the content of the long video and extract the core part, and output a short video that can summarize the entire long video content.
[0061] (2) Topics. Based on the subject content of the long video, the long video is split into multiple short videos according to the topic. Each output short video contains a complete topic.
[0062] (3) Wonderful examples. Extract descriptive information fragments about a story / news / hot spot / example that are not directly related to the topic. One example corresponds to one short video. The number of short videos output by this dimension is unlimited. When the large language model does not find any wonderful examples, the number of short videos output by this dimension is 0.
[0063] (4) Personal opinion. Organize the voice information of each speaker, condense the core content, and edit them separately to produce a short video for each speaker. When the input short video contains only one speaker, the editing logic is not enabled.
[0064] (5) Highlight clips. Output a collection of highlight clips. Each short video corresponds to a collection of highlight clips. Highlight clips in the same short video usually have no strong connection, but belong to the same type, for example: a collection of humorous highlight clips, a collection of thought-provoking highlight clips.
[0065] Furthermore, in some embodiments of the present application, it includes:
[0066] Perform a deduplication operation on the short video.
[0067] Specifically, after the long video is edited into a short video according to the text corresponding to the prompt word and the audio of the long video, it is necessary to deduplicate the results of the large language model editing to ensure the uniqueness and high quality of the output short video. The repetition threshold can be set to 0.6, for example.
[0068] Furthermore, in some embodiments of the present application, after performing the deduplication operation on the short video, the method further includes:
[0069] The short video is rated.
[0070] Specifically, the quality of the generated short videos is scored based on the large language model (the score range is 0-100) and the evaluation method is as follows:
[0071] Score=a*hook+b*flow+c*value+d*trend
[0072] Among them: a, b, c, d are all weight coefficients, which can be adjusted dynamically according to needs, and it is only necessary to ensure that the sum of the weight coefficients is 1; the value ranges of hook, flow, value, and trend are all 0-100. Hook refers to the score of the short video hook setting, including but not limited to continuous viewing appeal (viewpoint conflict, plot reversal), execution appeal, etc.; flow refers to the logic score, and clear narrative and strong logic can get higher scores; value refers to the short video value score, which is related to the quality and quantity of valuable learning points, ability improvement points, key information points, etc. provided in the full text; trend refers to the popularity and spreadability score of the short video. When the content / organization form of the short video conforms to the current hot trend or hits the hot controversial points, it can get a higher score.
[0073] It should be noted that after the long video is edited into short videos according to the above method, the content of each short video is summarized according to the dimension of short video editing based on the large language model, and short video sharing copy in different formats and styles is generated according to different short video content platforms (such as YouTube Shorts, TikTok). The content of the copy includes but is not limited to the short video title, short video content introduction, and short video tags.
[0074] In addition, the input long video and extracted audio are segmented and reorganized according to the output results of the large language model, and multiple short videos with audio are output. The output video ratio is automatically adjusted according to the preset output canvas ratio. The input is usually a 16:9 horizontal screen long video, and the output is usually a 9:16 vertical screen short video. Add text information as subtitles and add dynamic subtitle special effects. Output multiple short videos (with audio), output all edited short videos (including short video ratings and sharing copy), and you can export short videos to local, re-edit or distribute them to content creation platforms with one click.
[0075] Figure 2 2 is a structural block diagram of a short video generation device 200 provided in an embodiment of the present application. Figure 2 As shown, the short video generation device 200 of the embodiment of the present application includes: a long video acquisition module 210, a text conversion module 220 and a short video generation module 230, wherein:
[0076] A long video acquisition module 210 is used to acquire a prompt word and a long video with audio;
[0077] A text conversion module 220, configured to perform speech-to-text conversion on the audio of the long video to obtain text corresponding to the audio of the long video; and
[0078] The short video generation module 230 is used to edit the long video into a short video according to the prompt word and the text corresponding to the audio of the long video.
[0079] In the short video generation device provided in the embodiment of the present application, a prompt word and a long video with audio are first obtained, and the audio of the long video is converted from speech to text to obtain the text corresponding to the audio of the long video. Finally, according to the prompt word and the text corresponding to the audio of the long video, the long video is edited into a short video. Therefore, the short video generation method provided in the embodiment of the present application realizes a short video editing operation based on voice information, thereby realizing fully automated video editing and improving the work efficiency of short video editing.
[0080] It should be noted that the specific implementation method of the short video generation device of the embodiment of the present application is similar to the specific implementation method of the short video generation method of the embodiment of the present application. Please refer to the description of the method part for details, and no further details will be given here.
[0081] Figure 3 Schematic diagram of the structure of an electronic device 300 according to an embodiment of the present application.
[0082] like Figure 3As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage part 302 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0083] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.
[0084] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 309, and / or installed from a removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, the above-mentioned functions defined in the electronic device of the present application are executed.
[0085] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electronic device, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0086] In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction-executing electronic device, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0087] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based electronic device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0088] The units or modules involved in the embodiments of the present application may be implemented by software or hardware. The units or modules described may also be set in a processor, and the processor is used to implement the short video generation method when executing the program:
[0089] Get the cue words and long video with audio;
[0090] Converting the audio of the long video into speech-to-text to obtain text corresponding to the audio of the long video; and
[0091] The long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video.
[0092] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the short video generation method described in the present application:
[0093] Get the cue words and long video with audio;
[0094] Converting the audio of the long video into speech-to-text to obtain text corresponding to the audio of the long video; and
[0095] The long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video.
[0096] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiment; or may exist independently without being installed in the electronic device. The above computer program product stores one or more programs, and when the above programs are used by one or more processors to execute the short video generation method described in the present application:
[0097] Get the cue words and long video with audio;
[0098] Converting the audio of the long video into speech-to-text to obtain text corresponding to the audio of the long video; and
[0099] The long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video.
[0100] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. All equivalent structural changes made by using the contents of the present application specification and drawings under the application concept of the present application, or directly / indirectly used in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A short video generation method, characterized in that: The following steps are involved: Get the cue words and long video with audio; Convert the audio of the long video into speech-to-text, and obtain the text corresponding to the audio of the long video; and The long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video.
2. The short video generation method according to claim 1, characterized in that: Before converting the audio of the long video into speech-to-text and obtaining the text corresponding to the audio of the long video, the method further includes: Extracting audio from the long video; and The audio is cut into speech segments for speech-to-text conversion.
3. The short video generation method according to claim 2, characterized in that: After extracting the audio from the long video, the method further includes: Separating human voice from background sound in the audio; and Identify speaker information of a human voice in the audio.
4. The short video generation method according to claim 2, characterized in that: After extracting the audio from the long video, the method further includes: Identify the speech language of the audio.
5. The short video generation method according to claim 3 or 4, characterized in that: The step of editing the long video into a short video according to the prompt word and the text corresponding to the audio of the long video includes: The long video is edited into short videos through a large language model according to the prompt words, the text corresponding to the audio of the long video, the identified speaker information and the voice language of the audio.
6. The short video generation method according to claim 1, characterized in that: After the long video is edited into a short video according to the prompt word and the text corresponding to the audio of the long video, the method includes: Perform a deduplication operation on the short video.
7. The short video generation method according to claim 6, characterized in that: After the short video is deduplicated, the method further includes: The short video is rated.
8. A short video generation device, characterized in that: include: Long video acquisition module, used to acquire prompt words and long videos with audio; A text conversion module, used to convert the audio of the long video into speech-to-text, and obtain the text corresponding to the audio of the long video; and The short video generation module is used to edit the long video into a short video according to the prompt word and the text corresponding to the audio of the long video.
9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the short video generation method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the short video generation method according to any one of claims 1-7.