Video and audio intelligent editing method based on large language model and intelligent capability
Through the large language model and AI intelligent video editing method, short videos that meet the characteristics of different short video platforms are automatically generated, solving the time-consuming and labor-intensive problem of producers and achieving efficient and safe work production.
Patent Information
- Application Number
- CN202410321409.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-07-22
AI Technical Summary
The differences in characteristics of different short video platforms have caused producers to spend a lot of time and energy editing their works, affecting the efficiency and security of publishing.
Using a smart audio-visual editing method based on large language models and AI intelligent analysis, short videos that meet the characteristics of different platforms are automatically generated through user uploaded materials, AI information extraction, creative script generation and auxiliary content synthesis.
It realizes efficient and safe automatic generation of short videos that meet the characteristics of each platform, improving the efficiency and security of work production.
Smart Images

Figure CN120358369A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of short video editing, and particularly to a method for intelligent video and audio editing based on large language models and intelligent capabilities. Background Art
[0002] With the development of technology, the way people obtain information is changing rapidly. To facilitate users in obtaining valuable information in a short time, short videos have emerged as the times require.
[0003] Short videos are short film videos, which is a way of spreading Internet content. Generally, they are videos with a duration of less than 5 minutes spread on Internet new media; with the popularization of mobile terminals and the acceleration of the network, short, flat, and fast large-traffic dissemination content has gradually won the favor of major platforms and the like.
[0004] However, different short video platforms have their unique characteristics and tones, and their requirements for short videos are not the same. Currently, in order for short video operators to achieve different characteristics on different platforms, short video producers need to spend a lot of time and energy editing their works, which affects the release efficiency of the works and may also affect the security of the works.
[0005] Therefore, there is an urgent need for a method for intelligent video and audio editing based on large language models and intelligent capabilities to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for intelligent video and audio editing based on large language models and intelligent capabilities to solve the problems existing in the above-mentioned prior art.
[0007] To achieve the above purpose, the present invention provides the following solutions:
[0008] The present invention provides a method for intelligent video and audio editing based on large language models and intelligent capabilities, including the following steps:
[0009] S1. The user uploads the original video and audio materials and sets the preset information;
[0010] S2. Extract the information in the original video and audio materials through the AI intelligent analysis ability;
[0011] S3. The large language model combines the preset information and the information extracted from the original video and audio materials to generate an overall work creative script;
[0012] S4. Create the original video and audio materials according to the creative script to generate a short video work;
[0013] S5. Generate the auxiliary content corresponding to the short video work through the large language model.
[0014] Preferably, in step S1, the preset information includes the user's personalized requirements, publishing channels, and Internet hot topics.
[0015] Preferably, in step S2, the information in the original audio-visual material includes image content, voice content, and subtitle content.
[0016] Preferably, in step S3, the creative script is generated through fine-tuning prompt engineering, multi-level prompt engineering, character-level prompt language, and / or task-level prompt language functions.
[0017] Preferably, step S4 specifically includes:
[0018] S41. Select available shots based on the understanding of the content of the creative script and the original audio-visual material;
[0019] S42. Perform speech synthesis according to the creative script;
[0020] S43. Select a packaging template that conforms to the production line according to the creative script, the understanding of the content of the original audio-visual material, and the preset information;
[0021] S44. Use the selected shots and packaging template to synthesize a short video work.
[0022] Preferably, in step S5, the auxiliary content includes a title, an abstract, and topic content.
[0023] Preferably, it further includes:
[0024] S6. Publish the short video work on the short video platform after adding the auxiliary content.
[0025] The present invention has achieved the following beneficial technical effects compared with the prior art:
[0026] An audio-visual intelligent editing method based on a large language model and intelligent capabilities provided by the present invention includes the following steps: S1. A user uploads original audio-visual material and sets preset information; S2. Information in the original audio-visual material is extracted through AI intelligent analysis capabilities; S3. The large language model combines the preset information and the information in the extracted original audio-visual material to generate an overall work creative script; S4. The original audio-visual material is created according to the creative script to generate a short video work; S5. Auxiliary content corresponding to the short video work is generated through the large language model; S6. The short video work is published on the short video platform after adding the auxiliary content; automated intelligent editing is performed using the large language model and AI intelligent analysis capabilities, and the original single audio-visual content is automatically generated into short videos that conform to the characteristics of various publishing channels, while ensuring the efficiency and security of work production. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 It is a flow block diagram of an intelligent video and audio editing method based on a large language model and intelligent capabilities provided by the present invention. Specific embodiments
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0030] The purpose of the present invention is to provide an intelligent video and audio editing method based on a large language model and intelligent capabilities to solve the problems existing in the prior art.
[0031] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0032] Embodiment 1:
[0033] This embodiment provides an intelligent video and audio editing method based on a large language model and intelligent capabilities. As Figure 1 shown, it includes the following steps:
[0034] S1. The user uploads the original video and audio materials and sets preset information; the preset information includes the user's personalized requirements, publishing channels, Internet hot topics, and other contents;
[0035] S2. Extract the information in the original video and audio materials through AI intelligent analysis capabilities; the information in the original video and audio materials includes image content, voice content, subtitle content, etc.;
[0036] S3. The large language model combines the preset information and the information extracted from the original video and audio materials to generate an overall work creative script; the creative script is generated through functions such as fine-tuning prompt engineering, multi-level prompt engineering, role-level prompt words, and / or task-level prompt words;
[0037] S4. Create the original video and audio materials according to the creative script to generate a short video work; specifically including:
[0038] S41. Select available shots based on the understanding of the creative script and the content of the original audio-visual materials;
[0039] S42. Perform speech synthesis according to the creative script; Speech synthesis mainly aims to convert text into anthropomorphic speech, and can be trained and adjusted in terms of timbre, emotion, mood, volume, speech rate, etc. to make the speech synthesis closer to the speech of natural persons;
[0040] S43. Select a packaging template that conforms to the production line according to the understanding of the creative script, the content of the original audio-visual materials, and the preset information;
[0041] S44. Use the selected shots and packaging templates to synthesize short video works;
[0042] S5. Generate auxiliary content corresponding to the short video work through a large language model; The auxiliary content includes titles, abstracts, topics, etc.;
[0043] S6. Publish the short video work with auxiliary content added on a short video platform.
[0044] In the above method, in step S2, the AI intelligent analysis ability is applied to process the materials, and multi-modal understanding of the original audio-visual content can be realized in the process. Among them, speech recognition refers to the process of converting the speech in the audio into text through machine training. In this method, all the language texts in the audio-visual need to be extracted from the original materials through speech recognition; subtitle detection and recognition is to judge through images and record the text image content that appears in the picture as text; image understanding mainly analyzes the content of the picture and vectorizes the content to facilitate obtaining pictures with the same or similar themes during the automatic editing process; image recognition mainly recognizes scenes, people, landscapes, buildings, items, etc. in the video and supports the extraction of image semantic information at different dimensional levels.
[0045] In steps S3 and S5 of the above method, a large language model is used for creation. The large language model (abbreviated as LLM in English, Large Language Model) is an artificial intelligence model designed to understand and generate human natural language. After being trained with a large amount of text data, the large language model can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. At the same time, the large language model can perform various human thinking abilities including in-context learning, instruction understanding and execution, reasoning, etc.
[0046] In this embodiment, the application of the large language model in the generation link of the creative script is the core application of this method. The large language model will generate a creative script by synthesizing various requirements such as the user's personalized needs, the characteristics of the publishing channel, and the content characteristics of the original materials. The main requirements for the application of the large language model involve constraints in aspects such as character setting, safe and reliable control, and result output parameter requirements. At the same time, the large language model is required to output a creative script in a funny way. For the generation link of the auxiliary content, the large language model is required to comprehensively understand the Internet hotspots, the characteristics of the publishing channel, and the created result video, and comprehensively output the core content.
[0047] According to the requirements of the creative script of the work in the production process, the above method will synthesize the Internet hotspots and channel characteristics, combine the content that conforms to the Internet hotspots with the audio-visual packaging, such as the application of bullet screens, and synthesize and output a short video that conforms to the characteristics of the publishing channel.
[0048] This invention uses specific examples to expound the principle and implementation mode of this invention. The description of the above embodiments is only used to help understand the method and its core idea of this invention; at the same time, for those of ordinary skill in the art, according to the idea of this invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be construed as a limitation to this invention.
Claims
1. An intelligent video and audio editing method based on large language models and intelligent capabilities, characterized in that: It includes the following steps: S1. The user uploads the original video and audio materials and sets the preset information; S2. Extract the information in the original video and audio materials through the AI intelligent analysis ability; S3. The large language model combines the preset information and the information in the extracted original video and audio materials to generate the overall work creative script; S4. Create the original video and audio materials according to the creative script to generate a short video work; S5. Generate the auxiliary content corresponding to the short video work through the large language model.
2. The method for intelligent video and audio editing based on large language models and intelligent capabilities according to claim 1, wherein: In step S1, the preset information includes the user's personalized requirements, release channels, and Internet hot topics.
3. The video and audio intelligent editing method based on large language models and intelligent capabilities according to claim 1, characterized in that: In step S2, the information in the original video and audio materials includes image content, voice content, and subtitle content.
4. The video and audio intelligent editing method based on large language models and intelligent capabilities according to claim 1, characterized in that: In step S3, the creative script is generated through the functions of fine-tuning prompt engineering, multi-level prompt engineering, character-level prompt language, and / or task-level prompt language.
5. The video and audio intelligent editing method based on large language models and intelligent capabilities according to claim 1, characterized in that: Step S4 specifically includes: S41. Select available shots according to the understanding of the content of the creative script and the original video and audio materials; S42. Perform voice synthesis according to the creative script; S43. Select the packaging template that meets the production line according to the creative script, the understanding of the content of the original video and audio materials, and the preset information; S44. Use the selected shots and packaging templates to synthesize the short video work.
6. The video and audio intelligent editing method based on large language models and intelligent capabilities according to claim 1, wherein: In step S5, the auxiliary content includes the title, abstract, and topic content.
7. The video and audio intelligent editing method based on large language models and intelligent capabilities according to claim 1, characterized in that: It also includes: S6. Publish the short video work on the short video platform after adding the auxiliary content.