Video content editing method and electronic equipment

By combining AI image and language models, video footage is broken down and atomic tasks are performed, solving the problem of low efficiency in short video production and achieving efficient video creative production and distribution.

CN122002097APending Publication Date: 2026-05-08HANGZHOU ALIBABA INT INTERNET IND CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU ALIBABA INT INTERNET IND CO LTD
Filing Date
2025-12-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The production and distribution of short videos face challenges such as short lifecycles of online materials, the need for frequent replacement, insufficient supply of high-quality internal content, and high external procurement costs, resulting in low efficiency in video creative production and low return on investment.

Method used

By combining AI image understanding models and AI language models, the original video footage is broken down into keyframe image sequences, scene segmentation and image content understanding are performed, video editing schemes are generated, and multiple audio and video tools are used to execute atomic tasks, realizing the automation and efficient production of video creations.

Benefits of technology

It improves the efficiency and return on investment of video creative production, ensures that the generated videos meet specific needs and platform requirements, and realizes the automation and efficient execution of the process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002097A_ABST
    Figure CN122002097A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video clip generation method and electronic equipment. The method comprises the following steps: receiving an original video material submitted by a user and video creation demand information expressed through a natural language; respectively preprocessing the original video material and the video creation requirement; performing scene segmentation and picture content understanding on the key frame picture sequence through an AI image understanding model, and generating a video understanding result in a text format; reasoning the video creation requirement and the video understanding result through an AI language model, and generating a video editing scheme in combination with video editing knowledge; and determining a tool corresponding to the atomization task according to the video editing scheme, constructing required parameter information for the tool, and calling the tool to generate a target video. According to the embodiment of the invention, the efficiency and the commissioning ratio of video creative production can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video editing technology, and in particular to video content editing methods and electronic devices. Background Technology

[0002] For operators of product information service systems (also known as "e-commerce platforms"), video advertising on external platforms is frequently necessary to promote the system's application clients and increase user numbers. Because video content is more likely to attract user clicks and thus guide them to download the relevant application clients, it typically achieves higher conversion rates compared to image-based advertisements. Therefore, short video content has become a crucial engine for e-commerce marketing. The explosive growth of short video platforms has brought unprecedented exposure opportunities to short video content, but it has also placed higher demands on the efficiency of video creation. This is especially true for cross-border e-commerce platforms, which not only need to advertise on multiple platforms but also involve multiple countries and languages, requiring even greater quantity and production efficiency in their videos.

[0003] The main challenges currently faced in the production and distribution of short videos include the short lifecycle of online video materials, requiring frequent replacement; insufficient internal supply of high-quality content; and high external procurement costs, making large-scale coverage impossible. These factors restrict the efficiency and return on investment of video creative production and urgently need to be addressed through technological means. Summary of the Invention

[0004] This application provides video content editing methods and electronic devices that can improve the efficiency and return on investment in video creative production.

[0005] This application provides the following solution: A video content editing method, comprising: Receive raw video footage submitted by users, as well as video creation requirements expressed in natural language; The original video footage and video creation requirements are preprocessed separately to decompose the original video footage into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. The keyframe image sequence is segmented and its content is understood using an artificial intelligence (AI) image understanding model, and a video understanding result in text format is generated. The video creation requirements and video understanding results are inferred by an AI language model, and a video editing scheme is generated by combining video editing knowledge. The video editing scheme includes the composition structure of the target video to be generated, as well as multiple atomic tasks to be executed. Based on the video editing scheme, the tool corresponding to the atomization task is determined, and the required parameter information is constructed for the tool. Then, the tool is invoked to generate the target video.

[0006] This also includes: By utilizing natural language processing, semantic understanding technology, and user interaction, the system identifies and confirms users' video creation needs.

[0007] The preprocessing of the original video footage also includes: Add auxiliary information to the keyframe image, the auxiliary information being used to characterize the timing information of the keyframe in the original video footage.

[0008] The auxiliary information includes the position information of the keyframe on the timeline of the original video material; The atomization task includes a video segmentation task, and the parameters required for the video segmentation task include segmentation point location information, which is expressed through the location information on the time axis. When generating a video editing scheme, the AI ​​language model uses the timeline information in the keyframe images to fuzzily determine the position of the segment division points. The process of constructing the required parameter information for the tool includes: Based on the fuzzy determined segment segmentation point location information, multiple video frames within the target time range are determined from the original video material; By comparing the differences between consecutive frames of multiple video frames within the target time range, the location of key frames for scene transitions is determined, and the location of segmentation points is accurately determined based on the location of the key frames for scene transitions.

[0009] In accurately determining the location of the paragraph dividing point, the method further includes: The silent segments within the target time range of the original video footage are detected so that the segment segmentation point can be determined by combining the keyframe position of scene switching and the location of the silent segment.

[0010] The preprocessing of the original video footage also includes: The audio content is extracted from the original video material and converted into text content, so that the AI ​​image understanding model can combine the text content converted from the audio content to generate video understanding results.

[0011] The step of using an AI image understanding model to perform scene segmentation and content understanding on the keyframe image sequence, and generating a text-formatted video understanding result, includes: It generates overall textual descriptions of the original video footage and segmented video understanding results based on scene segmentation.

[0012] The video editing generation scheme includes: When generating the prompt information for the AI ​​language model, the required video editing knowledge is dynamically determined based on the video creation needs and added to the prompt information.

[0013] The original video footage consists of at least two clips, and the video creation requirements include the need to mix and edit at least two original video clips to generate more videos. The AI ​​language model is specifically used in generating video editing schemes for: Correlation analysis is performed on video segments segmented from different video materials to identify video segments suitable for cross-material combination, so as to generate multiple target videos by combining video segments from different video materials.

[0014] This also includes: After identifying the multiple atomic tasks that need to be executed, the multiple atomic tasks are also orchestrated and an editing workflow is generated so that the corresponding tools can be called according to the editing workflow.

[0015] This also includes: The client is provided with the execution progress information of the atomic task so that the client can display the execution progress information.

[0016] This also includes: A text summary is generated based on the video editing scheme corresponding to the target video and provided to the client so that the text summary can be displayed through the client.

[0017] A method for delivering video content, including The system receives at least two original video clips submitted by a user, as well as video creation requirements expressed in natural language; the video creation requirements include the requirement to mix and edit at least two original video clips to generate more videos. The original video footage and video creation requirements are preprocessed to decompose the original video footage into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. The keyframe image sequence is segmented and its content is understood using an artificial intelligence (AI) image understanding model, and a video understanding result in text format is generated. The video creation requirements and video understanding results are inferred using an AI language model, and a video editing scheme is generated by combining video editing knowledge. The video editing scheme includes the composition structure of multiple target videos to be generated, as well as multiple atomic tasks to be executed. The composition structure includes: after performing correlation analysis on video segments segmented from different video materials to determine the video segments suitable for cross-material combination between different video materials, the video segments, their arrangement order, and sources required for the target video are determined. Based on the video editing scheme, the tool corresponding to the atomization task is determined, and the required parameter information is constructed for the tool. The tool is then invoked to generate multiple target videos, which are then deployed to at least one target system.

[0018] The atomization task includes at least a video segmentation task and a video stitching task.

[0019] A video content editing system, the system comprising multiple intelligent agents, the multiple agents including: The demand analysis agent is used to interact with users after receiving the original video materials submitted by users and the video creation demand information expressed in natural language. It uses natural language processing and semantic understanding technology to identify and confirm the user's video creation demand and convert it into a machine-readable format. The video preprocessing agent is used to preprocess the original video material and the video creation requirements respectively, and to decompose the original video material into multiple keyframe image sequences. The video understanding agent is used to perform scene segmentation and content understanding on the keyframe image sequence using an artificial intelligence (AI) image understanding model, and generate video understanding results in text format. The creative planning agent is used to reason about the video creation requirements and the video understanding results through an AI language model, and generate a video editing plan by combining video editing knowledge. The video editing plan includes the composition structure of the target video to be generated, as well as multiple atomic tasks to be executed. The video editing agent is used to determine the tool corresponding to the atomic task based on the video editing scheme, construct the required parameter information for the tool, and then call the tool to generate the target video.

[0020] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of any of the preceding methods.

[0021] An electronic device, comprising: One or more processors; and A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of any of the preceding methods.

[0022] A computer program product includes a computer program / computer executable instructions that, when executed by a processor in an electronic device, implement the steps of any of the preceding methods.

[0023] According to the specific embodiments provided in this application, the following technical effects are disclosed: Through this embodiment, when a user needs to generate a video, they can submit original video footage and video creation requirements expressed in natural language. Then, the original video footage and video creation requirements can be preprocessed separately. The original video footage is broken down into multiple keyframe image sequences, and the video creation requirements expressed in natural language are converted into a machine-readable format. An AI image understanding model is used to perform scene segmentation and content understanding on the keyframe image sequences, generating a text-based video understanding result. Then, an AI language model is used to infer the video creation requirements and the video understanding result, and combined with video editing knowledge, a video editing scheme is generated. The video editing scheme includes the composition structure of the target video to be generated and multiple atomic tasks to be executed. Finally, based on the video editing scheme, the tools corresponding to the atomic tasks are determined, and the required parameter information is constructed for the tools. The tools are then invoked to generate the target video. In this way, multiple different AI models are combined and the video production process is broken down into multiple atomic tasks. Each atomic task corresponds to a specific audio and video tool. For example, the AI ​​image understanding model can be used to understand the image content of the original video material and generate text descriptions, while the AI ​​language model can generate specific video editing schemes. This can include determining the composition structure of the target video and breaking it down into atomic sub-tasks. Finally, by calling the corresponding tools, the goal of producing a video that meets the requirements is achieved. This automates and optimizes the process, improving the efficiency and return on investment of video creative production.

[0024] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic diagram of the system architecture provided in the embodiments of this application; Figure 2 This is a flowchart of the first method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the video editing process provided in an embodiment of this application; Figure 4 This is a flowchart of the second method provided in the embodiments of this application; Figure 5 This is a schematic diagram of the system provided in the embodiments of this application; Figure 6 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0028] In this embodiment, to improve the efficiency and return on investment of video creative production, AI (Artificial Intelligence) models can be used to assist in the video production process. Here, AI models refer to deep learning models containing massive amounts of parameters; that is, artificial intelligence models with a huge parameter scale trained on massive amounts of data based on "deep learning" technology. They are not designed for a single specific task, but possess general understanding and generation capabilities, like an "all-around" brain. Their characteristics include: massive scale (many parameters, huge amount of training data, high computational resource consumption); emergent ability (when the model scale reaches a certain level, it will exhibit capabilities that cannot be achieved by smaller models, such as complex reasoning, context learning, code generation, etc.); and strong versatility (a basic large model can be fine-tuned and applied to a variety of different downstream tasks).

[0029] AI models can generally be categorized into unimodal AI models and multimodal AI models. Unimodal AI models are specifically designed to process and generate a single type of information. "Modality" can be understood as the carrier or form of information, such as text, images, or audio. For example, typical unimodal AI models may include AI language models (processing text, with core capabilities in understanding and generating human language; typical tasks include question answering, translation, writing, summarizing, and code generation), AI vision models (processing images, with core capabilities in recognizing, understanding, and generating image content; typical tasks include image classification, object detection, image generation, and semantic segmentation), and AI audio models (processing audio, with core capabilities in recognizing, understanding, and generating sound and music; typical tasks include speech recognition, speech synthesis, and music composition). Unimodal models are typically quite powerful within their respective domains. However, they also suffer from the problem of "information silos." For instance, a pure text model cannot "see" an image, and a vision model may not be able to "read" the text description next to an image, and so on.

[0030] Multimodal AI models are models capable of simultaneously processing, understanding, and generating multiple types of information. They attempt to mimic humans, integrating visual, auditory, and linguistic sensory information to form a more comprehensive and profound understanding of the world. Multimodal AI models can be further subdivided into general-purpose multimodal models and task-specific multimodal models. General-purpose multimodal models aim to be versatile "all-rounders," capable of handling inputs and outputs of any modal combination. For example, a multimodal model might not only process text but also understand user-uploaded images and engage in dialogue based on mixed text and image information. Task-specific multimodal models can include text-to-image / video models (inputting text and generating corresponding images or videos), image-to-text models (typical tasks include image description and visual question answering), and so on.

[0031] In the scenario of this application embodiment, the ultimate goal is to produce video materials suitable for distribution as advertising or other forms. To achieve this goal, considering the aforementioned classification of AI models, there are various specific implementation methods. For example, one approach is to use a multimodal AI model of the "text-to-image / video" type. Users only need to input a description of their needs in natural language, and the required video can be generated by the "text-to-image / video" type multimodal AI model. However, since videos for external distribution often need to have certain requirements regarding the specific subject being displayed—for example, the subject might be a product published in a product information service system, and information about a promotional activity might also need to be added to attract user clicks. After redirecting to a specific app, the page content corresponding to the product in the video might also need to be displayed to ensure continuity in the user's access process, and so on. Therefore, if implemented solely through a "text-to-image / video model," it is difficult to associate the generated video with the specific product or other subject published on the platform, making it difficult to produce videos that meet actual needs.

[0032] Another approach is through "image + text-to-video generation," which involves inputting original image / video footage along with a text description, then using an AI model to generate the target video. This method, by incorporating the original image / video footage as visual guidance, provides information such as the video's starting point, style, composition, subject, and initial scene, ensuring a highly controllable and precise video generation process. Simultaneously, the text description serves as a dynamic instruction, describing how elements in the image should change, move, and transform. It can be used to define the video's "script," timeline, and other elements, guiding the model to generate dynamic content. Therefore, compared to simple "text-to-video generation," it offers greater control and solves the problem that pure text descriptions may not be able to accurately specify the video's visual style or initial frame.

[0033] However, the "image + text-to-video" model is far more complex than the single-modal model and the aforementioned "text-to-image / video" model. The main challenges include: spatiotemporal consistency: each frame of the generated video must not only be logically consistent within itself but also maintain physical logic and object appearance coherence with preceding and following frames; cross-modal alignment: the model must deeply understand the visual elements of the image (e.g., a seated person) and the dynamic instructions of the text (e.g., "this person stands up and walks to the window") and accurately match the two; dynamic understanding and generation: the model needs to "infer" a reasonable dynamic process from static images. For example, given an image of a calm sea and the text "a storm is coming," the model needs to generate dynamic effects such as surging waves and a darkening sky. Therefore, although "image + text-to-video" can combine static visual creativity with dynamic linguistic imagination, providing unprecedented dynamic content creation capabilities, this technology is still in its early stages and has limitations in terms of video length and logical consistency.

[0034] In summary, using a single AI model (whether it's a "text-to-image / video" model or an "image + text-to-video" model) may not meet the requirements of this application's embodiments. Therefore, this application's embodiments also provide a method that combines multiple different AI models and breaks down the video production process into multiple atomic tasks. Each atomic task can be completed using specific audio and video tools to collectively achieve the goal of producing a video that meets the requirements. These atomic tasks may include video segmentation (dividing a video into multiple segments), video splicing (merging multiple video segments into one video), adding text to the video frame, etc. Correspondingly, there can be multiple audio and video tools, such as video segmentation tools, video splicing tools, text addition tools, etc. Based on the aforementioned audio and video tools, the specific AI model's task is to fully understand the user's video creation needs and the original video materials, and based on this, combine pre-set video editing knowledge information to generate a corresponding video editing scheme. The specific scheme may include the composition of the target video (e.g., which segments it needs to consist of, whether transition shots need to be added, etc.), and which specific atomic tasks need to be executed. In addition, the tools corresponding to each atomization task can be identified, and the required parameter information can be constructed for the specific tools before calling the tools to finally generate the target video.

[0035] To better facilitate interaction between different AI models, multiple AI agents can be defined. This means a multi-agent collaborative system can be provided, covering the entire process from needs analysis, video understanding, creative planning to video editing. Through atomic task decomposition and link reuse, each agent can collaborate efficiently and focus on specific task modules, achieving process automation and efficient execution, and making the creative planning and editing of video content more precise and flexible.

[0036] Specifically, such as Figure 1 As shown, embodiments of this application can provide a system for generating video content (since it is based on original video material, it belongs to a kind of "secondary creation" of video, referred to as "video secondary creation"). This system may include a demand confirmation agent, an original video preprocessing agent, and a... Figure 1 (Not shown in the image), video understanding agent, creative planning agent, editing execution agent, etc. In specific implementations, these agents can also be combined into a single agent (for example, video understanding and creative planning can correspond to the same agent, such as...). Figure 1 In the example shown, these two parts are collectively referred to as the Creative Planning Agent. Of course, this Agent may involve calling different AI models. Alternatively, multiple Agents can be scheduled by setting up multiple Sub-Agents under the main Agent, and so on.

[0037] The requirement confirmation agent is primarily used to understand and verify user requests expressed in natural language after receiving them. This includes converting the video creation requirements expressed in natural language into a machine-readable format. Additionally, if the user's requirements are unclear or ambiguous, they can be clarified through AI dialogue. For example, if a user uploads two original video clips and describes their creation requirement as "Please edit these three videos into five videos," there might be a discrepancy between the actual number of uploaded video clips and the number stated in the requirement description. In this case, dialogue with the user can determine the exact number of original video clips needed. If it's determined that three are required, an error message can be displayed, prompting the user to upload one more video clip. It's important to note that the requirement confirmation agent may be used in multiple subsequent processing stages, including the final product delivery, where it can also be used to return the specific delivery result to the user.

[0038] The raw video preprocessing agent is primarily used to preprocess user-uploaded raw video footage to make it more suitable for AI models. Specifically, since AI models used for video understanding typically have limitations on the size of the input video, if the raw video footage is high-resolution, the frame rate can be lowered. For example, by extracting a keyframe at regular intervals, the raw video footage can be broken down into a sequence of images composed of multiple keyframes. Furthermore, since AI models used for video understanding often have poor time-dimension perception, the position information of specific keyframes on the timeline of the raw video footage can be marked during the preprocessing stage. For example, assuming a keyframe is located at 0.5 seconds in the raw video footage, information such as "0.5 seconds" can be marked in the content of that keyframe. Subsequent AI models can parse this timeline information from the content, thus helping them to determine the appearance and disappearance times of certain text in the video, facilitating the determination of video segmentation points, etc. In addition, if the original video footage includes audio content, the audio content can be extracted from the original video footage and converted into text content during the preprocessing stage. This allows the AI ​​model to combine the text content obtained from the audio content with the text content to generate a text description of the original video footage when performing video understanding.

[0039] Video understanding agents are primarily used to perform video understanding on pre-processed raw video footage. This video understanding task can be implemented using AI models such as "image-to-text" models (as mentioned earlier, a type of multimodal model). Specifically, it can perform scene segmentation and content understanding on the pre-processed keyframe image sequences, generating text-formatted video understanding results. Scene segmentation refers to dividing a long sequence into multiple sub-sequences according to a scene, and marking corresponding scene identifiers, such as frame-to-frame for "beginning," frame-to-frame for "product introduction," frame-to-frame for "waterfall flow," and so on. Additionally, it can understand the content, including subject recognition and the identification of textual content (including existing subtitles and timeline markers added during pre-processing). The results of scene segmentation and content recognition can be expressed in text form, which may include segmented video understanding results generated based on scene segmentation, or a comprehensive description of the video footage.

[0040] The Creative Planning Agent primarily generates video editing solutions based on the understanding of user creative needs, the understanding of original video footage, and pre-built video editing knowledge bases and tool libraries. Since the understanding of user creative needs and original video footage can be expressed in text form, and the video editing solutions can also be expressed in text form—meaning the model's input and output are both text content—this task can be accomplished using an AI language model. Specifically, the video editing knowledge base can include relevant knowledge information for editing under various specific editing needs, including required editing steps, which segments can be combined with which segments, and other key video editing knowledge points. This video editing knowledge base can exist in the form of a lower-level knowledge base, which can further include marketing and advertising-related knowledge, product knowledge, business knowledge, data, in-site product data, campaign data, trending topics, popular products, and other information for reference during creative planning. In addition, a tool library can be provided, offering tools for various basic video editing capabilities (e.g., video segmentation, video splicing) and advanced video processing capabilities (e.g., slow-motion frame interpolation, video style transfer). Furthermore, the tool library can also include tools for generating transition footage, as video creation may require such generation. Correspondingly, a media library can be provided, including base video footage and supplementary materials. Moreover, descriptive information about specific tools can be provided to the AI ​​model, including which tools are available and their respective functional descriptions (corresponding atomized task information). The AI ​​language model can then generate specific video editing schemes based on this information, including the composition of the target video and the atomized tasks it needs to be broken down into. For example, suppose a user needs to mix two original video clips. The system can identify which segments from each original video clip can be combined to create a new video, based on the segments included in each clip and the scene markers corresponding to each segment. This allows for the determination of specific atomic tasks to be performed. For instance, this might involve segmenting segment 1 from video 1 and segment 2 from video 2, then concatenating segment 1 and segment 2 into a single video. This requires two segmentation tasks and one concatenation task. Through this atomic task breakdown, AI scripts (which serve as the "blueprint" for video content, determining its structure) and AI-generated audio (which converts text scripts into speech) can also be generated. Subsequent implementation solutions can then be developed based on these AI scripts and AI-generated audio.

[0041] In other words, while the aforementioned AI models for video understanding can extract specific text from images and understand the image content, their reasoning ability is relatively poor. For example, after extracting the text, they may not be able to accurately and deeply understand its specific meaning. Therefore, it is advisable to first convert the image content into a text description, and then use an AI language model for reasoning to generate a specific video editing scheme.

[0042] The editing execution agent, after generating a video editing plan, selects editing tools according to the atomic tasks outlined in the plan. It then invokes these tools to perform specific editing operations and ultimately generate the target video. This invocation of tools involves the automated generation of tool input parameters. For example, for a video segmentation tool, specific input parameters include the original video footage identifier and the start and end points of the segmentation. Furthermore, if there are multiple atomic tasks, these tasks can be orchestrated to generate an editing workflow, which then schedules the editing tools accordingly.

[0043] Through the cooperation of these multiple agents, automated video production can be achieved. Furthermore, the system can include an AI management module, which can include functions such as Prompt management, tool management, and knowledge base management. The Prompt is information provided to the AI ​​model to guide its content generation. Tool management allows for the addition, modification, and deletion of tools, while knowledge base management allows for the addition, modification, and deletion of underlying knowledge. Moreover, a long-term memory network can be used to store information such as role-specific dialogue / interaction history, role-specific tool call history / status, and role-specific material library call history / status.

[0044] The specific implementation schemes provided in the embodiments of this application will be described in detail below.

[0045] Example 1 First, Embodiment 1 of this application provides a video content editing method, see [link to embodiment]. Figure 2 The method may specifically include: S201: Receive the original video footage submitted by the user, as well as the video creation requirements expressed in natural language.

[0046] In practice, a user interface can be provided, through which users can submit specific requirements, including uploading original video materials and describing specific creative needs in natural language.

[0047] S202: Preprocess the original video material and the video creation requirements respectively, so as to decompose the original video material into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format.

[0048] After receiving the original video footage and describing the specific creative requirements in natural language, preprocessing can be performed. This includes breaking down the original video footage into multiple keyframe image sequences to meet the input requirements of the AI ​​model. In addition, the video creation requirements expressed in natural language can be converted into a machine-readable format.

[0049] This process of breaking down the original video footage involves decomposing the complete video file into individual images. For example, breaking it down at a frame rate of 5 means extracting 5 frames per second. The keyframes are representative frames extracted from the video. In practice, one frame can be extracted from the original video footage at regular intervals, achieving an extraction frequency of n frames per second. The number n can be determined based on the processing requirements of the AI ​​model. Furthermore, if the original video footage has a high resolution, it can be compressed to reduce the size of each image, for example, compressing it from 1920×1080 to 640×360, and so on.

[0050] Furthermore, subsequent video editing may involve segmenting at specific points in time or inserting content at specific points in time. However, the exact location for these segmentations or insertions needs to be determined by the AI ​​model through image understanding and other processing. However, AI models are not sensitive to temporal information in videos. Therefore, during the preprocessing of the original video footage, auxiliary information can be added to the keyframe images. This auxiliary information represents the temporal sequence of the keyframes within the original video footage. There are several ways to add this auxiliary information. For example, one method is to mark the position of the keyframe image on the timeline of the original video footage within the image's content. The subsequent AI image understanding model can then identify the timeline information from the image content during video understanding, which can be used by the AI ​​language model in the creative planning process when generating video editing plans.

[0051] In addition, if there is audio content in the original video footage, the audio content can be extracted from the original video footage and converted into text content through speech-to-text technology. Subsequently, during the image understanding process, when the AI ​​image understanding model generates segmented video understanding results based on scene segmentation, it also combines the text content converted from the audio content to generate a text description of the original video footage as a whole.

[0052] Furthermore, in a preferred approach, natural language processing, semantic understanding technologies, and user interaction can be used to identify and confirm users' video creation needs before converting them into a machine-readable format. For example, if a user's expression is unclear or incomplete, AI dialogue can be used to interact with the user to further confirm their specific needs or complete their requirements, and so on.

[0053] S203: The keyframe image sequence is segmented and its content is understood using an artificial intelligence (AI) image understanding model, and a video understanding result in text format is generated.

[0054] After preprocessing the original video footage, an image understanding stage can be performed. In this stage, an AI model can be invoked through a corresponding agent to understand the video content. In this embodiment, the image content first needs to be converted into text for description; therefore, this stage can be implemented using an "image-to-text" model. The results of the original video footage preprocessing can be input into the aforementioned AI image-to-text model, which can perform scene segmentation and image content understanding. Scene segmentation refers to detecting scene changes in an image sequence and dividing it into multiple segments according to scene boundaries. Each segment can correspond to a scene, such as an intro, product description, etc.

[0055] The video understanding results can include segmented understanding results for each scene segment, such as scene identifiers, start and end points, and frame information for each segment. Additionally, it can include overall textual descriptions of the original video footage, including descriptions of the main elements within the video.

[0056] S204: The AI ​​language model is used to reason about the video creation requirements and the video understanding results, and combined with video editing knowledge to generate a video editing scheme. The video editing scheme includes the composition of the target video to be generated and multiple atomic tasks to be executed.

[0057] After completing video understanding, the pre-processed video creation requirements and understanding results can be input into the AI ​​language model. The AI ​​language model then combines this information with video editing knowledge to generate a video editing plan. The specific video editing plan can include the structural composition of the target video and multiple atomic tasks to be performed.

[0058] It's important to note that during the creative planning stage, specifically when calling the AI ​​language model, inputting prompts is involved. These prompts can include specific video editing knowledge. This knowledge can be provided to the AI ​​language model in various ways. For example, one approach is to store the editing knowledge in a video editing knowledge base (an external knowledge base of the AI ​​language model) and insert this knowledge into the prompts, thus providing the editing knowledge to the AI ​​language model through these prompts. As mentioned earlier, the video editing knowledge base can include relevant knowledge information for various specific editing needs, including required editing steps, which segments can be combined with which segments, and other key video editing knowledge points.

[0059] Since different editing needs may require different knowledge, the knowledge base is usually very large. In principle, all knowledge in the knowledge base could be provided to the AI ​​model via prompts; however, this might exceed the AI ​​model's input length limit. Furthermore, inputting the full amount of knowledge each time is unnecessary and could even increase the AI ​​model's processing burden. Therefore, in practice, methods such as retrieval-enhanced generation can be used. That is, based on the specific needs of the current request, relevant knowledge fragments can be retrieved from the knowledge base and used as prompts for the AI ​​language model, which can then generate an answer based on this context. Alternatively, in practice, data in the knowledge base can be converted into a format that can be efficiently retrieved through vectorization to facilitate the aforementioned retrieval-enhanced generation. Additionally, the knowledge base can encapsulate specific video editing knowledge for complex and high-frequency demand scenarios, and then dynamically determine the prompt information based on specific editing needs. For example, for highlight montage requirements, the specific prompt information could include which video segments correspond to the plot, which segments represent the site's internal storyline, and so on.

[0060] S205: Determine the tool corresponding to the atomization task according to the video editing scheme, construct the required parameter information for the tool, and then call the tool to generate the target video.

[0061] After generating the video editing plan, which includes multiple atomic tasks to be executed, the execution phase begins by selecting the corresponding tools based on these atomic tasks. Input parameters are then constructed for each tool before it is invoked. Each tool corresponds to a specific atomic task; one atomic task corresponds to one tool. Therefore, once the atomic tasks to be executed are determined, the specific tools to be invoked can be identified. For example, if an editing plan requires atomic tasks such as video segmentation and splicing, then corresponding video segmentation and splicing tools need to be invoked. If there are multiple atomic tasks, they can be orchestrated to generate an editing workflow. The workflow defines the execution order of each atomic task, thus determining the order in which tools are invoked, the input and output data streams of parameters, etc. The editing tools can then be scheduled according to this workflow to complete the video generation.

[0062] It's important to note that the specific parameters required vary depending on the tool, and can be generated according to the specific needs of the tool. For example, an atomic task might include video segmentation. In this case, the parameters required for video segmentation would include the location information of the segmentation points, which is expressed through their position on the timeline. When determining these segmentation points, the AI ​​language model can first perform a preliminary fuzzy determination during creative planning. Specifically, when generating a video editing plan, the AI ​​language model can fuzzily determine the segmentation point locations based on the timeline information in the keyframe images (which could be the timeline information inserted into the keyframe images during the aforementioned preprocessing stage). It's called fuzzy determination because during preprocessing, keyframes are extracted from the original video footage, and the data input to the AI ​​model is a sequence of keyframe images—that is, only a portion of the image frames from the original video footage. Therefore, the segmentation point locations determined by the AI ​​language model can only pinpoint the distance between keyframes. However, the actual original video footage may contain more suitable segmentation points. These more suitable locations are usually near the segmentation points identified from the keyframe sequence. Therefore, the segmentation point locations identified from the keyframe image sequence can be called fuzzily determined locations. Subsequently, based on the aforementioned fuzzily determined segmentation point location information, multiple video frames within a target time range can be determined from the original video footage. Then, by comparing the differences between consecutive frames within the target time range, the scene transition keyframe locations are determined, and the segmentation point locations are precisely determined based on these scene transition keyframe locations.

[0063] Specifically, when accurately determining the segment breakpoint location, silent segments can be detected within the target time range of the original video footage. This allows for a comprehensive determination of the segment breakpoint location based on the location of scene transition keyframes and the location of the silent segments. For example, if a scene transition keyframe location happens to fall within the silent segment range, then it is a suitable location to serve as a segment breakpoint, and so on.

[0064] In practical implementation, parameter requirements and other information can be described in the prompts. After determining the truly suitable locations for segmentation points, the parameters required by the video segmentation tool can be constructed based on this information. Of course, the specific input parameters can also include information such as the identifiers of the original video footage. For example, the parameter structure can be defined based on standard JSON schemes (JSON, or JavaScript Object Notation, is a specification for defining and validating the structure of JSON data). Then, the video segmentation tool can be called based on the constructed parameters to complete the video segmentation. Specifically, the tool can be called in various ways. For example, it can be called via function calls, that is, the tool can be defined as a callable function, allowing the AI ​​model to call the tool and obtain the processing results returned by the tool. Of course, other methods can also be used to call the tool, which are not limited here.

[0065] In practical implementation, to enable batch video generation, at least two original video clips can be input. Specific video creation requirements may include the need to mix and edit at least two original video clips. For example, different segments from different videos can be recombine to generate a new video. In this case, when generating video editing schemes using an AI language model, correlation analysis can be performed on video segments from different video clips to identify suitable segments for cross-cutting. This allows the target video to be generated by combining these segments. For example, if two original video clips both include an intro segment and a waterfall-style segment, their waterfall styles can be swapped to generate a new video, and so on.

[0066] It should be noted that, in specific implementations, the client can also provide users with information such as the progress of the video generation process. For example, since the specific video generation task is divided into multiple atomic tasks in this embodiment, the execution progress information of each atomic task can be provided to the client for display. This allows users to intuitively observe the specific execution status through the client.

[0067] The final video clip can be provided to users for download, or it can be published to relevant channel systems through pre-known system interfaces. Furthermore, in practical implementation, a text summary can be generated based on the video editing scheme corresponding to the target video and provided to the client for display. For example, the aforementioned text summary can be provided along with the download link for the target video, allowing users to gain a general understanding of the video's content, generation method, and structure before downloading, ensuring it meets their needs and avoiding resource waste.

[0068] To better understand the solution provided in the embodiments of this application, the complete implementation steps are described in detail below through an example.

[0069] In this example, assuming the user has uploaded two original video clips and needs to perform a "highlight" edit based on these two clips (e.g., video A and video B), the following steps can be taken after confirming the user's specific requirements: Step 1: As Figure 3 As shown in (A), this step involves using a video frame extraction tool to break down the two original video clips into keyframe image sequences at a certain frame rate (e.g., 5fps). If the resolution is too high, resolution compression can be performed to facilitate subsequent AI model processing. Additionally, if the original video clips contain audio content, Automatic Speech Recognition (ASR) can be performed and converted into text content (which can be called "voiceover script"). Through this processing, video A and video B are converted into image sequences composed of N keyframes (the number of keyframes extracted from different videos can vary), and the voiceover text information for both videos is obtained. To facilitate the subsequent acquisition of timeline information by the AI ​​model, the position information of the specific keyframe images on the timeline of the original video clips can be marked in the image content.

[0070] Step 2: As Figure 3As shown in (B), video understanding can be performed in this step. Specifically, based on the understanding of the video frame and structure, scene segmentation can be performed, and the content of the frame can be comprehensively identified, with detailed information of each segment recorded. For example, two scenes are identified in video A: an "eye-catching opening" and a "waterfall flow"; three scenes are identified in video B: an "eye-catching opening," a "product introduction," and a "waterfall flow." Additionally, overall video feature recognition and video structure recognition can also be performed. Afterward, a video text description can be generated. This can include a description of the video as a whole from the perspectives of video features and structure, as well as segmented text descriptions of different scenes. For example, segmented descriptions can include storyboard names, frame descriptions, etc. The frame descriptions can include specific product information, scene information, subject information, etc.

[0071] Step 3: As Figure 3 As shown in (C), creative planning can be performed in this step. Specifically, based on the video understanding results (in text form) and the user's editing requirements, a specific video editing plan is generated. In this example, since highlight mixing of different videos is required, the main focus when generating the video editing plan is to perform correlation analysis on the video segments of different videos, including the correlation between the visuals and the spoken text, the rationality of splicing, etc. Afterwards, suitable video segments or user-specified segments can be matched. That is, it determines which scene segments are included in each video and which segments can be mixed together. For example, in this example, video A can be divided into two segments: an "eye-catching opening" and a "waterfall flow," and video B can be divided into three segments: an "eye-catching opening," a "product introduction," and a "waterfall flow." Correlation and splicing rationality analyses can be performed on these segments to obtain the creative planning plan. For example, the structure of finished product 1: Video A - Eye-catching opening + Video B - Product introduction + Video A - Waterfall flow; the structure of finished product 2: Video B - Eye-catching opening + Video A - Waterfall flow, and so on. In addition, the specific atomic task information that needs to be executed can be determined based on the above planning scheme, and this task information can also be added to the planning scheme. For example, in this case, in order to obtain the above structure of segments 1 and 2, atomic tasks such as video segmentation and segment splicing can be broken down.

[0072] Step 4: See Figure 3As shown in (D), in this step and subsequent steps, specific editing processes can be performed according to the previously obtained planning scheme to obtain the final product. Specifically, based on the atomic task information included in the planning scheme, the corresponding tools can be determined, and the required input parameter information can be generated for the specific tools. In addition, if multiple atomic tasks are included, a workflow can also be generated. For example, first, a video segmentation tool needs to be called to perform segmentation, and then a segment splicing tool needs to be called. The segmentation of video A and video B can be processed in parallel, while the segment splicing needs to be performed after the segmentation is completed, and so on. When generating tool input parameters, they can be processed separately according to the needs of different tools. For example, for the atomic task of video segmentation, the specific tool requires input parameters including information such as the position of the segmentation point. When generating a planning scheme, the specific segmentation point can be fuzzily determined. For example, suppose the segment from the first frame to a certain keyframe X is the beginning of video A. Since the keyframe's content indicates the timeline position, this timeline position information can be extracted from the keyframe's content and used as the fuzzily determined segmentation point position. For example, if the timeline information of keyframe X is 5 seconds, then the fuzzily determined segmentation point position can be 5 seconds. In video B, the product introduction segment runs from keyframe M to keyframe N, where keyframe M is at 7 seconds and keyframe N is at 13 seconds. Therefore, the fuzzy determination of the segmentation point positions in video B is 7 seconds and 13 seconds. To more accurately determine the segmentation point positions, a range can be defined from the original video footage, and the segmentation points can be precisely determined within this range. For example, for keyframe X in video A, a more suitable segmentation point position can be found within the range of 4 to 6 seconds in video A. For keyframes M and N in video A, more suitable segmentation points can be found in the ranges of 6-8 seconds and 12-14 seconds in video B, respectively. Specifically, when searching for suitable segmentation points within the target range, edge detection can be performed, including comparing differences between preceding and following frames. The location with the largest difference between preceding and following frames (or exceeding a certain threshold) can be considered a suitable segmentation point. Alternatively, silence detection can be performed within the target range, and the segmentation point location can be determined based on the silence detection results. For example, if the difference between preceding and following frames at a certain location exceeds a certain threshold and falls within a silence range, it can be determined as a suitable segmentation point, and so on. For instance, in the aforementioned examples, the segmentation point for the "eye-catching opening" segment in video A can ultimately be determined as 5.5 seconds, and the segmentation point for the "product introduction" segment in video B can ultimately be determined as 6.5 seconds to 13.5 seconds, and so on.

[0073] Step 5: See Figure 3In section (E), after determining the specific segmentation point, a specific video segmentation tool can be called to perform video segmentation. For example, the "eye-catching opening" segment from 0 to 5.5 seconds can be segmented from video A, and the "product introduction" segment from 6.5 seconds to 13.5 seconds can be segmented from video B.

[0074] Step 6: See Figure 3 In section (F), after segmenting into specific segments, a video splicing tool can be used to complete the video splicing. That is, segments 0 to 5.5 seconds from video A and segments 6.5 seconds to 13.5 seconds from video B can be spliced ​​together. In addition, a transition segment can be generated between these two segments using the transition tool.

[0075] It should be noted that the various models used in the embodiments of this application, including traditional models or AI models, can be implemented using existing open-source models, or can be trained or fine-tuned based on existing models. For example, for video advertisements that need to be displayed in a product information service system, the specific video materials are usually related to product introductions, etc. Therefore, such videos can be used as training samples to fine-tune image understanding models, and so on. The specific training or fine-tuning process of the models will not be detailed here.

[0076] In summary, through the embodiments of this application, when a user needs to generate a video, they can submit original video footage and video creation requirements expressed in natural language. Then, the original video footage and video creation requirements can be preprocessed separately to decompose the original video footage into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. An AI image understanding model is then used to perform scene segmentation and content understanding on the keyframe image sequences, generating a text-based video understanding result. An AI language model is then used to infer the video creation requirements and the video understanding result, and combined with video editing knowledge, a video editing scheme is generated. This video editing scheme includes the compositional structure of the target video to be generated and multiple atomic tasks to be executed. Finally, based on the video editing scheme, the tools corresponding to the atomic tasks are determined, and the required parameter information is constructed for the tools before invoking them to generate the target video. In this way, multiple different AI models are combined and the video production process is broken down into multiple atomic tasks. Each atomic task corresponds to a specific audio and video tool. For example, the AI ​​image understanding model can be used to understand the image content of the original video material and generate text descriptions, while the AI ​​language model can generate specific video editing schemes. This can include determining the composition structure of the target video and breaking it down into atomic sub-tasks. Finally, by calling the corresponding tools, the goal of producing a video that meets the requirements is achieved. This automates and optimizes the process, improving the efficiency and return on investment of video creative production.

[0077] Example 2 This second embodiment protects a specific application scenario (i.e., video content delivery) of the video content generation method provided in this application. Specifically, this second embodiment provides a video content delivery method, see [link to relevant documentation]. Figure 4 The method may specifically include S401: Receive at least two original video clips submitted by the user, as well as video creation requirements expressed in natural language; the video creation requirements include: the requirement to mix and edit at least two original video clips to generate more videos; S402: Preprocess the original video material and video creation requirements respectively, so as to decompose the original video material into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format; S403: Perform scene segmentation and image content understanding on the keyframe image sequence using an artificial intelligence (AI) image understanding model, and generate video understanding results in text format; S404: The AI ​​language model is used to reason about the video creation requirements and the video understanding results, and a video editing scheme is generated by combining video editing knowledge. The video editing scheme includes the composition structure of multiple target videos to be generated, and multiple atomic tasks to be executed. The composition structure includes: after performing correlation analysis on video segments segmented from different video materials to determine the video segments suitable for cross-material combination between different video materials, the video segments, their arrangement order and source required for the target video are determined. S405: After determining the tool corresponding to the atomization task according to the video editing scheme and constructing the required parameter information for the tool, the tool is invoked to generate multiple target videos so that the multiple target videos can be delivered to at least one target system.

[0078] The atomization task includes at least a video segmentation task and a video stitching task.

[0079] The above methods enable AI systems to perform secondary creations based on original video footage, allowing more video content to be generated from at least two original video clips to meet the needs of video distribution to multiple platforms.

[0080] Example 3 This third embodiment provides a video content generation system from a system perspective. See [link to documentation]. Figure 5 The system includes multiple intelligent agents, and the multiple agents include: The demand analysis agent 501 is used to interact with users after receiving the original video materials submitted by users and the video creation demand information expressed in natural language. It uses natural language processing and semantic understanding technology to identify and confirm the user's video creation demand and convert it into a machine-readable format. Video preprocessing Agent 502 is used to preprocess the original video material and video creation requirements respectively, and to decompose the original video material into multiple keyframe image sequences. The video understanding agent 503 is used to perform scene segmentation and content understanding on the keyframe image sequence using an artificial intelligence (AI) image understanding model, and generate video understanding results in text format. Creative planning Agent 504 is used to reason about the video creation requirements and the video understanding results through an AI language model, and generate a video editing plan by combining video editing knowledge. The video editing plan includes the composition structure of the target video to be generated, as well as multiple atomic tasks to be executed. Video editing Agent 505 is used to determine the tool corresponding to the atomization task according to the video editing scheme, construct the required parameter information for the tool, and then call the tool to generate the target video.

[0081] For the parts of Embodiments 2 and 3 that are not detailed above, please refer to Embodiment 1 and other parts of this specification. They will not be repeated here.

[0082] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0083] Corresponding to Embodiment 1, this application also provides a video content generation apparatus, which may include: The request receiving unit is used to receive the original video materials submitted by the user, as well as the video creation request information expressed in natural language; The preprocessing unit is used to preprocess the original video material and the video creation requirements respectively, so as to decompose the original video material into multiple key frame image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. The image understanding unit is used to perform scene segmentation and image content understanding on the keyframe image sequence through an artificial intelligence (AI) image understanding model, and generate video understanding results in text format. The editing scheme generation unit is used to reason about the video creation requirements and the video understanding results through an AI language model, and generate a video editing scheme by combining video editing knowledge. The video editing scheme includes the composition structure of the target video to be generated, as well as multiple atomic tasks to be executed. The tool invocation unit is used to determine the tool corresponding to the atomization task according to the video editing scheme, construct the required parameter information for the tool, and then invoke the tool to generate the target video.

[0084] In a specific implementation, the device may further include: The requirement confirmation unit is used to identify and confirm users' video creation requirements by utilizing natural language processing, semantic understanding technology, and user interaction.

[0085] The preprocessing unit can also be used for: Add auxiliary information to the keyframe image, the auxiliary information being used to characterize the timing information of the keyframe in the original video footage.

[0086] Specifically, the auxiliary information includes the position information of the keyframe on the timeline of the original video footage; The atomization task includes a video segmentation task, and the parameters required for the video segmentation task include segmentation point location information, which is expressed through the location information on the time axis. When generating a video editing scheme, the AI ​​language model uses the timeline information in the keyframe images to fuzzily determine the position of the segmentation points. At this point, the tool invocation unit can specifically be used for: Based on the fuzzy determined segment segmentation point location information, multiple video frames within the target time range are determined from the original video material; By comparing the differences between consecutive frames of multiple video frames within the target time range, the location of key frames for scene transitions is determined, and the location of segmentation points is accurately determined based on the location of the key frames for scene transitions.

[0087] In addition to accurately determining the position of the segment segmentation point, silent segments can also be detected in the segments within the target time range of the original video material, so as to comprehensively determine the position of the segment segmentation point based on the position of the scene switching keyframe and the position of the silent segment.

[0088] In addition, during the preprocessing of the original video material, audio content can be extracted from the original video material and converted into text content, so that the AI ​​image understanding model can combine the text content converted from the audio content to generate video understanding results.

[0089] Specifically, the image understanding unit can be used for: It generates overall textual descriptions of the original video footage and segmented video understanding results based on scene segmentation.

[0090] The editing scheme generation unit can specifically be used for: When generating the prompt information for the AI ​​language model, the required video editing knowledge is dynamically determined based on the video creation needs and added to the prompt information.

[0091] In one specific manner, the original video footage consists of at least two pieces, and the video creation requirement includes the need to mix and edit at least two original video footage pieces to generate more videos; The AI ​​language model is specifically used in generating video editing schemes to: perform correlation analysis on video segments segmented from different video materials, determine video segments suitable for cross-material combination in different video materials, so as to generate multiple target videos by combining video segments from different video materials across materials.

[0092] Additionally, the device may also include: The workflow generation unit is used to, after determining the multiple atomic tasks to be executed, to arrange the multiple atomic tasks and generate an editing workflow so that the corresponding tools can be called according to the editing workflow.

[0093] An execution progress information providing unit is used to provide the client with the execution progress information of the atomic task so that the client can display the execution progress information.

[0094] The text summary content generation unit is used to generate text summary content based on the video editing scheme corresponding to the target video, and provide it to the client so that the text summary content can be displayed through the client.

[0095] Corresponding to Embodiment 2, this application also provides a video content casting device, which may include... The raw material and demand receiving unit is used to receive at least two raw video materials submitted by the user, as well as video creation demand information expressed in natural language; the video creation demand includes: the demand to mix and edit at least two raw video materials to generate more videos; The preprocessing unit is used to preprocess the original video material and the video creation requirements respectively, so as to decompose the original video material into multiple key frame image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. The image understanding unit is used to perform scene segmentation and image content understanding on the keyframe image sequence through an artificial intelligence (AI) image understanding model, and generate video understanding results in text format. The editing scheme generation unit is used to reason about the video creation requirements and the video understanding results through an AI language model, and generate a video editing scheme by combining video editing knowledge. The video editing scheme includes the composition structure of multiple target videos to be generated, and multiple atomic tasks to be executed. The composition structure includes: after performing correlation analysis on video segments segmented from different video materials to determine the video segments suitable for cross-material combination between different video materials, the video segments, their arrangement order, and sources required for the target video are determined. The tool invocation unit is used to determine the tool corresponding to the atomization task according to the video editing scheme, construct the required parameter information for the tool, and then invoke the tool to generate multiple target videos so as to deliver the multiple target videos to at least one target system.

[0096] The atomization task includes at least a video segmentation task and a video stitching task.

[0097] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0098] And an electronic device, comprising: One or more processors; and A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.

[0099] A computer program product includes a computer program / computer executable instructions that, when executed by a processor in an electronic device, implement the steps of the method described in the foregoing method embodiments.

[0100] in, Figure 6 An exemplary architecture of an electronic device is shown, which may include a processor 610, a video display adapter 611, a disk drive 612, an input / output interface 613, a network interface 614, and a memory 620. The processor 610, video display adapter 611, disk drive 612, input / output interface 613, network interface 614, and memory 620 can communicate with each other via a communication bus 630.

[0101] The processor 610 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to achieve the technical solution provided in this application.

[0102] The memory 620 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 620 can store the operating system 621 for controlling the operation of the electronic device 600, and the basic input / output system (BIOS) for controlling the low-level operations of the electronic device 600. Additionally, it can store a web browser 623, a data storage management system 624, and a video editing system 625, etc. The aforementioned video editing system 625 can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when implementing the technical solution provided in this application through software or firmware, the relevant program code is stored in the memory 620 and is called and executed by the processor 610.

[0103] Input / output interface 613 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0104] Network interface 614 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0105] Bus 630 includes a pathway for transmitting information between various components of the device, such as processor 610, video display adapter 611, disk drive 612, input / output interface 613, network interface 614, and memory 620.

[0106] It should be noted that although the above-described device only shows the processor 610, video display adapter 611, disk drive 612, input / output interface 613, network interface 614, memory 620, bus 630, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the solution of this application, and does not necessarily include all the components shown in the figures.

[0107] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0108] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0109] The video content editing method and electronic device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A video content editing method, characterized in that, include: Receive raw video footage submitted by users, as well as video creation requirements expressed in natural language; The original video footage and video creation requirements are preprocessed separately to decompose the original video footage into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. The keyframe image sequence is segmented and its content is understood using an artificial intelligence (AI) image understanding model, and a video understanding result in text format is generated. The video creation requirements and video understanding results are inferred by an AI language model, and a video editing scheme is generated by combining video editing knowledge. The video editing scheme includes the composition structure of the target video to be generated, as well as multiple atomic tasks to be executed. Based on the video editing scheme, the tool corresponding to the atomization task is determined, and the required parameter information is constructed for the tool. Then, the tool is invoked to generate the target video.

2. The method according to claim 1, characterized in that, Also includes: By utilizing natural language processing, semantic understanding technology, and user interaction, the system identifies and confirms users' video creation needs.

3. The method according to claim 1, characterized in that, The preprocessing of the original video footage also includes: Add auxiliary information to the keyframe image, the auxiliary information being used to characterize the timing information of the keyframe in the original video footage.

4. The method according to claim 3, characterized in that, The auxiliary information includes the position information of the keyframe on the timeline of the original video material; The atomization task includes a video segmentation task, and the parameters required for the video segmentation task include segmentation point location information, which is expressed through the location information on the time axis. When generating a video editing scheme, the AI ​​language model uses the timeline information in the keyframe images to fuzzily determine the position of the segment division points. The process of constructing the required parameter information for the tool includes: Based on the fuzzy determined segment segmentation point location information, multiple video frames within the target time range are determined from the original video material; By comparing the differences between consecutive frames of multiple video frames within the target time range, the location of key frames for scene transitions is determined, and the location of segmentation points is accurately determined based on the location of the key frames for scene transitions.

5. The method according to claim 4, characterized in that, In accurately determining the location of the paragraph dividing point, the method further includes: The silent segments within the target time range of the original video footage are detected so that the segment segmentation point can be determined by combining the keyframe position of scene switching and the location of the silent segment.

6. The method according to claim 1, characterized in that, The preprocessing of the original video footage also includes: The audio content is extracted from the original video material and converted into text content, so that the AI ​​image understanding model can combine the text content converted from the audio content to generate video understanding results.

7. The method according to claim 1, characterized in that, The step of using an AI image understanding model to perform scene segmentation and content understanding on the keyframe image sequence, and generating a text-formatted video understanding result, includes: It generates overall textual descriptions of the original video footage and segmented video understanding results based on scene segmentation.

8. The method according to claim 1, characterized in that, The video editing generation scheme includes: When generating the prompt information for the AI ​​language model, the required video editing knowledge is dynamically determined based on the video creation needs and added to the prompt information.

9. The method according to claim 1, characterized in that, The original video footage consists of at least two pieces, and the video creation requirements include the need to mix and edit at least two original video footage to generate more videos. The AI ​​language model is specifically used in generating video editing schemes for: Correlation analysis is performed on video segments segmented from different video materials to identify video segments suitable for cross-material combination, so as to generate multiple target videos by combining video segments from different video materials.

10. The method according to any one of claims 1 to 9, characterized in that, Also includes: After identifying the multiple atomic tasks that need to be executed, the multiple atomic tasks are also orchestrated and an editing workflow is generated so that the corresponding tools can be called according to the editing workflow.

11. The method according to any one of claims 1 to 9, characterized in that, Also includes: The client is provided with the execution progress information of the atomic task so that the client can display the execution progress information.

12. The method according to any one of claims 1 to 9, characterized in that, Also includes: A text summary is generated based on the video editing scheme corresponding to the target video and provided to the client so that the text summary can be displayed through the client.

13. A method for delivering video content, characterized in that, include Receive at least two original video clips submitted by the user, as well as video creation requirements expressed in natural language; The video creation requirements include: the need to mix and edit at least two original video clips to generate more videos; The original video footage and video creation requirements are preprocessed to decompose the original video footage into multiple keyframe image sequences and convert the video creation requirements expressed in natural language into a machine-readable format. The keyframe image sequence is segmented and its content is understood using an artificial intelligence (AI) image understanding model, and a video understanding result in text format is generated. The video creation requirements and video understanding results are inferred using an AI language model, and a video editing scheme is generated by combining video editing knowledge. The video editing scheme includes the composition structure of multiple target videos to be generated, as well as multiple atomic tasks to be executed. The composition structure includes: after performing correlation analysis on video segments segmented from different video materials to determine the video segments suitable for cross-material combination between different video materials, the video segments, their arrangement order, and sources required for the target video are determined. Based on the video editing scheme, the tool corresponding to the atomization task is determined, and the required parameter information is constructed for the tool. The tool is then invoked to generate multiple target videos, which are then deployed to at least one target system.

14. The method according to claim 13, characterized in that, The atomization task includes at least video segmentation and video stitching tasks.

15. A video content editing system, characterized in that, The system includes multiple intelligent agents, which include: The demand analysis agent is used to interact with users after receiving the original video materials submitted by users and the video creation demand information expressed in natural language. It uses natural language processing and semantic understanding technology to identify and confirm the user's video creation demand and convert it into a machine-readable format. The video preprocessing agent is used to preprocess the original video material and the video creation requirements respectively, and to decompose the original video material into multiple keyframe image sequences. The video understanding agent is used to perform scene segmentation and content understanding on the keyframe image sequence using an artificial intelligence (AI) image understanding model, and generate video understanding results in text format. The creative planning agent is used to reason about the video creation requirements and the video understanding results through an AI language model, and generate a video editing plan by combining video editing knowledge. The video editing plan includes the composition structure of the target video to be generated, as well as multiple atomic tasks to be executed. The video editing agent is used to determine the tool corresponding to the atomic task based on the video editing scheme, construct the required parameter information for the tool, and then call the tool to generate the target video.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program performs the steps of the method described in any one of claims 1 to 14.

17. An electronic device, characterized in that, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 14.

18. A computer program product comprising a computer program / computer-executable instructions, characterized in that, When the computer program / computer executable instructions are executed by a processor in an electronic device, they implement the steps of the method according to any one of claims 1 to 14.