A method, device and storage medium for automatically creating a video using multiple AI models

By integrating different AI models into the video creation workflow within the visual editing interface, the problem of frequent platform switching in AI video generation is solved, achieving efficient video creation and quality assurance.

CN122513635APending Publication Date: 2026-08-04GUANGZHOU LANHAO ADVERTISING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU LANHAO ADVERTISING CO LTD
Filing Date
2026-05-27
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing AI video generation technologies, users need to frequently switch between different platforms, resulting in fragmented workflows, data incompatibility, cumbersome operations, and a high risk of image quality loss.

Method used

By deploying work nodes of the video creation workflow in the visual editing interface, including media material input nodes, output nodes, and AI model processing nodes from different vendors, configuring connection relationships to form a directed acyclic graph, and automatically executing the video creation workflow, the system integrates the capabilities of different AI models.

Benefits of technology

It ensures video quality while improving video creation efficiency, avoiding frequent switching and manual import operations, and increasing video creation efficiency by more than 300%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513635A_ABST
    Figure CN122513635A_ABST
Patent Text Reader

Abstract

The application provides a video automatic creation method and device based on multiple AI models, equipment and a storage medium. In a visual editing interface, a plurality of work nodes required for video creation workflow are deployed in response to a node deployment instruction. The work nodes include an input node of media material, an output node, and at least two processing nodes for calling AI models of different manufacturers. In response to a node editing instruction, the working parameters of the work nodes and the connection relationship between the input node, the output node and the processing nodes are configured to form a directed acyclic graph representing the video creation workflow. In response to a workflow execution instruction, the video creation workflow is automatically executed according to the directed acyclic graph, and the video creation result is output through the output node. The method is highly integrated, does not require frequent switching and frequent manual import on different platforms, ensures video quality, and improves video creation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia, and in particular to a method, apparatus, device and storage medium for automatic video creation using multiple AI models. Background Technology

[0002] Currently, AI video generation technology is developing rapidly, with various AI models from different manufacturers such as Keling, Jimeng, and Veo emerging. Because different AI models have different characteristics and advantages, the video creation process requires the use of corresponding videos at different stages. This necessitates users frequently switching between different web platforms or software, manually exporting and importing intermediate materials, resulting in a highly fragmented workflow, data incompatibility, cumbersome operations, and a high risk of image quality loss. Summary of the Invention

[0003] This application provides a method, apparatus, device, and storage medium for automatic video creation using multiple AI models, to address at least one problem existing in related technologies. The technical solution is as follows: In a first aspect, embodiments of this application provide a method for automatic video creation using multiple AI models, including: In the visual editing interface, in response to the node deployment command, several working nodes required for the video creation workflow are deployed; the working nodes include input nodes and output nodes for media materials, as well as at least two processing nodes for calling AI models from different vendors; In response to node editing instructions, the working parameters of the working nodes and the connection relationships between the input nodes, the output nodes and the processing nodes are configured to form a directed acyclic graph representing the video creation workflow; In response to the workflow execution command, the video creation workflow is automatically executed according to the directed acyclic graph, and the video creation result is output through the output node.

[0004] In one implementation, the deployment of several work nodes required for the video creation workflow in response to a node deployment command includes: In response to the node deployment command, select the corresponding working node from the node list area of ​​the visual editing interface; Move the selected work node to the workflow deployment area of ​​the visual editing interface; Each processing node in the node list area is pre-encapsulated with a corresponding AI model version number, AI model call address, authentication information, and response parsing logic.

[0005] In one implementation, configuring the working parameters of the worker node in response to a node editing command includes: In response to a node editing command, the media material to be processed is configured in the input node; In the processing node, feature parameters based on the AI ​​model are configured. The feature parameters include at least one of the following: the AI ​​model's thinking mode, prompt words, the format, style, aspect ratio, resolution, language conversion, and video super-resolution. The prompt words are either user-input content or the processing results automatically transmitted from the previous working node.

[0006] In one implementation, the step of automatically executing the video creation workflow based on the directed acyclic graph in response to the workflow execution instruction, and outputting the video creation result through the output node, includes: In response to the workflow execution command, the directed acyclic graph is parsed to determine the execution order between the various work nodes; According to the execution order, the media material of the input node is passed to the first processing node. The AI ​​model is called to process the media material through the first processing node and its feature parameters to determine the first result as the input of the second processing node. The AI ​​model is called to determine the second result as the input of the next processing node through the second processing node and its feature parameters until the last processing node outputs the video creation result. The video creation result is passed to the output node, and the video creation result is output through the output node.

[0007] In one implementation, the step of calling an AI model to process media materials using a first processing node and its feature parameters, and then calling the AI ​​model to process the media materials to determine a first result as input to a second processing node, includes: The first processing node accesses the AI ​​model's calling address and performs authentication based on the authentication information. Once authentication is successful, the AI ​​model is invoked based on the feature parameters of the first processing node to process the media material and obtain preliminary processing results; According to the response parsing logic, the preliminary processing result is converted into a first result compatible with the next working node, and the first result is used as the input of the second processing node.

[0008] In one embodiment, the method further includes: In response to workflow execution instructions, creation tasks that automatically execute video creation workflows are generated in the task queue; The task status of the creation task is monitored and displayed in real time. The task status includes the execution status and progress information of each work node, wherein each work node provides an individual execution start and stop function. The connection time of the processing node to the AI ​​model is recorded, as well as the feedback time from the processing node sending the request to call the AI ​​model to obtaining the preliminary processing result and the polling time from successfully connecting to the AI ​​model to outputting the video creation result. Timeout prompts are given for the connection time, the feedback time and the polling time respectively.

[0009] In one embodiment, the method further includes: On the output node, a thumbnail of the video creation result is displayed; In response to the first interactive command on the thumbnail, the media player is invoked to open the video creation result; or, In response to a second interactive command on the thumbnail, the video creation result is saved to a specified storage location or shared to a third-party application.

[0010] Secondly, embodiments of this application provide a multi-AI model video automatic creation device, comprising: The first module is used to deploy several working nodes required for the video creation workflow in response to node deployment instructions in the visual editing interface; the working nodes include input nodes and output nodes for media materials, as well as at least two processing nodes for calling AI models from different manufacturers; The second module is used to respond to node editing instructions, configure the working parameters of the working node and the connection relationship between the input node, the output node and the processing node, and form a directed acyclic graph representing the video creation workflow; The third module is used to respond to workflow execution instructions, automatically execute the video creation workflow according to the directed acyclic graph, and output the video creation result through the output node.

[0011] In one implementation, the third module is further configured to: In response to workflow execution instructions, creation tasks that automatically execute video creation workflows are generated in the task queue; The task status of the creation task is monitored and displayed in real time. The task status includes the execution status and progress information of each work node, wherein each work node provides an individual execution start and stop function. The connection time of the processing node to the AI ​​model is recorded, as well as the feedback time from the processing node sending the request to call the AI ​​model to obtaining the preliminary processing result and the polling time from successfully connecting to the AI ​​model to outputting the video creation result. Timeout prompts are given for the connection time, the feedback time and the polling time respectively.

[0012] In one implementation, the third module is further configured to: On the output node, a thumbnail of the video creation result is displayed; In response to the first interactive command on the thumbnail, the media player is invoked to open the video creation result; or, In response to a second interactive command on the thumbnail, the video creation result is saved to a specified storage location or shared to a third-party application.

[0013] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, wherein the memory stores instructions that are loaded and executed by the processor to implement the methods in any of the above-described embodiments.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed, implements the methods in any of the above-described embodiments.

[0015] The beneficial effects of the above technical solution include at least the following: In the visual editing interface, in response to node deployment commands, several working nodes required for the video creation workflow are deployed. These working nodes include input nodes for media materials, output nodes, and at least two processing nodes for calling AI models from different vendors. In response to node editing commands, the working parameters of the working nodes and the connection relationships between input nodes, output nodes, and processing nodes are configured to form a directed acyclic graph (DAG) representing the video creation workflow. In response to workflow execution commands, the video creation workflow is automatically executed according to the DAG, and the video creation results are output through the output nodes. This highly integrated approach eliminates the need for frequent switching between different platforms and frequent manual imports, ensuring video quality while improving video creation efficiency.

[0016] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, these aspects, embodiments, and features will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0017] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.

[0018] Figure 1This is a flowchart illustrating the steps of a multi-AI model video automatic creation method according to an embodiment of this application; Figure 2 This is a schematic diagram of a visual editing interface according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating the configuration of feature parameters of a processing node according to an embodiment of this application; Figure 4 This is a schematic diagram of a directed acyclic graph according to an embodiment of this application; Figure 5 This is a schematic diagram illustrating the configuration of authentication information according to an embodiment of this application; Figure 6 This is a structural block diagram of a multi-AI model video automatic creation device according to an embodiment of this application; Figure 7 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0019] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0020] Reference Figure 1 The flowchart illustrates a method for automatic video creation using multiple AI models according to an embodiment of this application. This method may include at least steps S100-S300: S100. In the visual editing interface, respond to the node deployment command and deploy several working nodes required for the video creation workflow.

[0021] Optionally, the working nodes include an input node for media materials, an output node, and at least two processing nodes for calling AI models from different vendors, with each processing node corresponding to an AI model.

[0022] S200, responding to node editing instructions, configures the working parameters of the working nodes and the connection relationships between input nodes, output nodes and processing nodes, forming a directed acyclic graph representing the video creation workflow.

[0023] S300 responds to workflow execution instructions and automatically executes the video creation workflow based on the directed acyclic graph, outputting the video creation results through output nodes.

[0024] The technical solution of this application embodiment deploys several working nodes required for the video creation workflow in a visual editing interface in response to node deployment instructions. The working nodes include input nodes for media materials, output nodes, and at least two processing nodes for calling AI models from different vendors. In response to node editing instructions, the working parameters of the working nodes and the connection relationships between the input nodes, output nodes, and processing nodes are configured to form a directed acyclic graph representing the video creation workflow. In response to workflow execution instructions, the video creation workflow is automatically executed according to the directed acyclic graph, and the video creation result is output through the output node. This highly integrated approach eliminates the need for frequent switching between different platforms and frequent manual imports, ensuring video quality while improving video creation efficiency.

[0025] It should be noted that the instructions involved in the embodiments of this application can be at least one of keyboard, mouse, touch, and voice, and are not specifically limited.

[0026] In one implementation, step S100 includes steps S110-S120: S110. In response to the node deployment command, select the corresponding working node from the node list area of ​​the visual editing interface.

[0027] like Figure 2 As shown, optionally, the visual editing interface has a node list area on the left, a workflow deployment area (central canvas) in the middle, and an operation description display area on the right (showing what operations the user performed, displayed in text order). Therefore, users can select the corresponding work nodes from the node list area of ​​the visual editing interface based on their needs through node deployment commands, selecting one work node at a time. For example, image upload and video upload are input nodes; "Keling Video," "Jimeng Video," "Banana Image," "Veo Video," image editing, and text recognition are processing nodes; text display and video (image) output are output nodes; text display is an output midway through the workflow, and video (image) output is the final output of the workflow.

[0028] It should be noted that the visual editing interface also supports editing operations such as copying, pasting, deleting, undoing / redoing work nodes.

[0029] S120. Move the selected work node to the workflow deployment area of ​​the visual editing interface.

[0030] Optionally, the selected work nodes can be dragged and dropped to the workflow deployment area of ​​the visual editing interface, thereby displaying all selected work nodes that have been moved to the workflow deployment area.

[0031] It should be noted that the number of input nodes in the workflow deployment area can be more than one, and the input of a processing node can be multiple. For example, the input of a processing node can be the output of several processing nodes, or it can be the output of an input node and the output of a processing node. The processing node of the same AI model can also be used multiple times, because the same AI model can also output content in different dimensions, without specific limitations.

[0032] In this embodiment, all working nodes in the node list area are pre-encapsulated. For example, each processing node pre-encapsulates the version number of a corresponding AI model, the calling address of the AI ​​model, authentication information, and response parsing logic. This is equivalent to establishing a unified adaptation interface for APIs from different vendors such as Keling AI (Kuaishou), Jimeng AI (ByteDance), Gemini (Google), Veo3 (Google), and GPT-5.2 (OpenAI). Optionally, if an AI model has different version numbers, a processing node can encapsulate the different version numbers of the AI ​​model, thus allowing for selection of different version numbers. For AI models from the same vendor, if no new functions are added, there is no need to add new working nodes; simply add the corresponding model name to the corresponding list in the Python file (which contains the encapsulated content of all working nodes). If models from the same company have different functions, corresponding processing nodes can be newly designed and added. For example, in Jimeng, the new Seedance 2.0 adds functions such as video modification and image reference. At this time, new processing nodes can be added based on these new functions. That is, an AI model can be understood not only as a model with multiple functions, but also as a processing function corresponding to an AI model.

[0033] For example, a portion of the encapsulation process is shown: A. Configure the entry point through a unified node base class interface. All nodes inherit from NodeItem. The node's port and parameter area are defined within the node class.

[0034] For example: 1. TextVisionNode (text-to-image recognition) usage: - add_multi_input(“reference image”) - add_output(“text”) 2. GeminiAPINode (banana raw image) usage: - add_multi_input(“reference image”) - add_output(“image”) 3. VeoAPINode (veo live video) dynamically switches according to the mode: - add_input(“image”) or - add_input(“first frame image”) - add_input(“tail frame image”) And unify add_output(“video”) The node port layer has been standardized: - The image input node outputs the image path or a reference to the image content; - The text visual node outputs text; - The raw image node outputs the image; - Output video from the video node.

[0035] B. The node parameter area (i.e., the area in the worker node where feature parameters are configured) only exposes business parameters, not underlying protocol details: The `_setup_content()` method of each API node constitutes the node's parameter configuration area. This area configures business parameters that users can understand, rather than underlying HTTP fields.

[0036] For example: 1. The parameter area of ​​TextVisionNode (text-to-image recognition) includes: - Model_combo: gpt-5.4, gemini-3.1-pro-preview - Thinking Combo - Output format: format_combo - prompt_edit - Temperature temp_slider - Image sorting thumbnail_strip 2. The parameter area of ​​GeminiAPINode (banana raw image) includes: - model_combo - prompt_edit - Aspect Ratio Combo - resolution_combo - Reference image sorting thumbnail_strip 3. The parameter area of ​​VeoAPINode (veo raw video) includes: - Generation mode mode_combo (image-based video / first and last frames) - model_combo - Aspect Ratio ar_combo - Chinese to English enhancement_combo - Video super-resolution upsample_combo - prompt_edit 4. The parameter area of ​​JimengAPINode (i.e., Dream Video) includes: - Generation mode mode_combo - model_combo - prompt_edit - resolution_combo - Duration_combo - Random seed_spin These parameter areas are essentially "upstream configuration entry points for unified adaptation interfaces." What users fill in in the nodes is not the underlying request JSON, but general business parameters; when the nodes are executed, these parameters are then passed to the corresponding API class to be converted into real platform requests.

[0037] C. The API class is responsible for mapping node parameters to the target platform protocol. This layer is the closest to the actual code implementation of the "unified adaptation interface".

[0038] 1. TextVisionNode corresponds to TextVisionAPI (text-to-image recognition). The `get_params()` method of `TextVisionNode` outputs a unified parameter dictionary, for example: - prompt - model - format_mode - thinking_mode - image_order -temperature When the model is actually invoked, the node does not construct the HTTP request itself, but instead delegates it to the TextVisionAPI in api\text_vision_api.py.

[0039] TextVisionAPI internally switches based on model name: - If the model starts with gpt, then call _call_gpt_api(). - Otherwise, call _call_gemini_api() This is precisely the unified adaptation logic in the testdir program: The same "text image recognition node" appears as just one node on the interface, but at the underlying level it can be adapted to the GPT or Gemini protocol based on the model field.

[0040] Furthermore, the differences between the two protocols have already been resolved internally within the TextVisionAPI, for example: - GPT uses / v1 / chat / completions - Gemini uses / v1beta / models / {model}:generateContent - The GPT image field is image_url - Gemini image field is inline_data - GPT uses reasoning_effort (thinking mode) Gemini uses thinking-level thinking patterns. However, these differences are not visible to TextVisionNode. Therefore, TextVisionAPI is a unified adaptation interface implementation configured within TextVisionNode.

[0041] 2. GeminiAPINode corresponds to GeminiAPI (banana image) The get_params() method of GeminiAPINode outputs the following uniformly: - prompt - model - aspect_ratio - image_size (only appears with supported models) The GeminiAPI in api\gemini_api.py is responsible for: - Convert the reference image to base64 - Organize the reference diagrams in the nodes into contents.parts - Map node parameters to generationConfig.imageConfig - Determine if imageSize is supported based on model capabilities. - Parse the returned candidates, parts, and inlineData Finally, write the returned image to a local file. This means that the unified adaptation interface for the "banana raw image node" is not displayed separately in the node UI, but is completed by the combination of GeminiAPINode + GeminiAPI.

[0042] 3. VeoAPINode corresponds to Veo3API VeoAPINode provides unified node parameters: - mode - model - aspect_ratio - enhance_prompt - enable_upsample - prompt The Veo3API is responsible for converting these fields into the Veo request body. - prompt - model - enhance_prompt - enable_upsample - aspect_ratio - images In one implementation, the text-based image recognition node corresponds to the program class TextVisionNode, the model capability class TextVisionAPI, and the model function multimodal visual understanding + text generation. 1. Node Function It is used to read the input image and outputs the understanding result of the image. If the prompt requires "generate storyboard", "output storyboard script" or "decompose the image into shot description", then it outputs the storyboard text.

[0043] 2. Specific processing within the node TextVisionNode primarily performs the following processing: - Accepts multiple reference images as input, with the port name being "Reference Images"; - Allows users to set models, thinking modes, format modes, prompts, and temperatures; - Sort multiple input images using thumbnail bars; - Call TextVisionAPI.estimate_tokens() to estimate the cost; - During execution, parameters such as prompt, model, thinking_mode, format_mode, temperature, and image_order are passed to TextVisionAPI.

[0044] 3. This node does not use a fixed model at its underlying level; instead, it distributes traffic based on the selection from the dropdown menu. - When gpt-5.4 is selected, GPT's visual understanding and text generation capabilities are invoked; - When you select gemini-3.1-pro-preview, you can use Gemini's multimodal understanding and text generation capabilities.

[0045] In one implementation, the text display node corresponds to the program class: TextDisplayNode, and the corresponding function type: text receiving and display node. This node is typically used to receive the text results output by TextVisionNode and display them in the visual editing interface for easy viewing, confirmation, copying, or further modification.

[0046] In one implementation, the banana-shaped image node corresponds to the program class: GeminiAPINode, the model capability class: GeminiAPI, and the model function: reference graph-driven image generation / image editing generation. 1. Node Function This step is used to generate new image results based on "reference image + text prompts". In the workflow, this step combines the "original image" and the "storyboard text output by the text image recognition node" to continue generating storyboards that better meet the storyboard requirements.

[0047] 2. Specific processing within the node - Accepts multiple reference images as input; - Use thumbnail bars to maintain image order; - Read the prompt words entered by the user; - Set the model, aspect ratio, and resolution; - Pass these parameters to GeminiAPI to make the actual request.

[0048] 3. Which model function is used at the underlying level? The models available for selection in GeminiAPINode (different versions of Gemini) include: -gemini-3.1-flash-image-preview - gemini-3-pro-image-preview In one implementation, the Veo-generated video node corresponds to the program class: VeoAPINode, the corresponding model capability class: Veo3API, and the corresponding model function: image-generated video / first and last frame video generation. 1. Node Function Used to generate video results based on input images and cue words. The storyboard generated by GeminiAPINode will be used as input for VeoAPINode, and then combined with the storyboard text or shot cues to generate the final video.

[0049] 2. Specific processing within the node - Offers two modes: raw video and first / last frame; - Receive a single image, or receive the first and last frames of the image; - Set the model name; - Set the aspect ratio; - Configure whether to switch from Chinese to English; - Configure whether to enable video super-resolution; - Set prompt words; - Pass these parameters to Veo3API.

[0050] Further internal improvements to the Veo3API: - Convert the input image into a base64 data URL; - Assemble the request body, including prompt, model, aspect_ratio, images, enhance_prompt, and enable_upsample; - Submit the task by calling / v1 / video / create; - Get the task ID; - Poll the video generation status; - Returns the video_url upon completion; - Use the results for downloading, displaying, or saving.

[0051] 3. Which model function is used at the underlying level? The models that can be selected for this node include: -veo3.1 - veo3.1-fast - veo3.1-pro - veo3.1-4k - veo3.1-pro-4k - veo3 - veo3-pro In one implementation, the output node corresponds to the program class: OutputNode, and its function type is: result receiving and exporting node, without directly calling the large model. 1. Node Function It is used to receive video results generated upstream and to perform functions such as output display, status feedback and result saving.

[0052] 2. Specific processing procedure The node mainly performs the following tasks: - Receive video path; - Displays the result status; - Allows the entire workflow to be executed; - Use the project manager to write the project directory; - As a downstream result node, it closes the entire link.

[0053] 3. Function The final video is typically output from the Veo node to the OutputNode, where it is saved as the final video creation result file.

[0054] In one implementation, step S200, in response to a node editing command, configures the working parameters of the working node, including steps S210-S230: S210, In response to node editing instructions, configure the media material to be processed in the input node.

[0055] Optionally, through node editing commands, users can configure the input media materials to be processed, such as images or videos, in the input node as needed. This application embodiment uses images as an example for illustration. For example, the input node corresponds to the program class ImageNode, which can receive local image files, record image paths, generate thumbnail displays, and pass image paths or image references to subsequent worker nodes.

[0056] S220. In the processing node, configure the feature parameters based on the AI ​​model.

[0057] Optionally, the feature parameters include at least one of the following: the AI ​​model's thinking mode, prompt words, the format, style, aspect ratio, resolution, language conversion, and video super-resolution; wherein, the prompt words are the content input by the user, or are configured as the processing results automatically passed from the previous working node, that is, they can be obtained from the output of the previous working node.

[0058] For example, such as Figure 3 As shown, some AI models may have normal mode, deep thinking mode, etc. Users can select the corresponding thinking mode according to their needs, input prompt words or automatically extract prompt words from the output of the previous working node (the extraction template can be preset), select the format of the returned processing results (it can be a unified standard format so that the output can be applicable to the input of all working nodes), select style, aspect ratio, resolution, language conversion (e.g., Chinese to English), whether to perform video super-resolution, etc. The feature parameters of different AI models will be different and are pre-packaged for users to select accordingly.

[0059] In one implementation, in step S200, after configuring the working parameters of all working nodes, the connection relationships between input nodes, output nodes, and processing nodes are determined by connecting lines, thereby ultimately generating a directed acyclic graph (DAG) representing the video creation workflow. Figure 4 As shown, the exemplarily formed directed acyclic graph follows this order: 1. Input node (image upload): Upload media materials, such as images; 2. Processing nodes (text image recognition): The AI ​​model can be GPT-5.4 or Gemini. The input is the image of the input node, and the output is scene description text, character description text, and shot storyboard text. 3. Processing nodes (text-based image recognition): The input is the image of the input node and the prompt words (manual input or preset templates), and the output is scene description text, character description text, and shot storyboard text. 4. Processing nodes (banana raw images): The AI ​​model can be Gemini. The input is scene description text, character description text, and shot storyboard text. The output is structured storyboard JSON and raw image prompts. 5. Processing node (text image recognition): The input is structured storyboard JSON and raw image prompts; the output is a sequence of storyboard images and candidate storyboard images. 6. Processing node (Veo video), which can also be Keling or Jimeng, takes as input storyboard image sequence, candidate storyboard image, video prompt words, and structured storyboard JSON, and outputs as final video generation request package (video creation result), shot timing description, and transition control parameters. 7. Output node: Both input and output are the video creation results.

[0060] It should be noted that the output of the previous working node serves as the input of the next working node. This input can be used as prompt words and as the basis for processing by the AI ​​model. Prompt words can also be pre-inputted or configured based on preset templates, without specific limitations. In addition, the specific processing content of the processing node depends on user needs and can achieve standardized encapsulation of capabilities including but not limited to image-to-video generation, first and last frame video generation, image enlargement, ultra-high definition processing, partial redrawing, and text visual understanding. For example, GPT-5.4 and Gemini-3.1-pro-preview for text image recognition are used for image back-engineering and prompt word optimization. Prompt words for storyboard generation and prompt word optimization for tone maps are generated. Banana image generation (gemini-3.1-flash-image-preview and geimini-3-pro-image-preview) is used for image optimization. Text image generation is used for: generating storyboards and tone maps based on storyboards; and VEO video generation, which can generate videos flexibly, i.e., dream-generated videos. After the tone map is determined, the video is generated from the tone map and finally compiled into a video using editing software.

[0061] In one implementation, step S300 includes steps S310-S330: S310. In response to workflow execution instructions, the directed acyclic graph is parsed to determine the execution order between each work node.

[0062] Optionally, after inputting workflow execution instructions, the directed acyclic graph is automatically parsed to determine the content of each work node (such as media materials, feature parameters, pre-packaged content, etc.) and the execution order between each work node.

[0063] S320. According to the execution order, the media material of the input node is passed to the first processing node. The AI ​​model is called to process the media material through the first processing node and its feature parameters to determine the first result as the input of the second processing node. The AI ​​model is called to determine the second result as the input of the next processing node through the second processing node and its feature parameters until the last processing node outputs the video creation result.

[0064] Optionally, the media materials from the input node are automatically passed to the first processing node according to the execution order. Using the first processing node and its feature parameters, an AI model is invoked to process the media materials to determine the first result, which is then used as the input to the second processing node. Similarly, using the second processing node and its feature parameters, an AI model is invoked to determine the second result, which is then used as the input to the next processing node, until the final processing node outputs the video creation result. Figure 4The example shown here generates the video creation result for the last processing node (Veo video).

[0065] Optionally, using the first processing node and its feature parameters, an AI model is invoked to process the media material to determine a first result as input to the second processing node, including: S3201: Access the AI ​​model's calling address through the first processing node and perform authentication based on the authentication information.

[0066] like Figure 5 As shown, each processing node has a pre-encapsulated call address, allowing access to the AI ​​model based on that address. Each processing node's AI model is configured with corresponding authentication information, such as an API key, enabling automatic authentication by the processing node. It's worth noting that when configuring the API key for a processing node, a test connection can be triggered beforehand to verify the key's correctness and prevent issues during subsequent use.

[0067] S3202. Once authentication is successful, the AI ​​model is invoked based on the feature parameters of the first processing node to process the media material and obtain preliminary processing results.

[0068] Optionally, once authentication is successful, the first processing node can then invoke the AI ​​model based on its feature parameters to process the media material and obtain preliminary processing results.

[0069] S3203. According to the response parsing logic, the preliminary processing result is converted into a first result compatible with the next working node, and the first result is used as the input of the second processing node.

[0070] Optionally, the response parsing logic may include a compatible format. For example, a certain format may be configured in the feature parameters (ensuring compatibility among worker nodes during configuration) or a benchmark compatible format supported by all worker nodes may be pre-defined. In this case, the first processing node, according to the response parsing logic, converts the preliminary processing result into a first result compatible with the next worker node, which is then used as the output of the first processing node, and the first result is used as the input of the second processing node. It is understood that the processing steps of other processing nodes are similar and will not be elaborated further.

[0071] S330: Pass the video creation result to the output node, and output the video creation result through the output node.

[0072] Optionally, the last processing node (Veo video) generates the video creation result and passes it to the output node, so that the output node outputs and displays the video creation result.

[0073] It should be noted that the output node is pre-packaged with display and interactive functions, specifically: S3301. On the output node, display a thumbnail of the video creation result.

[0074] S3302, In response to the first interactive command on the thumbnail, call the media player to open the video creation result.

[0075] Optionally, a first interactive command can be generated by clicking a thumbnail. The output node responds to the first interactive command by calling a media player to open and play the video creation result. For example, clicking a thumbnail once plays the video creation result within the output node, and clicking a thumbnail twice calls the local default media player to play the video creation result. If the final output of the output node is an image, clicking the image will open it using the local default image player.

[0076] Alternatively, S3303, in response to a second interactive instruction on the thumbnail, saves the video creation result to a specified storage location or shares it to a third-party application.

[0077] Optionally, in response to a second interactive command on the thumbnail, such as a right-click menu selection or drag-and-drop, the video creation result can be saved to a specified storage location (e.g., a folder) or shared to a third-party application, such as directly sending it to social media, without needing to download, thus facilitating the use of the video creation result. Additionally, the output node provides a right-click menu function, allowing users to choose to execute the current output node or execute all work nodes, thus completing the video creation workflow.

[0078] In one embodiment, steps S410-S430 are also included: S410, in response to workflow execution instructions, generates creation tasks in the task queue that automatically execute video creation workflows.

[0079] Optionally, in response to a workflow execution instruction, a creation task for automatically executing a video creation workflow is generated in a task queue, the status of which reflects the execution status of the video creation workflow.

[0080] S420: Monitors and displays the task status of creation tasks in real time. The task status includes the execution status and progress information of each work node, and each work node provides an individual start and stop function.

[0081] Optionally, the task status of the creation task can be monitored and displayed in real time, or monitored and displayed at preset time intervals (e.g., 5 seconds). For example, the task status includes the execution status and progress information of each work node, such as execution status (in progress, execution completed) and progress information (e.g., percentage completed). In addition, each work node provides an individual start / stop function, which can control whether a work node is executed or not.

[0082] It should be noted that when a creation task is added to the task queue, several creation tasks can be performed simultaneously in the task queue. The task queue displays creation task time statistics and monitors several creation tasks at the same time. Clicking on the task queue or the corresponding creation task will display the corresponding task queue or creation task in a pop-up window.

[0083] S430: Record the connection time of the processing node to the AI ​​model, and record the feedback time from the processing node sending the request to call the AI ​​model to obtaining the preliminary processing result and the polling time from successfully connecting to the AI ​​model to outputting the video creation result. Timeout prompts are given for the connection time, feedback time and polling time respectively.

[0084] Optionally, the connection time of the processing node to the AI ​​model, i.e., the time of successful connection to the AI ​​model, is recorded. The time from sending the AI ​​model call request to receiving the initial processing result and the polling time from successfully connecting to the AI ​​model to outputting the video creation result are also recorded. Then, different timeout thresholds are set for the connection time, feedback time, and polling time for comparison. The timeout thresholds for connection time and feedback time can also be different for different AI models. For example, TextVisionAPI requires a timeout threshold of 10s for connection time and 30s for feedback time, while GeminiAPI requires a timeout threshold of 15s for connection time and 900s for feedback time, etc. When a time exceeds its corresponding timeout threshold, a timeout prompt is given, and a failure is returned to terminate the task. Therefore, wireless LED strip issues caused by slow network connections, prolonged server unresponsiveness, and excessively long result polling times can be avoided. It also prevents interface freezing and indefinite thread suspension, allowing the program to promptly report errors or retry in abnormal network environments.

[0085] In one implementation, this application embodiment can create a project based on the user's request in a visual editing interface, establish a standardized project directory structure in a specified storage path, and then create the above-mentioned creation task to execute the video creation workflow under the project. The video creation workflow configuration is automatically persisted and stored, and a 2-second anti-shake mechanism is used to avoid frequent IO operations. At the same time, it provides lifecycle management functions such as project list management, opening, and deletion. The final video creation result is automatically stored in the project directory, and it provides historical record management and version tracking of the final video generation result.

[0086] This application integrates AI models from different vendor platforms into a single application, improving video creation efficiency by over 300% and avoiding the operational costs associated with switching between multiple platforms. It also enhances flexibility: the node-based workflow allows for the arbitrary combination of different AI model capabilities, enabling complex creation logic that traditional single-platform solutions cannot handle, such as a fully automated chain from image generation to intelligent image expansion, video generation, and content understanding. Furthermore, the workflow can be saved as a template for repeated use, and the output of each work node can be cached and reused, avoiding the time and cost waste caused by repeated API calls. A unified experience is provided: all AI models use a consistent operation and interaction method, eliminating the need for users to learn operation methods on multiple different platforms, reducing learning costs by 80%. Batch processing is supported: a task queue mechanism allows for the batch submission of multiple generation tasks, with automatic background scheduling and execution, eliminating the need for manual intervention. Because each work node is an independently encapsulated modular design, users can add new AI model work nodes according to actual needs without modifying the core workflow engine code. A custom node extension mechanism is also supported, allowing third-party developers to write plugins to add new functional nodes. In addition to desktop implementation, it is also applicable to other operating environments such as web browsers and mobile devices.

[0087] Reference Figure 6 The diagram illustrates a structural block diagram of a multi-AI model video automatic creation device according to an embodiment of this application. The device may include: The first module is used to deploy several work nodes required for the video creation workflow in response to node deployment instructions in the visual editing interface; the work nodes include input nodes and output nodes for media materials, as well as at least two processing nodes for calling AI models from different vendors; The second module is used to respond to node editing instructions, configure the working parameters of the working nodes and the connection relationships between input nodes, output nodes and processing nodes, and form a directed acyclic graph representing the video creation workflow. The third module is used to respond to workflow execution instructions, automatically execute the video creation workflow based on the directed acyclic graph, and output the video creation results through the output node.

[0088] In one implementation, the third module is further configured to: In response to workflow execution instructions, creation tasks that automatically execute video creation workflows are generated in the task queue; The task status of creation tasks is monitored and displayed in real time. The task status includes the execution status and progress information of each work node, and each work node provides an individual start and stop function. Record the connection time of the processing node to the AI ​​model, and record the feedback time from the processing node sending the request to call the AI ​​model to obtaining the preliminary processing result and the polling time from successfully connecting to the AI ​​model to outputting the video creation result. Timeout prompts are given for the connection time, feedback time and polling time respectively.

[0089] In one implementation, the third module is further configured to: The output node displays a thumbnail of the video creation result; In response to the first interactive command on the thumbnail, the media player is invoked to open the video creation result; or, In response to a second interactive command on the thumbnail, the video creation result can be saved to a specified storage location or shared to a third-party application.

[0090] The functions of each module in the device of this application embodiment can be found in the corresponding description in the above method, and will not be repeated here.

[0091] Reference Figure 7 The diagram illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device includes a memory 310 and a processor 320. The memory 310 stores instructions that can be executed on the processor 320. The processor 320 loads and executes these instructions to implement the multi-AI model video automatic creation method described in the above embodiment. The number of memories 310 and processors 320 can be one or more.

[0092] In one embodiment, the electronic device further includes a communication interface 330 for communicating with external devices and exchanging data. If the memory 310, processor 320, and communication interface 330 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0093] Optionally, in a specific implementation, if the memory 310, processor 320 and communication interface 330 are integrated on a single chip, the memory 310, processor 320 and communication interface 330 can communicate with each other through an internal interface.

[0094] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-AI model video automatic creation method provided in the above embodiments.

[0095] This application also provides a chip, which includes a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the method provided in this application.

[0096] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.

[0097] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting the Advanced Reduced Instruction Set Computing (RISC) machine (ARM) architecture.

[0098] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0099] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0100] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0101] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0102] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0103] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0104] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0106] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for automatic video creation using multiple AI models, characterized in that, include: In the visual editing interface, responding to node deployment commands, several work nodes required for the video creation workflow are deployed; The working nodes include input nodes and output nodes for media materials, as well as at least two processing nodes for calling AI models from different manufacturers. In response to node editing instructions, the working parameters of the working nodes and the connection relationships between the input nodes, the output nodes and the processing nodes are configured to form a directed acyclic graph representing the video creation workflow; In response to the workflow execution command, the video creation workflow is automatically executed according to the directed acyclic graph, and the video creation result is output through the output node.

2. The video automatic creation method using multiple AI models according to claim 1, characterized in that: The response to the node deployment command, deploying the required work nodes for the video creation workflow includes: In response to the node deployment command, select the corresponding working node from the node list area of ​​the visual editing interface; Move the selected work node to the workflow deployment area of ​​the visual editing interface; Each processing node in the node list area is pre-encapsulated with a corresponding AI model version number, AI model call address, authentication information, and response parsing logic.

3. The video automatic creation method using multiple AI models according to claim 1 or 2, characterized in that: The configuration of the working parameters of the working node in response to the node editing command includes: In response to a node editing command, the media material to be processed is configured in the input node; In the processing node, feature parameters based on the AI ​​model are configured. The feature parameters include at least one of the following: the AI ​​model's thinking mode, prompt words, the format, style, aspect ratio, resolution, language conversion, and video super-resolution. The prompt words are either user-input content or the processing results automatically transmitted from the previous working node.

4. The video automatic creation method using multiple AI models according to claim 3, characterized in that: The automatic execution of the video creation workflow in response to the workflow execution command, based on the directed acyclic graph, and the output of the video creation result through the output node, includes: In response to the workflow execution command, the directed acyclic graph is parsed to determine the execution order between the various work nodes; According to the execution order, the media material of the input node is passed to the first processing node. The AI ​​model is called to process the media material through the first processing node and its feature parameters to determine the first result as the input of the second processing node. The AI ​​model is called to determine the second result as the input of the next processing node through the second processing node and its feature parameters until the last processing node outputs the video creation result. The video creation result is passed to the output node, and the video creation result is output through the output node.

5. The video automatic creation method using multiple AI models according to claim 4, characterized in that: The process of calling an AI model to process media materials using the first processing node and its feature parameters, and then using the AI ​​model to process the media materials to determine the first result as the input to the second processing node, includes: The first processing node accesses the AI ​​model's calling address and performs authentication based on the authentication information. Once authentication is successful, the AI ​​model is invoked based on the feature parameters of the first processing node to process the media material and obtain preliminary processing results; According to the response parsing logic, the preliminary processing result is converted into a first result compatible with the next working node, and the first result is used as the input of the second processing node.

6. The video automatic creation method using multiple AI models according to claim 5, characterized in that: The method further includes: In response to workflow execution instructions, creation tasks that automatically execute video creation workflows are generated in the task queue; The task status of the creation task is monitored and displayed in real time. The task status includes the execution status and progress information of each work node, wherein each work node provides an individual execution start and stop function. The connection time of the processing node to the AI ​​model is recorded, as well as the feedback time from the processing node sending the request to call the AI ​​model to obtaining the preliminary processing result and the polling time from successfully connecting to the AI ​​model to outputting the video creation result. Timeout prompts are given for the connection time, the feedback time and the polling time respectively.

7. The video automatic creation method using multiple AI models according to claim 1, characterized in that: The method further includes: On the output node, a thumbnail of the video creation result is displayed; In response to the first interactive command on the thumbnail, the media player is invoked to open the video creation result; or, In response to a second interactive command on the thumbnail, the video creation result is saved to a specified storage location or shared to a third-party application.

8. A video automatic creation device with multiple AI models, characterized in that, include: The first module is used to deploy several work nodes required for the video creation workflow in response to node deployment instructions in the visual editing interface; The working nodes include input nodes and output nodes for media materials, as well as at least two processing nodes for calling AI models from different manufacturers. The second module is used to respond to node editing instructions, configure the working parameters of the working node and the connection relationship between the input node, the output node and the processing node, and form a directed acyclic graph representing the video creation workflow; The third module is used to respond to workflow execution instructions, automatically execute the video creation workflow according to the directed acyclic graph, and output the video creation result through the output node.

9. An electronic device, characterized in that, include: A processor and a memory, wherein instructions are stored in the memory and loaded and executed by the processor to implement the method as claimed in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program therein, which, when executed, implements the method as described in any one of claims 1-7.