Video editing method, device, storage medium, and program product

By downloading and modifying video templates on local terminal devices and adjusting rendering rules and positional relationships using video editing software, the problem of high cost and low efficiency in existing video production technologies has been solved, enabling efficient and low-cost diversified video production.

CN120186431BActive Publication Date: 2026-06-02BEIJING 58 INFORMATION TTECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING 58 INFORMATION TTECH CO LTD
Filing Date
2025-04-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current video production technologies require multiple complex steps, rely on specialized equipment and personnel, resulting in high costs and low efficiency, and failing to meet diverse video needs.

Method used

By downloading video templates to local terminal devices and using video editing software to modify the templates, adjust rendering rules and positional relationships, secondary editing of the video can be achieved.

Benefits of technology

It has improved the efficiency and diversity of video production, reduced the reliance on professional equipment and human intervention, and met the diverse video needs of high efficiency and low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186431B_ABST
    Figure CN120186431B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video editing method, device, storage medium and program product. In the embodiments of the present application, a first video template for describing rendering rules and hierarchical relationships of a group of video materials is downloaded to a local terminal device and imported into a video editing software; the first video template is opened in an editing area of the video editing software, and a first video generated based on the first video template is displayed in a preview area of the video editing software; further, by responding to a modification operation on the first video template, modified rendering rules in any target materialization template and / or modified position relationships between any two target materialization templates are determined, and the first video is edited according to the modified rendering rules and / or the modified position relationships to obtain a second edited video, which can improve the efficiency and diversity of video production, reduce the dependence on professional equipment and manual intervention, and meet the diversified video demand of high efficiency and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a video editing method, device, storage medium, and program product. Background Technology

[0002] With the rapid development of digital media technology and the diversification of content consumption demands, video production plays a crucial role in numerous fields such as advertising, film and television creation, online education, and short video platforms. For example, in advertising production, engaging video ads can be created based on product characteristics and target audiences, thereby increasing brand awareness and product sales.

[0003] In the existing technology, in order to meet flexible and diverse video needs, video production requires multiple complex steps such as shooting, post-editing, and adding special effects. Each step requires professional equipment and personnel, as well as a lot of time and effort to complete. As a result, video production costs are high and production efficiency is low, which cannot meet the ever-increasing video demand. Summary of the Invention

[0004] Embodiments of this application provide a video editing method, device, storage medium, and program product to improve the efficiency and diversity of video production, reduce reliance on professional equipment and manual intervention, and meet the diverse video needs of high efficiency and low cost.

[0005] This application provides a video editing method, including: downloading a first video template to a local terminal device. The first video template is a video template for rendering a set of video materials to generate a first video. The first video template includes: a target material template adapted to the material type in the set of video materials from a variety of material templates. The various material templates are obtained by varying the initial video template. Each material template is used to describe the rendering rules of a video material. The first video template is used to describe the rendering rules and hierarchical relationship of a set of video materials. The hierarchical relationship is reflected in the positional relationship between the target material templates in the first video template. Running video editing software on the terminal device, importing the first video template into the video editing software, opening the first video template in the editing area of ​​the video editing software, and displaying the first video generated based on the first video template in the preview area of ​​the video editing software. In response to the modification operation of the first video template, determining the modified rendering rules in any target material template and / or the modified positional relationship between any two target material templates, and editing the first video according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video.

[0006] This application also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is able to implement the various steps in the video editing method provided in this application.

[0007] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the various steps in the video editing method provided in this application.

[0008] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, enable the processor to implement the various steps in the video editing method provided in this application.

[0009] In this embodiment, a first video template describing the rendering rules and hierarchical relationships of a set of video materials is downloaded to a local terminal device, and the first video template is imported into video editing software running on the terminal device. The first video template is opened in the editing area of ​​the video editing software, and a first video generated based on the first video template is displayed in the preview area of ​​the video editing software. Further, by responding to the modification operation of the first video template, the modified rendering rules in any target material template and / or the modified positional relationship between any two target material templates are determined, and the first video is edited according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video. Based on the existing first video template, the second video is edited by combining the localized video editing software, which improves the efficiency of video production, reduces the dependence on professional equipment and manual intervention, and thus reduces costs. Furthermore, since the first video template is obtained by personalized combination of the video materials used to generate the video, the combined first video template has a corresponding relationship with the first video, which enriches the types of video templates to a certain extent, thereby meeting the diverse video needs of generated video content. Attached Figure Description

[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0011] Figure 1 A flowchart illustrating a video editing method provided for an exemplary embodiment of this application;

[0012] Figure 2 An interactive schematic diagram of a video batch generation method provided as an exemplary embodiment of this application;

[0013] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.

[0016] Additionally, it should be noted that when user interaction operations or triggering operations are involved in the embodiments of this application, these operations include, but are not limited to, various interaction methods such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations. Touch operations include, but are not limited to, click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations. Swipe operations include, but are not limited to, straight-line swipes and curved-line swipes.

[0017] Furthermore, it should be noted that, in the cases where the embodiments of this application involve jumping between the first interface and the second interface, the jumping methods involved in the embodiments of this application include, but are not limited to: jumping directly from the first interface to the second interface, or jumping from the first interface to the task interface and completing the corresponding task operation on the task interface before jumping to the second interface; completing the corresponding task operation on the task interface includes, but is not limited to: completing the game operation on the game interface when the task interface is implemented as a game interface; completing identity authentication on the identity authentication interface when the task interface is implemented as an identity authentication interface; completing the recharge operation on the recharge interface when the task interface is implemented as a recharge interface; and so on.

[0018] In the existing technology, in order to meet flexible and diverse video needs, video production requires multiple complex steps such as shooting, post-editing, and adding special effects. Each step requires professional equipment and personnel, as well as a lot of time and effort to complete. As a result, video production costs are high and production efficiency is low, which cannot meet the ever-increasing video demand.

[0019] To address the problems existing in the prior art, in this embodiment, a first video template describing the rendering rules and hierarchical relationships of a set of video materials is downloaded to a local terminal device, and the first video template is imported into video editing software running on the terminal device. The first video template is opened in the editing area of ​​the video editing software, and a first video generated based on the first video template is displayed in the preview area of ​​the video editing software. Furthermore, by responding to modification operations on the first video template, the modified rendering rules in any target material template and / or the modified positional relationship between any two target material templates are determined, and the first video is edited according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video. Based on the existing first video template, the second video is edited using localized video editing software, which improves the efficiency and diversity of video production, reduces reliance on professional equipment and manual intervention, thereby reducing costs. Furthermore, since the first video template is obtained by personalized combination of the video materials used to generate the video, the combined first video template corresponds to the first video, enriching the types of video templates to a certain extent, thus meeting the diverse video needs of generated video content. The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0020] Figure 1 This is a flowchart illustrating a video editing method provided for an exemplary embodiment of this application. Figure 1 As shown, the method includes:

[0021] S11. Download the first video template to the local terminal device. The first video template is a video template used to render a set of video materials to generate the first video. Among multiple material templates, there is a target material template that matches the material type in a set of video materials. Multiple material templates are obtained by changing the initial video template. Each material template is used to describe the rendering rules of a video material. The first video template is used to describe the rendering rules and hierarchical relationship of a set of video materials. The hierarchical relationship is reflected in the positional relationship of the target material templates among the first video templates.

[0022] S12. Run the video editing software on the terminal device, import the first video template into the video editing software, open the first video template in the editing area of ​​the video editing software, and display the first video generated based on the first video template in the preview area of ​​the video editing software;

[0023] S13. In response to the modification operation on the first video template, determine the modified rendering rules in any target material template and / or the modified positional relationship between any two target material templates, and edit the first video according to the modified rendering rules and / or the modified positional relationship to obtain the edited second video.

[0024] In this application embodiment, the specific form or deployment method of the local terminal device is not limited. Any device with video processing and editing capabilities and capable of implementing video editing methods can be used as the local terminal device in this application embodiment. For example, the local terminal device can be a smartphone or tablet computer running a mobile operating system (such as an iOS / Android device) that completes editing operations through touch interaction; it can also be a desktop computer or workstation equipped with a professional GPU (Graphics Processing Unit) that supports 4K / 8K high-resolution video processing; or it can be an embedded device (such as a smart TV, in-vehicle entertainment system, or drone control terminal).

[0025] In this embodiment, the storage location of the first video template is not limited. Any cloud or edge computing node with storage resources can store the first video template in this embodiment. For example, the first video template can be stored in a CDN (Content Delivery Network) network, and downloading the first video template from the CDN network can reduce transmission latency; it can also be stored on the server, and the first video template can be downloaded from the server; it can also be stored on an edge computing node, and network slicing technology can be used to ensure the transmission quality of the first video template.

[0026] In this embodiment, the first video template is a video template used to render a set of video materials to generate a first video. The first video refers to the video obtained by rendering a set of video materials according to the rendering rules provided by the first video template. A set of video materials includes video materials of multiple material types. Material type refers to the type of different video elements (i.e., video materials) that constitute the video content, including but not limited to: audio, video, background image, subtitles, digital humans, etc. For example, a set of video materials may include a background video, an audio clip of a digital human narrating, a subtitle layer for display, and a background image.

[0027] In this embodiment, the first video template includes a target material template that matches the material type of a set of video materials from among multiple material templates. The target material template refers to a material template that matches the material type of a set of video materials, obtained by combining multiple material templates corresponding to multiple video materials according to a combination logic. The combination logic includes a hierarchical relationship, which refers to the layer stacking order of multiple video materials during the video generation process, used to determine the front-to-back overlapping relationship and occlusion logic of multiple video materials in video generation. For example, in generating a video, subtitles can be located on a relatively upper layer in the layers, the background image can be located on a relatively lower layer in the layers, and the digital human can be superimposed on the upper layer of the background image and located below the subtitle layer.

[0028] In addition, the first video template also includes at least one access link for video material. This access link refers to the network address or storage path used to locate and retrieve the video material from the CDN network during the rendering process. For example, when the video material type is a digital human, the access link could point to a digital human stored in the CDN network, such as https: / / model-cdn.example.com / avatar_001.glb.

[0029] In this embodiment, multiple material templates refer to templates obtained by variableizing an initial video template. The initial video template refers to an existing video production framework. This framework can be pre-defined or exported from video editing software. Regardless of the method, the initial video template includes rendering rules for various video materials required to generate the video, as well as the hierarchical relationships between these materials. It should be noted that the initial video template does not contain the video materials themselves, but rather the rendering rules and hierarchical relationships. The rendering rules correspond to the types of video materials, and the hierarchical relationships describe the display order and occlusion logic between different video materials.

[0030] Optionally, the initial video template can be implemented as an MLT (Media Lovin' Toolkit) template. An MLT template is an open-source framework for multimedia processing. It is an XML-based file that records all parameters for video editing, such as video clips on the timeline, audio tracks, filter effects, transitions, etc., and can be used for video editing.

[0031] In this embodiment, the variable processing is partly reflected in the structural decomposition and parameter information templating of the initial video template based on the material type as the splitting dimension to obtain a material-based template. There is a correspondence between the material-based template and the material type. For example, material types include audio, video, background images, subtitles, digital humans, etc. The corresponding material-based template types include, but are not limited to: audio material-based templates, video material-based templates, background image material-based templates, subtitle material-based templates, digital human material-based templates, etc.

[0032] In this embodiment, each material template describes the rendering rules for a type of video material. A first video template describes the rendering rules and hierarchical relationships of a set of video materials. The hierarchical relationship is reflected in the positional relationship between target material templates within the first video template. Rendering rules control the presentation of a video material in the generated video, or the rendering result. Rendering rules define the presentation method of the video material through combinations of rendering parameter information. Rendering parameter information refers to specific parameters related to the rendering of the video material. Rendering parameter information can describe at least one attribute of a material file of a certain material type. Each attribute can be used as a rendering parameter, and the attribute value of each attribute in the initial video template can be used as the default parameter value for the corresponding rendering parameter. Rendering parameters with default parameter values ​​form a set of rendering parameter information corresponding to the material file of that material type. "Material file" refers to various resource files used in the material template, mainly including various video materials, such as images, videos, and audio. Rendering parameter information can control the rendering logic of the material file so that the material file can achieve the desired rendering effect in the final generated video. Rendering parameter information quantifies and describes the attributes of the source files, enabling the conversion of video source file attributes into rendering parameter attributes. This rendering parameter information controls the rendering effect of the video source files in the final generated video. In other words, the attributes of the video source files are described by the rendering parameter information as corresponding rendering parameters, and rendering parameters with default values ​​define the attribute values ​​of a certain attribute in the source file.

[0033] The rendering parameter information can be extracted from the information fragments corresponding to each of the various material types. The rendering parameter information for each material type includes at least one rendering parameter and its corresponding default parameter value. Each rendering parameter describes the attributes of a material file of a certain material type. In this embodiment, the specific attributes corresponding to the rendering parameters for each of the various material types are not limited.

[0034] The hierarchical relationship refers to the positional relationship between the target material templates and the first video template. Specifically, it refers to the layer stacking order of different material templates during video generation. For example, in a video, the subtitle material template might be set to the top layer, the background image material template to the bottom layer, and the digital human material template to the middle layer. This positional relationship determines the display order and occlusion logic of multiple video materials in the final video; that is, the top layer material will cover the bottom layer material, while the middle layer material may partially occlude the bottom layer material and be occluded by the top layer material. In this way, the hierarchical relationship ensures the reasonable presentation of video content, allowing various elements in the video to overlap and interact with each other in the expected manner.

[0035] In this embodiment, video editing software refers to an application that can run on a terminal device and has a video editing area and a preview area. The video editing area is an interface for importing, editing, adjusting, and combining video materials or templates. A first video template can be imported into the video editing software, opened in the editing area, and modified. For example, the order of video materials can be adjusted, rendering rules modified, or layer relationships changed. The preview area is a visual interface that displays the real-time results of the current editing operation. It displays the first video generated based on the first video template. For example, after a user modifies parameters in the editing area, the preview area immediately displays the effect of a second video generated based on the updated first video template, ensuring the user can instantly verify the editing results. This embodiment does not limit the specific form of the video editing software. For example, the video editing software can be open-source and free software, such as Shotcut, which supports basic editing on multiple platforms; it can also be other custom-developed dedicated editing tools, such as enterprise-level video production platforms or industry-customized software.

[0036] In this application embodiment, the video editing software is implemented as Shotcut, and the specific implementation method for editing the first video is described. The method of editing the first video template can modify the rendering rules involved in the target material template and / or the positional relationship between any two target material templates as needed. This will be described in detail below through embodiments.

[0037] Option 1: The method of editing the first video template can be customized by adjusting the rendering rules involved in the target material template as needed. Specifically, this can be achieved by customizing the rendering parameter information of the video material to be modified. For example, customizing the rendering parameter information of the video material can involve adjusting the attribute values ​​of the video material.

[0038] In one optional embodiment, custom adjustments to the rendering parameters of video footage can be made to the attribute values ​​of specific attributes related to the modified video footage type in the first video template, in order to optimize image quality, performance, or adapt to specific needs. Specifically, Shotcut can provide multiple candidate parameter values ​​for various video footage attributes, and one of the candidate parameter values ​​can replace the original attribute value of the specific video footage type. For example, for video footage, the resolution attribute value can be adjusted from "1080p" to "720p," and the brightness attribute value can be replaced from "50" to "60."

[0039] Option 2: The method of editing the first video template can be customized according to the positional relationship between any two target material templates that are to be modified. Modifying the positional relationship between any two target material templates can be achieved by adjusting the hierarchical relationship between any two video clips, or by changing the playback order of any two video clips.

[0040] In an optional embodiment, the editing of the first video template can also involve modifying the hierarchical relationship between any two video clips. Adjusting the hierarchical relationship refers to controlling the overlapping relationship and presentation priority of different clips in a multi-track timeline by changing the track position or stacking speed of the clips. In this embodiment, the method of adjusting the hierarchical relationship is not limited. For example, Shotcut can be used to adjust the timeline hierarchy. Each video track on the timeline (also known as the timeline, a tool that connects different video clips in chronological order) forms an independent hierarchy. You can reorder the tracks by dragging the track labels. The content of the top track always covers the content of the lower track. For example, you can move the subtitle track above the video track to ensure that the subtitles are always visible. You can also use Shotcut to adjust the timeline through masking and compositing. In the timeline, newly added video clips will cover the area of ​​previous video clips by default. You can select "promote" or "demote" the video clips, or directly drag the clips above / below the target track to adjust the priority. You can also use Shotcut to adjust the timeline hierarchy by controlling the keyframe animation. For dynamic effects containing keyframes (such as scaling and rotation), adjusting their track positions can ensure that the animation effects interact with other video clips during compositing.

[0041] In another optional embodiment, the editing of the first video template can also be done by adjusting the playback order of any two video clips. Adjusting the playback order of video clips involves changing their arrangement within the same track to reorganize the narrative logic or timeline structure of the video. For example, moving a dialogue audio clip before the background audio adjusts the playback order of two video segments; or adjusting the playback order of multiple video clips.

[0042] In this embodiment, the first video is edited based on modifications to the rendering rules and / or the modified positional relationships of the first video template to obtain an edited second video. A second video, with different visual effects or structures, is regenerated based on the user's modifications to the rendering rules and / or the modified positional relationships of the first video. For example, in the original first video, the subtitles are located at the bottom layer and are obscured by the background image, resulting in unclear subtitles. By modifying the positional relationships, the subtitles are moved to the upper layer, making them more clearly visible, thus generating the second video.

[0043] In one optional embodiment, one way to determine the modified rendering rules in any target material template in response to a modification operation on the first video template includes: displaying the target material template to be modified in the editing area in response to a positioning operation on the first video template; and obtaining the modified rendering rules in response to a code editing operation in any target material template. The positioning operation involves the system loading the corresponding target material template into the editing area when the user selects any material type to be modified through interface interaction, thus visualizing the rendering rules of the current target material template. After loading the corresponding target material template into the editing area, the current rendering rules of the target material template can be seen in the editing area. For example, if the video material type is subtitles, the target material template may include the font size, color, transparency, etc., of the subtitles. The rendering rules can be adjusted by directly modifying the attribute values ​​of the corresponding attributes of the video material in the code or rendering parameter information. For example, if the target material template is a background image, the user can directly change "filter: none" to "filter: blur(2px)" in the first video template code, which means changing the background image from clear to slightly blurred to highlight the foreground content (such as digital figures or subtitles); or, the user can change the code value of the background image transparency from "0.5" to "0.8" in the editing area to make the background image clearer and emphasize the background image content.

[0044] Optionally, another implementation of determining the modified rendering rules in any target material template in response to a modification operation on the first video template includes: displaying the first editing window in response to a trigger operation that modifies any target material template in the first video template; determining the name of any target material template and the modified rendering rules in response to a first editing operation in the first editing window; and refreshing the first video template in response to a first edit submission operation based on the name of any target material template and the modified rendering rules. The trigger operation is used by the user to initiate editing. For example, double-clicking the background image on the timeline or right-clicking and selecting "Edit Parameters". The first editing window is the interface in the video editing software used to modify the target material template. It may include input boxes for inputting the corresponding attributes or attribute values ​​of the video material in the rendering parameter information, as well as sliders or drop-down menus. For example, the editing window for a target material template whose video material type is a background image may display a transparency adjustment slider. The first editing operation is the specific modification operation performed by the user on the target material template in the first editing window, thereby determining the name of the modified target material template and the modified rendering rules. For example, dragging the slider changes the opacity from 0.5 to 0.8, setting the target template name to "Background Image Template" and the modified rendering rule to "opacity: 0.8". The first edit submission is the user's confirmation of changes after completion. For example, the first edit submission could be clicking the "Save" or "Submit" button. Refreshing refers to updating the rendering parameters of the first video template based on the target template name and the modified rendering rules, and updating the preview of the edited video effect in real time.

[0045] For example, a user is using video editing software to edit a first video template that includes a background image, subtitles, and a digital human. The user wants to adjust the display effect of the subtitles to make them more eye-catching. The specific steps are as follows: The user selects the target material template for the subtitles in the video editing software, and the system detects this trigger. The system displays a first editing window, which contains the name of the subtitle material template (e.g., "Subtitle Template 1") and the current rendering rules (e.g., white font color, 24px font size). The user modifies the subtitle rendering rules in the first editing window, changing the font color to yellow, the font size to 32px, and adding a shadow effect. The system detects the user's editing operation, confirms the name of the target material template as "Subtitle Template 1," and records the modified rendering rules (font color, size, and shadow effect). The user clicks the "Save" button to submit the changes. The system refreshes the first video template according to the modified rendering rules, and the subtitle display effect is updated accordingly, resulting in the second video. Through the above operations, the user successfully adjusts the subtitle display effect and generates a second video that meets the requirements. This process enables efficient and accurate video template modification through individual video material templates and real-time feedback that refreshes immediately after submission. Users can intuitively adjust rendering rules through a graphical interface without manually modifying the underlying code, improving the efficiency and diversity of video production, reducing reliance on professional equipment and manual intervention, and meeting diverse video needs at high efficiency and low cost.

[0046] In one optional embodiment, one way to determine the modified positional relationship between any two target material templates in response to a modification operation on the first video template includes: in response to a drag operation on any target material template in the first video template, moving any target material template according to the drag trajectory of the drag operation, and determining another target material template according to the position when the drag operation ends; filling the structural position of the other target material template into the structural position of the first target material template, and filling the structural position of the other target material template into the structural position of the first target material template, to obtain the modified positional relationship between any two target material templates. Here, a drag operation refers to a user dragging a target material template in the editing area. For example, a user drags any target material template using a mouse or touch gesture to adjust its position. The system moves the visual position of the target material template in real time according to the drag trajectory. When the user's drag operation ends, the system can determine another selected target material template based on the final termination of the drag operation. For example, when a user drags a background image template in the editing area, when the user's drag operation ends, another selected target material template, such as a subtitle template, is determined based on the final release position.

[0047] In this context, a structural bit describes the position of the corresponding target material template within the first video template. It's a pre-defined placeholder within the target material template, used to identify where the target material template can be inserted. Each structural bit corresponds to the filling of a target material template for a specific material type. The positional relationship between multiple structural bits can reflect the hierarchical relationship between various video materials. The information of a structural bit can include its number or name, and it determines the occlusion logic of the target material template. For example, the structural bit information might be: layer_1 (background layer) is at the bottom; layer_2 (middle layer) covers it; and layer_3 (foreground layer) is at the top.

[0048] After determining any two target template materials, one template can be filled into the structural position of the other, and vice versa, thus swapping the structural positions of the two template materials to obtain the modified positional relationship between them. This positional relationship refers to the hierarchical relationship or track order of the two template materials in the video, which determines their occlusion logic. Ultimately, the hierarchical or track positions of the two template materials are reversed, thereby altering their occlusion logic in the final generated second video. For example, the occlusion logic could be such that one template material overlaps the other.

[0049] For example, a user is using video editing software to create a first video template containing a background image, subtitles, and a digital human. The user wants to place the background image on top of the subtitles. The specific steps are as follows: The user drags the selected target template, designated "Background Image Template," upwards. When the user drags the Background Image Template to the location of the "Subtitle Template" and releases the mouse, the system identifies another target template, designated "Subtitle Template." The system determines the structural information of the Background Image Template as Layer 1 and the Subtitle Template as Layer 2, swapping the Layer 1 and Layer 2 information. This results in a modified positional relationship between the two target templates: the new structural information for the Background Image Template is now "Layer 2," and the new structural information for the Subtitle Template is "Layer 1." Ultimately, in the generated second video, the background image will cover the subtitles, causing the subtitles to be obscured. This process, through a visual drag-and-drop operation, allows users to intuitively swap the structural positions of the target templates, thereby modifying their positional relationship. This interactive method simplifies the steps of complex layer adjustments, allowing users to quickly reverse the layer order without manually editing code, improving video production efficiency, reducing reliance on professional equipment and manual intervention, and meeting the needs of high-efficiency, low-cost video production.

[0050] Optionally, another implementation of determining the modified positional relationship between any two target material templates in response to a modification operation on the first video template includes: displaying a second editing window in response to a trigger operation that adjusts the positional relationship between any two target material templates in the first video template; determining the names of any two target material templates and their swapped structural positions in response to a modification operation in the second editing window; and refreshing the first video template based on the names of any two target material templates and their swapped structural positions in response to a modification submission operation, to obtain the modified positional relationship between any two target material templates. The second editing window can display the names of all target material templates and their current structural positions. Furthermore, the second editing window includes a drop-down menu or checkboxes for the user to select all target material templates that can be swapped. Additionally, the second editing window includes an input box for the user to add new target material templates.

[0051] The specific modification method for determining the names of any two target material templates and their swapped structural positions in the second editing window is not limited. For example, a visual modification mode can be used by utilizing a drop-down menu or checkbox in the second editing window that allows the user to select all target material templates for swapping, thereby determining the names of the two target material templates and their swapped structural positions. Alternatively, a manual modification mode can be used by utilizing an input box in the second editing window that allows the user to add new target material templates, thereby determining the names of any two target material templates and their swapped structural positions. In this embodiment, the process of refreshing the first video template in response to the modification submission operation, based on the names of any two target material templates and their swapped structural positions, to obtain the modified positional relationship between the two target material templates, has been described in detail in the foregoing embodiments and will not be repeated here.

[0052] For example, a user might want to move the subtitle template from Layer 2 to Layer 1, covering the background image. The user can use the visual mode to select both the "Background Image Template" and the "Subtitle Template" in the editing interface and click the "Swap Positions" button. The second editing window, as shown below, displays the names of all target template materials and their current structural positions:

[0053] Target material template Information of current structure bits Background image template Layer 1 Subtitle template Layer 2 Digital Human Template Layer 3

[0054] Users can select "Subtitle Template" and "Background Image Template" and then choose the "Swap Position" function from the drop-down menu.

[0055] Optionally, users can also manually input the name of the target template A: Subtitle Template → New Structure Position: Layer 1, and the name of the target template B: Background Image Template → New Structure Position: Layer 2. The system will determine that the target templates to be modified are the Subtitle Template and the Background Image Template, and their new structure positions: Subtitle → Layer 1 and Background Image → Layer 2.

[0056] After the user clicks "Submit," the hierarchical relationship in the first video template is updated to generate the second video, where the subtitle template is located on layer 1 (top layer), covering the background image template (layer 2). This process, through the interactive design of the second editing window, not only supports the selection of a visual mode but also allows manual input, enabling users to flexibly adjust the structural positions of the two target material templates, thereby changing their positional relationship. This design balances intuitiveness and flexibility, making it suitable for complex hierarchical management scenarios. It can accurately update the first video template, ensuring that the modified hierarchical relationship takes effect in real time.

[0057] In one optional embodiment, downloading the first video template to the local terminal device includes: in response to an access operation to the batch video generation result page, displaying the batch video generation result page, which displays access links on the CDN network or server for each of the N video instance identifiers corresponding to the target video templates, access links on the CDN network or server for each of the N video instance identifiers corresponding to a set of video materials, and access links on the CDN network or server for each of the N video instance identifiers corresponding to the videos; where N is an integer ≥ 2; and in response to a trigger operation on the access link of any target video template on the CDN network or server, downloading any target video template as the first video template from the CDN network or server to the local terminal device. The batch video generation result page is a page that centrally displays the batch-generated videos and their related information.

[0058] In this embodiment, the N video instance identifiers are unique identifiers generated based on the batch generation task. Each of the N video instance identifiers is distinct, as each identifier can represent a video to be generated, and each identifier has its own corresponding target video template. Therefore, for each of the N video instance identifiers, the batch video generation result page can display access links for the target video templates corresponding to each of the N video instance identifiers on the CDN network or server.

[0059] In this process, each video to be generated has a set of associated video materials. These materials are used to generate the video corresponding to that video instance identifier. Each video instance identifier serves to track the generation process of the video to be generated; in other words, it also acts as a unique identifier for the video during the generation process, corresponding to the generation of a single video. Therefore, for N video instance identifiers, the batch video generation results page can display the access links for the set of video materials corresponding to each of the N video instance identifiers on the CDN network or server, as well as the access links for the videos corresponding to each of the N video instance identifiers on the CDN network or server.

[0060] The method for generating the N video instance identifiers is not limited. For example, it can include, but is not limited to, numbers and strings. For instance, it can start from any integer and form an increasing sequence with a fixed step size to obtain N integers as the N video instance identifiers; or it can be a preset set of N strings, etc.

[0061] In one optional embodiment, for any video instance identifier, when the various video materials corresponding to the video instance identifier include target subtitles, target audio, target digital human green screen video and target background image, the access links of the target subtitles, target audio, target digital human green screen video and target background image on the CDN network or server, the access links of the video corresponding to the video instance identifier and the target video template corresponding to the video instance identifier on the CDN network or server are displayed on the batch video generation result page.

[0062] For example, when a user clicks the "View Generation Results" button, they are taken to the batch video generation results page. The page displays three columns of information for three video instances:

[0063]

[0064] In an optional embodiment, as shown above, the batch video generation results page also includes access links to the video materials. Based on this, in addition to local secondary editing of the first video, multiple video materials can be downloaded from a CDN network or server to the local terminal device for local synthesis of a new video.

[0065] In one optional embodiment, in response to a triggered operation of accessing multiple video materials on a CDN network or server, multiple video materials are downloaded from the CDN network or server to a local terminal device; wherein, the multiple video materials are distributed in one or more groups of video materials; video editing software on the terminal device is run, the multiple video materials are imported into the video editing software, and the multiple video materials are displayed in the editing area of ​​the video editing software; in response to the editing operation of the multiple video materials, a third video is generated, and a second video template corresponding to the third video is generated according to the attribute information and hierarchical relationship of the multiple video materials after editing, the second video template includes the rendering rules of the multiple video materials and the positional relationship of the multiple video materials in the second video template.

[0066] Users can trigger actions to download multiple video clips from the CDN network or server to their local terminal devices. The downloaded video clips can be distributed across one or more groups, and users can select clips from the same or different groups to download.

[0067] For example, the first set of video footage includes:

[0068]

[0069] The second set of video footage includes:

[0070]

[0071] Users can choose from the following video materials: background images from the first group (e.g., wedding-themed backgrounds such as churches, gardens, starry skies) and subtitle templates (e.g., preset wedding subtitle styles with dynamic "Forever & Always" text effects) to download as various video materials corresponding to the third video from the CDN network; or they can choose from the following video materials: background images from the first group (e.g., wedding-themed backgrounds such as churches, gardens, starry skies) and audio materials from the second group (e.g., wedding background music such as piano pieces, symphonies, pop song clips) and subtitle templates (e.g., romantic subtitle animations such as petal falling effects) to download as various video materials corresponding to the third video from the CDN network. In this embodiment, users can freely choose different groups of video materials to combine according to their creative needs, without being restricted by video material grouping, thus improving the flexibility of video generation. Furthermore, the CDN network or server provides a rich variety of video material types (video, audio, background images, subtitles, etc.), which users can freely combine to create unique video content. This flexible video material download and combination function allows for efficient acquisition of required video materials from the CDN network or server, and personalized creation on local terminal devices, meeting diverse video production needs.

[0072] After downloading, run the video editing software on your terminal device. You can then import the selected video footage and display and edit it within the editing area. After editing these footage, a new video will be generated; the final video file generated from the edited footage is the third video. Simultaneously, based on the edited footage's attribute information and hierarchical relationships, the system will generate a corresponding second video template. This second video template contains the rendering rules for the multiple video footage and their positional relationships within the template, allowing for the rapid generation of similar videos later.

[0073] In this embodiment, the specific implementation principle of generating a third video in response to editing multiple video materials is the same as that of editing the first video in response to modifying the first video template in the foregoing embodiments to obtain an edited second video, and will not be repeated here.

[0074] Based on the attribute information and hierarchical relationships of multiple edited video clips, a second video template corresponding to the third video is generated. The system automatically records the rendering rules and positional relationships of the third video after editing, forming a reusable second video template for quickly generating similar videos in the future.

[0075] Example scenario: A user creates multiple product advertising videos, each with a different background and subtitles, but sharing a unified animation effect.

[0076] First, users download the following video content types from the CDN network:

[0077] Background group: Background 1.mp4 (beach), Background 2.mp4 (city) Subtitle Group: Product A Subtitles.srt, Product B Subtitles.srt Animation Team: Product Showcase Animation.mp4 (Same animation template)

[0078] Users import three video clips into the video editing software, drag "Background 1.mp4" to the bottom of the timeline, add "Product A subtitle" to the top, and set the animation transparency to 70%. The result is an advertising video featuring a beach background, Product A subtitle, and a semi-transparent animation. The system saves the edited settings (such as background transparency and subtitle position) to create a new template, the second video template. This process combines distributed material acquisition with local editing to generate video templates, achieving efficient and flexible video production: users can flexibly access different types of video materials from the CDN network, create personalized combinations to obtain the second video template, and enhance the diversity of video generation; furthermore, they can quickly generate new videos locally and save them as video templates, improving video production efficiency, reducing reliance on professional equipment and manual intervention, and meeting the demand for efficient and low-cost video production.

[0079] The following is a detailed description of one implementation method for batch video generation based on multiple material templates provided in the embodiments of this application.

[0080] In one optional embodiment, in response to input operations on the video generation page regarding the number of videos and video categories, a batch video generation task is generated, which includes the number of videos N and the video categories. Based on the batch video generation task, N video instance identifiers are generated, and a set of video materials related to the video category is generated for each video instance identifier. Multiple material templates are obtained by variableizing an initial video template. The initial video template includes rendering rules for multiple video materials required for video generation and the hierarchical relationship between these materials. For each video instance identifier, at least one target material template is determined from the multiple material templates based on the material type in the set of video materials corresponding to the video instance identifier. Based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target material template is combined to obtain the target video template corresponding to the video instance identifier. The target video template describes the rendering rules and hierarchical relationship of the set of video materials. Video generation processing is performed based on the target video templates corresponding to each of the N video instance identifiers and the set of video materials to obtain N videos under the video category.

[0081] In this embodiment, the executing entity of the above-described video batch generation method is not limited. For example, the method can be implemented as a service product, which can adopt a client-server architecture. In the case of a client-server architecture, on the one hand, the client provides a video generation page for batch video generation to the user, receiving the user's input on the number and category of videos, and then initiating a batch video generation task to the server. On the other hand, the server responds to the client's initiation of a batch video generation task through the video generation service page, and uses the server's computing resources, network bandwidth, storage resources, and other resources to generate videos in batches, which helps to improve the speed of video generation.

[0082] For example, as the processing power of client-side hardware increases, the above method can also be executed by the client. The client can provide a video generation page to the user and respond to the user's input on the number and category of videos on the video generation page, generating batch video generation tasks. Then, the client's own computing and storage resources are used to generate videos in batches. In this case, when the batch generation tasks are deployed and executed on the client side, there is no need to transmit data to the server, which can save network latency.

[0083] In this embodiment, the batch video generation task includes the number of videos N and video categories, where N is an integer ≥ 2. The video category describes the theme of the generated video content; the specific implementation of the video category is not limited. For example, it includes, but is not limited to: product introductions, tutorials, science popularization, and beauty and skincare, etc. Optionally, the video category can be implemented as a single-level video category, such as business registration or legal consultation. Optionally, the video category can also be implemented as a multi-level video category. For example, the first-level video category could be business registration; the second-level video categories of this first-level category could be cleaning or food business, etc.

[0084] In this embodiment, N video instance identifiers are generated according to the batch generation task. Each of the N video instance identifiers is different. Each video instance identifier can be used to uniquely represent a video to be generated, so as to facilitate the tracking of the video materials and video templates required for the video to be generated. That is to say, the video instance identifier can also serve as a unique identifier for the relevant content (such as video materials) of the video to be generated.

[0085] The method for generating the N video instance identifiers is not limited. For example, it can include, but is not limited to, numbers and strings. For instance, it can start from any integer and form an increasing sequence with a fixed step size to obtain N integers as the N video instance identifiers; or it can be a preset set of N strings, etc.

[0086] In this embodiment, a set of video materials related to a video category is generated for each video instance identifier. This set of video materials is used to generate the video corresponding to that video instance identifier. Each set of video materials includes video materials of at least one material type. In some embodiments of this application, video materials of one material type are simply referred to as one type of video material.

[0087] The material type refers to the type of media resources that constitute the video, including but not limited to: audio, video, background images, subtitles, digital humans, and other material types.

[0088] In this embodiment, the implementation method of generating a set of video materials related to the video category for each video instance is not limited.

[0089] In one optional implementation, for any video instance identifier, video materials can be randomly extracted from multiple material types stored in the basic material library. At least one extracted video material is then used as a set of video materials for that video instance identifier. The basic material library stores multiple video materials under various video categories.

[0090] In another optional implementation, for any video instance identifier, the semantic similarity between at least one video material in the basic material library and the video category is calculated to obtain multiple similarity information; the video materials that meet the similarity conditions among the multiple similarity information are taken as a group of video materials for that video instance identifier.

[0091] Furthermore, in this embodiment, multiple material templates are obtained by performing variable processing on the initial video template.

[0092] In this embodiment, the initial video template can be an existing video production framework. This framework can be pre-defined or exported from video editing software. This embodiment does not limit the specific method of obtaining the initial video template. For example, the initial video template can be a pre-defined general video template provided by the system, a custom video template created by the user according to their needs, or a video template imported from other external sources.

[0093] The initial video template includes rendering rules for various video materials required to generate the video, as well as the hierarchical relationship between these video materials. In this embodiment, the rendering rules for each type of video material describe the rendering logic followed by that type of video material during the video rendering process, in order to achieve the desired rendering effect in the generated video. The hierarchical relationship between the various video materials within the initial video template refers to the layer stacking order of the various video materials in the initial video template during the video generation process, used to determine the overlapping relationship and occlusion logic of the various video materials in the video generation process.

[0094] In this embodiment, the variable processing of the initial video template is partly reflected in the structured decomposition of the initial video template and the templated rendering parameter information based on the material type as the splitting variable, to obtain multiple material-based templates. There is a correspondence between the material-based templates and the material types. For example, material types include audio, video, background images, subtitles, digital humans, etc. The corresponding material-based template types include, but are not limited to: audio material-based templates, video material-based templates, background image material-based templates, subtitle material-based templates, digital human material-based templates, etc. In other words, each material type corresponds to one material-based template, and each material-based template describes the rendering rules for that type of video material.

[0095] In this embodiment, the timing of the variable processing is not limited. For example, the initial video template can be variableized in advance. Alternatively, the initial video template can be variableized dynamically. For details on how to perform variable processing, please refer to subsequent embodiments.

[0096] In this embodiment, the variable-processed material templates can be modularly recombined. For each video instance identifier, based on the material types in a set of video materials corresponding to that video instance identifier, at least one target material template is determined from multiple material templates. There is a correspondence between the material types in this set of video materials and the material templates.

[0097] For example, if the set of video materials includes audio, video, background images, and subtitles, then the target material template can include the target material templates corresponding to each of the audio, video, background images, and subtitles. Similarly, if the set of video materials includes audio, video, background images, subtitles, and digital human materials, then the target material template can include the target material templates corresponding to each of the audio, video, background images, subtitles, and digital human materials.

[0098] Furthermore, given at least one target material template corresponding to a set of video materials, based on the hierarchical relationship between the various video materials included in the initial video template, the at least one target material template is combined to obtain the target video template corresponding to the video instance identifier. The target video template describes the rendering rules and hierarchical relationship of a set of video materials.

[0099] The initial video template includes a hierarchical relationship among various video materials. This hierarchy represents the layer stacking order of multiple material templates and can be used to organize and integrate target material templates, thereby forming a target video template with a clear hierarchy and corresponding rendering rules for the video materials. The target video template describes the hierarchical relationship of a group of video materials, which is consistent with the hierarchical relationship of that group of video materials in the initial video template.

[0100] In this embodiment, each of the N video instance identifiers corresponds to its own target video template. That is, each group of video materials has its own corresponding target video template, enriching the variety of target video templates used in batch video generation. Based on the rendering rules described by the N target video templates, video generation processing is performed on the grouped video materials corresponding to the N video instance identifiers, ensuring that the style of each batch-generated video matches the adopted video template, thereby increasing the diversity of the batch-generated video content.

[0101] Given the target video templates corresponding to each video instance identifier, video generation processing is performed based on the target video templates corresponding to each of the N video instance identifiers and a set of video materials to obtain N videos under the video category. Video generation processing refers to the process of filling and rendering the target video templates with video materials based on the target video templates corresponding to each video instance identifier and a set of video materials to obtain the video corresponding to that video instance identifier. For example, for N video instance identifiers, the set of video materials corresponding to each of the N video instance identifiers can be filled into N target video templates to obtain filled target video templates. Then, the filled target video templates can be rendered to obtain the videos under the video category.

[0102] In one optional embodiment, the video generation corresponding to each of the N video instance identifiers can be performed in batches. Each batch of video generation processes M videos in parallel, where M is less than N and M is an integer. After the M videos are generated, the process continues to generate the next M videos until all the videos corresponding to the N video instance identifiers have been processed. By performing batch video generation, the resource utilization of the server is improved, while avoiding server overload caused by high concurrency.

[0103] In this embodiment, during batch video generation, various material templates of different types are obtained by variable processing of the initial video template. These material templates are then recombined to obtain the video templates required for video generation. Furthermore, for each video generation, a corresponding video instance identifier is generated. Each video instance identifier is bound to a corresponding set of video materials. Based on the set of video materials corresponding to the video instance identifier, a target material template for recombination is determined from various material templates. Combining the hierarchical relationship between the material templates provided by the initial template, the target material template is organized and integrated to obtain the video template corresponding to each video instance identifier. This template is then used for video generation processing of the set of video materials corresponding to each video instance identifier. Since the video templates can be personalized by combining the video materials of the generated video, the flexibility of video generation is improved. Additionally, the recombined video templates have a corresponding relationship with the video instance identifiers, enriching the variety of video templates to a certain extent, so that the style of each batch-generated video matches the adopted video template, thus increasing the diversity of batch-generated video content.

[0104] In one optional embodiment, the target video template, a set of video materials, and the video corresponding to each of the N video instance identifiers can be uploaded to a CDN network or a server for storage. This ensures that the target video template, video materials, and the video corresponding to each of the N video instance identifiers can be quickly accessed and distributed, while reducing the storage pressure on local terminal devices. The CDN network can efficiently distribute the target video template, video materials, and the video corresponding to each of the N video instance identifiers to nodes near the user, improving access speed and stability. Server-side storage can centrally manage the target video template, video materials, and the video corresponding to each of the N video instance identifiers, supporting larger-scale batch video generation needs. By storing the target video template, video materials, and the video on a CDN network or server, users can access these resources at any time via access links for further editing or viewing, thereby improving the efficiency and flexibility of video generation.

[0105] In the embodiments of this application, the method of variable processing will be described in detail in the subsequent embodiments.

[0106] In this embodiment, the initial video template is variableized to obtain multiple material templates. During batch video generation, for each video instance identifier, at least one target material template is determined from the multiple material templates based on the material type in a set of video materials corresponding to the video instance identifier.

[0107] In one optional embodiment, when determining at least one target material template from multiple material templates based on the material type in a set of video materials corresponding to a video instance identifier, the method includes: identifying at least one material type contained in a set of video materials corresponding to a video instance identifier; selecting at least one initial material template from multiple material templates based on the at least one material type contained in the set of video materials, wherein each material type corresponds to one initial material template; and adjusting the parameters of at least some of the selected initial material templates to obtain at least one target material template.

[0108] In this embodiment, the material templates included in the multiple material templates obtained by variable processing of the initial video template are referred to as initial material templates. When at least one initial material template corresponding to a certain group is determined from the multiple initial material templates, the rendering rules described by the initial material templates can be referred to as initial rendering rules. Furthermore, parameters are adjusted for at least some of the initial material templates to obtain at least one target material template. Each initial material template, after parameter adjustment, yields a corresponding target material template. The target material template after parameter adjustment contains target rendering rules, which are different from the initial rendering rules described by the initial material templates.

[0109] In this embodiment, since the initial material template is obtained through variableization, the variableization process can transform the fixed rendering parameters in the initial video template into variable parameters that can be dynamically assigned values, thereby achieving flexible configuration of the template content. Different values ​​of the optional parameters will result in different rendering rules. Optionally, each initial material template includes at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values ​​and default parameter values.

[0110] In this embodiment, parameter adjustment of the initial material template is performed to assign different candidate parameter values ​​to the variable parameters of the initial material template, thereby obtaining multiple target material templates with different candidate parameter values. A target material template for a specific material type can be used to render video materials grouped from different video instance identifiers. Different variable parameter values ​​indicate different rendering rules described by the target material templates for the same material type, resulting in different rendering results for batch-generated videos, creating differentiation between different videos, and enhancing the richness of video content. The following will describe how to adjust the parameters of the initial material template.

[0111] In one optional embodiment, adjusting the parameters of at least a portion of the initial material templates in at least one initial material template to obtain at least one target material template includes: determining the number of templates whose parameters need to be adjusted, wherein the number of templates is less than or equal to the number of at least one initial material template; selecting initial material templates to be adjusted from the at least one initial material template according to the number of templates to be adjusted; determining variable parameters to be adjusted from the initial material templates to be adjusted; randomly determining a target parameter value from a plurality of candidate parameter values ​​associated with the variable parameter to be adjusted; and assigning the target parameter value to the variable parameter to be adjusted to obtain the target material template.

[0112] In this embodiment, the method for determining the number of templates whose parameters need to be adjusted is not limited. For example, parameters can be adjusted for each initial video template, in which case the number of templates to be adjusted is the number of at least one initial material template. Alternatively, a random integer within a preset range can be generated using a random number generation algorithm as the number of templates to be adjusted. The preset range refers to the number of templates being less than or equal to the number of at least one initial material template. The random number generation algorithm is not limited and includes, but is not limited to, methods such as the Linear Congruential Generator (LCG) and the Mersenne Twister algorithm.

[0113] Further, based on the number of templates to be adjusted, initial material templates to be adjusted are selected from at least one initial material template. In an optional embodiment, the initial material templates to be adjusted are selected based on the priority of the material type and the number of templates to be adjusted. The priority of the material type refers to the importance of the material type to the presentation effect of the generated video. For example, subtitle material types generally have a lower importance to the presentation effect, so the priority of subtitle material types can be set to a lower priority; relatively speaking, the priority of background images can be higher than that of subtitles, so they can be set to a medium priority; and digital humans have a higher importance to the presentation effect, so they can be set to a higher priority. Furthermore, based on the priority, templates for adjusting parameters can be selected first from initial material templates with higher material type priority. If the number of these initial material templates is less than the previously determined number of templates to be adjusted, then initial material templates with lower material type priority can be selected, and the final number of selected templates to be adjusted is less than or equal to the number of at least one initial material template.

[0114] Furthermore, from the initial material template to be adjusted, the variable parameters to be adjusted are determined. In one optional embodiment, all variable parameters in the initial material template to be adjusted can be used as the variable parameters to be adjusted. In another optional embodiment, a target variable parameter in the initial material template to be adjusted is used as the optional parameter to be adjusted, and the target variable parameter is a pre-selected optional parameter.

[0115] Then, the target parameter value is randomly determined from multiple candidate parameter values ​​associated with the variable parameter to be adjusted. In some embodiments, the N video instance identifiers each have at least one initial material template with the same variable parameter to be adjusted. The multiple candidate parameter values ​​associated with the variable parameter to be adjusted can be randomly selected as the target parameter values ​​for the N video instance identifiers, so that the optional parameter values ​​of the N video instance identifiers are as different as possible.

[0116] Furthermore, the target parameter value is assigned to the variable parameter to be adjusted to obtain the target material template.

[0117] Once the target material template is obtained, it is combined to obtain the target video template. This embodiment does not limit the combination method. Two combination methods are provided below, but are not limited to these.

[0118] In one optional embodiment, a basic video template is generated based on the hierarchical relationship between various video materials included in the initial video template. This basic video template serves as the framework for the target video template and includes multiple blank structural positions corresponding to various video materials. The blank structural positions are placeholders for target material templates preset in the basic video template, used to identify the positions where target material templates can be inserted. Each blank structural position corresponds to the filling of a target material template of a certain material type. The positional relationship between the multiple blank structural positions reflects the hierarchical relationship between various video materials. As described in the above embodiment, the hierarchical relationship represents the order in which various materials are superimposed in the generated video. In some embodiments, this hierarchical relationship is extracted from the initial video template; alternatively, it can be preset based on at least one target material template, meaning the hierarchical relationship can be set as needed. For example, the structural position corresponding to subtitles is located at the top layer, the structural position corresponding to digital humans is located at the middle layer, and the structural position corresponding to background images is located at the bottom layer. Further, at least one target material template is inserted into the corresponding blank structural positions in the basic video template to obtain the target video template corresponding to the video instance identifier.

[0119] In another optional embodiment, based on at least one target material template, the structural bits containing the rendering rules of video materials of the same material type in the initial video template are overwritten to obtain the target video template corresponding to the video instance identifier. The difference between the structural bits and the aforementioned blank structural bits is that each type of video material's rendering rule occupies one structural bit in the initial video template, while blank structural bits are empty. The positional relationship between the structural bits reflects the hierarchical relationship between multiple video materials.

[0120] Given a target video template, video generation processing is performed based on the target video template corresponding to each video instance identifier and a set of video materials. In an optional embodiment, when performing video generation processing based on the target video templates corresponding to N video instance identifiers and a set of video materials to obtain N videos under the video category, the process includes: for each video instance identifier, filling the set of video materials corresponding to the video instance identifier into the target material template of the target video template corresponding to the video instance identifier; for the filled target video template, rendering the set of video materials according to the rendering rules and hierarchical relationships of the set of video materials described in the filled target video template to obtain a video under the video category.

[0121] The target material template includes at least one placeholder for the corresponding material type, which is used to fill in video material of that material type.

[0122] In one optional embodiment, generating a set of video materials related to the video category for each video instance identifier includes: obtaining a set of video material description information related to the video category for each video instance identifier, wherein each set of video material description information includes description information of multiple video materials; for each video instance identifier, based on the set of video material description information corresponding to the video instance identifier, calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously uploading the multiple video materials corresponding to the video instance identifier to the content delivery network; correspondingly, before performing video generation processing based on the target video templates and a set of video materials corresponding to N video instance identifiers to obtain N videos under the video category, the method further includes: in response to a batch video generation trigger event, obtaining multiple sets of video materials corresponding to the N video instance identifiers from the content delivery network according to the N video instance identifiers. Hosting video materials through the content delivery network reduces server-side storage pressure, enabling the server to efficiently render a large number of videos and improve user experience.

[0123] In this embodiment, the descriptive information of each group of video materials is used to generate video materials with corresponding video instance identifiers. Each group of video materials contains video materials of various types. Material type refers to the type of different video elements that constitute the video content, including but not limited to: audio, video, background images, subtitles, digital humans, etc.

[0124] In this embodiment, the implementation method of generating a set of video material description information related to the video category for each video instance identifier is not limited.

[0125] In one optional implementation, for any video instance identifier, description information of video materials can be randomly extracted from multiple material types stored in the basic material library, and the extracted description information of multiple video materials can be used as the description information of a group of video materials for any video instance identifier. The basic material library stores description information of multiple video materials under multiple video categories.

[0126] In another optional implementation, for any video instance identifier, the semantic similarity between the description information of multiple video materials in the basic material library and the video category is calculated to obtain multiple similarity information; the description information of the video materials that meet the similarity conditions among the multiple similarity information is used as the description information of a group of video materials for that video instance identifier.

[0127] In another optional implementation, for any video instance identifier, a material description information generation model is invoked based on the video category. This model is used to generate description information for various video materials related to that video category. Specifically, this material description information generation model is trained using description information from a large number of sample video materials across different video categories and material types. By learning the semantic relevance between the description information of different sample video categories and material types, the model can selectively combine different video categories to generate description information for various related video materials.

[0128] Furthermore, for each video instance identifier, based on a set of video material description information corresponding to that video instance identifier, multiple material generation models based on artificial intelligence are invoked to generate multiple video materials corresponding to that video instance identifier, and these multiple video materials corresponding to that video instance identifier are simultaneously uploaded to the content delivery network.

[0129] One material generation model can generate at least some of the video materials from multiple sources. For example, one material generation model can generate only one type of video material. The following explanation uses this as an example, but is not limited to it.

[0130] In this embodiment, when generating video materials with video instance identifiers, a target video template corresponding to each video instance identifier is determined based on the multiple video materials corresponding to each video instance identifier. The target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to that video instance identifier. The method for determining the target video template corresponding to each video instance identifier can be referred to the above embodiment, and will not be repeated here.

[0131] In this embodiment, in response to the batch video generation trigger event, various video materials corresponding to the N video instance identifiers are obtained from the content delivery network according to the N video instance identifiers; videos are generated according to the various video materials corresponding to the N video instance identifiers and the target video template to obtain N videos under the video category.

[0132] In this embodiment, the specific implementation of the batch video generation trigger event is not limited and can be flexibly configured according to actual application needs. For example, it can be when all the various materials corresponding to N video instance identifiers have been generated; or, it can be when multiple materials corresponding to each video instance identifier are generated, in which case video generation will be performed on the various video materials of that video instance identifier; or, the batch video generation trigger event can also be a preset trigger time, such as a period of time after the input operation on the video generation page regarding the number of videos and video categories, for example, 2 hours, 1 day, or 1 week, etc., with no limitation on the time span.

[0133] It should be noted that each video instance identifier corresponds to the generation of one video, and N video instance identifiers can correspond to the generation of N videos. When generating videos for the various video materials corresponding to each of the N video instance identifiers, the generation process of the video corresponding to each video instance identifier is asynchronous, and the generation of videos corresponding to each video instance identifier does not affect each other, so as to improve the generation efficiency of N videos.

[0134] Further, optionally, the descriptive information for each set of video footage includes, but is not limited to: subtitle description information, audio type description information, digital human description information, and background image description information. For example... Figure 2As shown, based on a set of video material description information corresponding to the video instance identifier, multiple material generation models based on artificial intelligence are invoked to generate various video materials corresponding to the video instance identifier. This includes: generating text information using a generative language model based on subtitle description information to obtain target subtitles; converting target subtitles into target audio that matches the audio type description information using a text-to-speech model based on the target subtitles and audio type description information; selecting a target digital human based on the digital human description information using a multimodal model based on the target audio and the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain a green screen video of the target digital human; and generating a background image using a text-to-image model based on the background image description information to obtain a target background image.

[0135] In this embodiment, the APIs (Application Programming Interfaces) of multiple material generation models are associated with endpoints. In this embodiment, the video generation page is the presentation of the endpoint's front-end code. The endpoint is used for generating video materials, managing batch video generation tasks, and controlling the video generation process.

[0136] The endpoint includes front-end code and back-end code, adopting a client-server structure as described in the above embodiment. The front-end code refers to the video generation page built on the front-end framework. This page runs on the client side, interacts with the user, receives user input, and initiates batch video generation tasks to the server. The back-end code runs on the server side and is derived from the API of existing video editing software, such as Shortcut. In existing video editing software, the front-end UI code and video rendering function code are highly coupled. In this embodiment, the code is separated according to function to decouple the front-end UI and video rendering function code of the video editing software. The video rendering function code of the existing video editing software is then encapsulated into an independent API, allowing external code to call the video rendering function through a standard API, achieving automated video generation.

[0137] like Figure 2 In one example, the endpoint's front-end code could be a video generation page built on Astro, responsible for receiving callback notifications. For instance, when the material generation model is complete, it can send a notification to the endpoint to indicate that the subsequent video generation process can continue. The endpoint's back-end code could be obtained by API-izing the video rendering code of video editing software, exposing it externally for external calls via an API interface.

[0138] In one optional embodiment, the process of generating text information by calling a generative language model based on the subtitle description information to obtain target subtitles includes: matching corresponding keywords according to the video category, calling a pre-designed prompt word template, filling the prompt word template with keywords of the video category to obtain prompt words for the video category; inputting the prompt words of the video category into the generative language model to process the text information and obtain target subtitles related to the video category.

[0139] Keywords describe the theme of the video category; different video categories can correspond to different keywords, and different video types can also have different prompt word templates. Target subtitles include all the text content required for each video. By calling a generative language model, target subtitles can be automatically generated based on the video category without manual intervention, improving generation efficiency.

[0140] Optionally, based on the target subtitles and audio type description information, a text-to-speech model is invoked to convert the target subtitles into target audio that matches the audio type description information. The target subtitles provide the text content, and the audio type description information specifies the audio type of the generated target audio. Audio types include, but are not limited to, audio format, audio language, audio tone, and audio quality. Any audio type that can be used to specify the sound effects of the target audio is applicable to this embodiment. Based on the target subtitles and audio description information, the text-to-speech model is invoked to generate the corresponding target audio, which is then uploaded to a content delivery network, improving the automation of audio generation.

[0141] Furthermore, a multimodal model is invoked to generate a green screen video of a digital human. Here, "digital human" refers to a virtual character generated based on AI technology, capable of synchronously simulating the appearance, voice, lip movements, and other behaviors of a real person. It can be used in scenarios such as intelligent customer service, short video production, and virtual anchors, but is not limited to these. In this embodiment, by combining multiple video materials corresponding to each video instance identifier, a video with a realistic speaking effect can be generated.

[0142] In one optional embodiment, the digital human can be a pre-recorded video or image of a real person. In subsequent embodiments, the video and image are collectively referred to as video frames, and the number of video frames can be one or more. In this case, the descriptive information for different digital humans can be implemented as identification information. The identification information serves as a unique identifier for the digital human and is used to obtain the video frames of the digital human corresponding to the identification information. In another optional embodiment, the video frames of the digital human can be dynamically generated based on a multimodal model. In this case, the descriptive information of the digital human can be a prompt word used to generate the digital human, and the descriptive information of the digital human corresponding to each video instance can be different.

[0143] Optionally, when calling a multimodal model based on the target audio and digital human description information, selecting a target digital human based on the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain a green screen video of the target digital human, the process includes: obtaining the corresponding target digital human video frames based on the digital human description information; calling a multimodal model based on the target digital human video frames and the target audio to extract multidimensional features from the target audio to obtain multidimensional speech features; wherein, the multidimensional speech features include, but are not limited to: speech content features and speech emotion features; determining the lip-sync control parameters of the target digital human based on the speech content features; determining the facial expression control parameters of the target digital human based on the speech emotion features; determining the body movement control parameters of the target digital human based on the speech content features and the speech emotion features; generating a green screen video of the target digital human based on the lip-sync control parameters, facial expression control parameters, and body movement control parameters of the target digital human; and matching the lip-sync, facial expression, and body movement of the target digital human's green screen video with the target audio.

[0144] In this embodiment, speech content features are used to reflect the semantic information in the target audio, such as lexical content, grammatical structure, speech intent, and speech rhythm, and are mainly used to drive the lip movements of the digital human to be semantically synchronized with the target audio. Speech emotion features are used to reflect the emotional state of the target audio, including but not limited to tone of voice, speech rate, and intonation changes, and are mainly used to drive the facial expressions and body movements of the digital human.

[0145] Furthermore, based on the characteristics of speech content, the lip-sync control parameters of the target digital human are determined; based on the characteristics of speech emotion, the facial expression control parameters of the target digital human are determined; and based on the characteristics of speech content and speech emotion, the body movement control parameters of the target digital human are determined.

[0146] Having obtained lip-sync control parameters, facial expression control parameters, and body movement control parameters, the target digital human is driven to perform motion rendering based on these parameters, generating a corresponding green screen video of the target digital human. This allows for flexible replacement of the background image later. Since the green screen video of the target digital human is generated based on multi-dimensional speech features extracted from the target audio content, it ensures that the dynamic performance of the target digital human's lip-sync, facial expressions, and body movements in the green screen video matches the target audio semantically, temporally, and emotionally. This ensures precise lip-sync alignment and enhances the realism and viewing experience of the generated video.

[0147] Furthermore, in this embodiment, the target audio and target subtitles are aligned. For example, the target subtitles are time-stamped and aligned using the timestamps of the speech content in the target audio to ensure that the subsequent display of the target subtitles is completely synchronized with the target audio.

[0148] In this embodiment, the background image is generated in two ways: one is to use a text-based image model to generate the background image based on the background image description information, thereby obtaining the target background image; the other is to directly obtain a pre-generated or captured background image, which is not limited to this method.

[0149] In one optional embodiment, a set of video material generation states is maintained for each of the N video instance identifiers. The generation state of any video material of any type within each set includes: ready to generate, generating in progress, successfully generated, and failed generated. For example, for any video instance identifier, taking the multiple video materials that need to be generated for that video instance identifier, including: target subtitles, target audio, target digital human green screen video, and target background image, then the generation states for the target subtitles, target audio, target digital human green screen video, and target background image can be maintained separately.

[0150] Specifically, for each video instance identifier, the various video materials corresponding to that video instance identifier are marked as ready for generation. When calling multiple AI-based material generation models to generate the various video materials corresponding to that video instance identifier, if any material generation model returns a "generating in progress" response message, the video materials that the material generation model should generate are updated to the "generating in progress" state; if any material generation model returns a "generating successfully" message, the video materials that the material generation model should generate are updated to the "generating successfully" state; if any material generation model returns a "generating failed" message, the video materials that the material generation model should generate are updated to the "generating failed" state. For example... Figure 2 As shown, a subscription service is provided that can notify each material generation model of the generation status of the video material for that model, such as a successful generation status. The subscription service then returns a successful generation message to the server to inform it that the corresponding video material has been successfully generated. Figure 2 This example only uses the notification process for the target digital person as an example, but is not limited to this.

[0151] Optionally, if the generation status of the video footage is updated to a generation failure status, a failure notification message is output to the user who initiated the input operation. The failure notification message includes the footage type of the video footage that has been updated to a generation failure status and its corresponding video instance identifier. If a regeneration operation triggered by the user is received, the corresponding description information of the video footage is obtained based on the video instance identifier, and the corresponding AI-based footage generation model is invoked to regenerate the corresponding video footage.

[0152] In this embodiment, when multiple video materials corresponding to each video instance identifier are generated, these multiple video materials can be simultaneously uploaded to the content delivery network (CDN), and access links for the video materials corresponding to each video instance identifier can be obtained within the CDN. Continuing with the example above, for any video instance identifier, if the multiple video materials corresponding to that video instance identifier include target subtitles, target audio, target digital human green screen video, and target background image, the target subtitles, target audio, target digital human green screen video, and target background image are uploaded to the CDN respectively, and access links for the target subtitles, target audio, target digital human green screen video, and target background image are obtained.

[0153] Furthermore, in response to the batch video generation trigger event, based on the access links of the various video materials corresponding to each of the N video instance identifiers on the content delivery network, the various video materials corresponding to the N video instance identifiers are obtained respectively; and, based on the various video materials corresponding to each of the N video instance identifiers and the target video template, videos are generated to obtain N videos under the video category.

[0154] Following the example above, in response to the batch video generation trigger event, based on the access links of the target subtitles, target audio, target digital human green screen video, and target background image corresponding to each of the N video instance identifiers in the content delivery network, the target subtitles, target audio, target digital human green screen video, and target background image corresponding to each of the N video instance identifiers are obtained respectively; and, based on the target subtitles, target audio, target digital human green screen video, target background image, and target video template corresponding to each of the N video instance identifiers, videos are generated to obtain N videos under the video category.

[0155] Optionally, the videos and target video templates corresponding to the N video instance identifiers are uploaded to the content delivery network, and access links for the videos and target video templates corresponding to the N video instance identifiers in the content delivery network are obtained. The access links are added to the video generation results page. In response to the viewing operation of the video results page, the video generation results page is displayed. The video generation results page contains a set of video materials, videos, and access links for the target video templates corresponding to at least one of the N video instance identifiers in the content delivery network.

[0156] In this embodiment, each video instance identifies its own video materials, video, and target video template, which are stored on the content delivery network. This eliminates the need for the server to store a large number of files, reducing disk I / O load and ensuring service stability. Furthermore, storing access links for each video instance's identified video materials, video, and target video template on the content delivery network avoids data bloat, reduces query pressure, and supports larger-scale data management.

[0157] In this optional embodiment, in response to a triggered operation of an access link in the content delivery network corresponding to at least one video instance identifier, the set of video materials, videos, and / or target video templates corresponding to at least one video instance identifier is accessed. By hosting videos and video templates through a CDN, the video generation result page can directly load videos from the content delivery network. Compared to pulling videos from the server, this consumes less bandwidth, loads faster, and significantly improves performance, making it suitable for large-scale video browsing scenarios.

[0158] The following section introduces the methods for handling variables.

[0159] In one optional embodiment, the initial video template is parsed using the material type as a splitting variable to obtain information fragments corresponding to multiple material types; from the information fragments corresponding to each of the multiple material types, the rendering parameter information corresponding to each of the multiple material types is extracted; and the rendering parameter information corresponding to each of the multiple material types is templated to obtain multiple material templates.

[0160] Furthermore, after obtaining the initial video template, the template is parsed using the material type as a splitting variable. Here, material type refers to the different categories of video materials, used as a classification standard to distinguish different video materials, while the splitting variable refers to the independent information segments obtained by dividing the initial video template according to the material type during parsing. For example, if the initial video template contains both image and audio video materials, it will be split into independent image and audio information segments based on the material type, and then processed separately.

[0161] Here, an information fragment refers to an information segment extracted from the initial video template that corresponds to the material type. An information fragment can contain rendering parameter information for the corresponding material type. For example, an information fragment can contain rendering parameter information such as path information, playback time, and effects.

[0162] In this embodiment, rendering parameter information describes at least one attribute of a material file of a certain material type. Each attribute can be used as a rendering parameter, and the attribute value of each attribute in the initial video template can be used as the default parameter value of the corresponding rendering parameter. Rendering parameters with default parameter values ​​form a rendering parameter information corresponding to the material file of that material type. Here, "material file" refers to various resource files used in the materialization template, mainly including various video materials, such as images, videos, and audio. Rendering parameter information can be used to control the rendering logic of the material file so that the material file can achieve the rendering effect presented in the final generated video. Rendering parameter information quantifies and describes the attributes of the material file, enabling the conversion of video material attributes into rendering parameter information attributes, and controlling the rendering effect of the video material in the final generated video through the rendering parameter information. In other words, the attributes of the video material are described by the rendering parameter information as corresponding rendering parameters, and rendering parameters with default parameter values ​​define the attribute value of a certain attribute of the material file.

[0163] The rendering parameter information can be extracted from the information fragments corresponding to each of the various material types. The rendering parameter information for each material type includes at least one rendering parameter and its corresponding default parameter value. Each rendering parameter describes the attributes of a material file of a certain material type. In this embodiment, the specific attributes corresponding to the rendering parameters for each of the various material types are not limited.

[0164] In this embodiment, templating can transform the rendering parameter information corresponding to various material types from fixed-configuration default parameter values ​​into a set of variable parameters with dynamically replaceable parameter values. Specifically, templating the rendering parameter information corresponding to various material types yields multiple material templates for each material type. Each material template describes the rendering rules for a particular material type, and these rules include at least one variable parameter. Each variable parameter is associated with multiple candidate parameter values ​​to control the rendering effect of the generated video.

[0165] In this embodiment, the material template is the result of templated processing of rendering parameter information corresponding to various material types. It includes path information placeholders for the material files, as well as variable parameters and their associated candidate parameter values. The material template describes the rendering rules of the corresponding material type in the generated video. The rendering rules describe the rendering logic followed by the video material during the video rendering process, thereby controlling the expected rendering effect of the material file in the generated video. Specifically, by transforming the rendering parameter information from fixed-configuration rendering parameters with default values ​​into a material template formed by a dynamically replaceable set of variable parameters, different candidate parameter values ​​can be assigned to the variable parameters to flexibly adjust the rendering effect of the material files in the generated video in different scenarios. The templated rendering parameter information forms multiple modularly recombinable material templates. Different target video templates can be recombined for different material types, enabling the batch generation of diverse videos without the need for separate video template design and adjustment for each video. This not only saves time and manpower but also improves the flexibility and content diversity of video template generation, thereby achieving efficient batch generation of diverse style videos.

[0166] For example, for the same material template, by assigning different candidate parameter values ​​to the variable parameters, the attribute values ​​of various properties corresponding to the material type of the material template can be flexibly set, thereby achieving fine-grained control over the rendering rules of the material files. Material templates improve the flexibility of video generation, significantly improve the efficiency and quality of video generation, and meet diverse creative needs.

[0167] In this embodiment, it is not limited to the specific content of the candidate parameter values ​​associated with the variable parameters obtained after the rendering parameter information corresponding to various material types is templated.

[0168] By associating multiple candidate parameter values ​​with variable parameters, the variable parameters can be flexibly adjusted according to needs, improving the flexibility and diversity of video template generation, and thus achieving efficient batch generation of videos with diverse styles.

[0169] In one optional embodiment, the initial video template is parsed using the material type as a splitting variable to obtain information fragments corresponding to various material types. This includes: loading the XML document corresponding to the initial video template, where the XML document includes a root element and multiple non-root elements connected to the root element. Each non-root element includes multiple specific elements, each describing the rendering rules for a particular material type; traversing the non-root elements in the XML document starting from the root element to identify the multiple specific elements; and extracting multiple information fragments containing the multiple specific elements as information fragments corresponding to the various material types. The initial video template's XML document can be loaded into memory to form a parsable tree structure. In this embodiment, a DOM (Document Object Model) parser can be used to load the XML document into memory, forming a tree-structured parsing method.

[0170] The tree structure includes a root element and non-root elements. The root element is the top-level node of the XML document, and non-root elements are directly or indirectly nested under the root element. An XML document has one and only one root element, which is the first element of the XML document and can be used as the starting point of the XML document.

[0171] Non-root elements are child elements directly or indirectly nested under the root element, used to divide the initial video template into different information segments according to the material type. Non-root elements include multiple specific elements. By traversing the non-root elements in the XML document starting from the root element, multiple specific elements can be identified. Specific elements are elements that directly describe the rendering rules for a certain type of material, and each specific element corresponds to a specific material type.

[0172] By extracting these information fragments, the system can quickly identify the attributes of the source files and render them based on their attribute values. This design for extracting information fragments significantly improves the flexibility and diversity of video generation, meeting diverse creative needs, and providing a technological foundation for efficient batch video generation.

[0173] In an optional embodiment, starting from the root element, the non-root elements in the XML document are traversed to identify multiple specific elements, including: S1, starting from the root element, traversing the non-root elements in the XML document; S2, for the currently traversed non-root element, obtaining the element tag contained in the currently traversed non-root element; S3, if the element tag is a specific tag, determining whether the currently traversed non-root element contains child elements; S4, if the currently traversed non-root element contains child elements, then the child element is taken as the currently traversed non-root element, and the execution step S2 is returned; S5, if the element tag is a non-specific tag, then the next non-root element is encountered, and the execution step S2 is returned; S6, if the currently traversed non-root element does not contain child elements, then the currently traversed non-root element is taken as a specific element.

[0174] In step S1, the non-root elements of the XML document are traversed one by one, starting from the root element. In this embodiment, the specific implementation strategy of the traversal is not limited. For example, the traversal can be implemented using a depth-first search algorithm or a breadth-first search algorithm. In this embodiment, the process of traversing the non-root elements in the XML document to identify multiple specific elements is described in detail using a depth-first search algorithm as an example.

[0175] Next, step S2 is executed. For the currently traversed non-root element, the element tags contained in the currently traversed non-root element are obtained. In an XML document, element tags are identifiers enclosed in angle brackets (<>), used to mark the type and semantic meaning of an element. Different material types are distinguished by the name of the element tags; the specific location of the material is located by the hierarchical relationship of the element tags.

[0176] After retrieving the element tags contained in the currently traversed non-root element, determine whether the element tag is a specific tag. The specific tag can be a predefined set of key tags representing element tags in the XML document that require special handling, used to identify the type of material to be extracted.

[0177] Next, proceed to either step S3 or S5. If step S5 is executed, i.e., the element tag is a non-specific tag, then proceed to the next non-root element and return to step S2.

[0178] If step S3 is executed, i.e., the element tag is a specific tag, then it continues to determine whether the currently traversed non-root element contains child elements. Here, child elements can be other elements nested within the currently traversed non-root element.

[0179] Next, execute either step S4 or S6. If step S4 is executed, meaning the currently traversed non-root element contains child elements, then the child element is treated as the currently traversed non-root element, and the process returns to step S2. If step S6 is executed, meaning the currently traversed non-root element does not contain child elements, then the currently traversed non-root element is treated as a specific element. Step S4 is the recursive processing when child elements are present; the child element of the currently traversed non-root element is set as the new traversal starting point, and steps S2-S6 are re-executed for each child element. After the recursion ends, the traversal of other non-root elements continues. Step S6 is the end-collection when there are no child elements; the currently traversed non-root element is treated as a specific element.

[0180] In one optional embodiment, rendering parameter information corresponding to each of the various material types is extracted from information fragments corresponding to each material type, including: for each information fragment, extracting the path information of the material file and at least one attribute value of the material file from the information fragment, the attribute value being used to render the material file; reading the material file according to the path information, determining the material type described by the information fragment according to the extension of the material file; using at least one attribute to which at least one attribute value belongs as at least one rendering parameter, and using at least one attribute value as the default parameter value of at least one rendering parameter, so as to obtain the rendering parameter information corresponding to the material type described by the information fragment.

[0181] The information fragment is a portion extracted from the XML document that is related to the material type, containing the path information of the material file and at least one attribute value of the material file. The material file can be used to construct various media files for the final generated video. In this embodiment, the material file may include, but is not limited to: files in formats such as MP4 and AVI, which contain moving images and audio, and are video files used to display a series of continuous images; and files in formats such as WAV and MP3, which provide audio files containing background music, narration, or special effects sounds.

[0182] In this context, at least one attribute value extracted from the source file is a specific value corresponding to the attribute, describing the specific state of the attribute. Attributes can describe the rendering effect of the source file. For example, for a video file, attributes can be video duration, video speed, filter effects, etc. If the attribute is video duration, the attribute value can be 2 minutes, 1 hour, 1 day, etc.

[0183] In one optional embodiment, the rendering parameter information corresponding to the material type described by the information fragment includes at least one rendering parameter and a default parameter value for at least one rendering parameter. The default parameter value can be the attribute value of the attribute corresponding to the rendering parameter before quantization. Specifically, at least one attribute extracted from the material file in the information fragment can be used as at least one rendering parameter, and the at least one attribute value can be used as the default parameter value for at least one rendering parameter to obtain the rendering parameter information. That is, the rendering parameter and the default parameter value are combined into a key-value pair set to control the rendering logic of the material file in the video to achieve the desired rendering effect.

[0184] In one optional embodiment, the rendering parameter information corresponding to various material types is templated to obtain multiple material templates. This includes: for each material type, selecting at least one variable parameter from the rendering parameter information corresponding to that material type; associating multiple candidate parameter values ​​with the at least one variable parameter; and generating a material template corresponding to the material type based on the multiple candidate parameter values ​​associated with the at least one variable parameter. The variable parameter is obtained by variability processing of rendering parameters with default values. Each variable parameter is associated with multiple candidate parameter values, and different candidate parameter values ​​correspond to different rendering logics to produce different rendering effects in the generated video. This can be used to generate diverse material templates. For each material type, at least one variable parameter is selected from its corresponding rendering parameter information. These variable parameters can vary within an adjustment range to obtain different material templates. For example, when the material file is video material, playback speed and filter effects can be selected as variable parameters. Associating the variable parameter with multiple candidate values ​​limits the adjustment range of the variable parameter, ensuring that the material file of the generated video meets the design requirements and avoiding invalid settings.

[0185] In one optional embodiment, when selecting at least one variable parameter from the rendering parameter information corresponding to the material type, two selection methods are provided: one is to treat all rendering parameters as variable parameters, and the other is to select some rendering parameters as variable parameters based on weight values. When selecting to treat all rendering parameters as variable parameters, there is no need to compare weight values, and all rendering parameters can be adjusted as variable parameters.

[0186] In one optional embodiment, selecting some rendering parameters as variable parameters based on weight values ​​can be achieved by selecting at least one variable parameter from the rendering parameter information corresponding to the material type. This includes: pre-configuring weight values ​​for each rendering parameter for the material type; parsing each rendering parameter from the rendering parameter information corresponding to the material type; and selecting at least one rendering parameter whose weight value is greater than a set weight threshold as at least one variable parameter. The weight value can serve as an importance score assigned to each rendering parameter, quantifying the influence of the rendering parameter on the rendering effect of the material file in video generation. This allows for the selection of at least one rendering parameter that has a significant impact on user perception or application goals as at least one variable parameter. The weight threshold is a predefined critical value. Users can filter rendering parameters with weight values ​​higher than the weight threshold, selecting parameters with weight values ​​higher than the weight threshold as variable parameters. This avoids using too many irrelevant rendering parameters as variable parameters, preventing configuration conflicts or rendering logic confusion.

[0187] In this embodiment, the method of pre-configuring the weight values ​​of each rendering parameter is not limited. For example, the weight values ​​of each rendering parameter can be pre-configured based on empirical annotation or automatically generated through user behavior data analysis.

[0188] Optionally, selecting at least one variable parameter from the rendering parameter information corresponding to the material type can also be done randomly based on the set number of variable parameters. In an optional embodiment, generating a materialized template corresponding to the material type based on multiple candidate parameter values ​​associated with at least one variable parameter includes: adding each rendering parameter and its default value to a preset template file corresponding to the material type; adding multiple candidate parameter values ​​associated with at least one variable parameter to the preset template file; and adding a placeholder to the preset template file to carry the material file corresponding to the material type, thereby obtaining the materialized template corresponding to the material type. The preset template file is a basic video template containing basic configurations, including the basic structure of the video template and some preset rendering parameters. Adding each rendering parameter and its default value to the preset template file corresponding to the material type ensures that the final materialized template has complete rendering parameter information.

[0189] In this embodiment, in addition to adding the rendering parameters corresponding to each material type and their default values ​​to the preset template file, at least one variable parameter associated with multiple candidate parameter values ​​is added to the preset template file. These candidate parameter values ​​offer a variety of choices, allowing the user or system to select different candidate parameter values ​​according to their needs.

[0190] In one optional embodiment, based on adding at least one variable parameter associated with multiple candidate parameter values ​​to a preset template file, placeholders can be added to the preset template file to carry the material files corresponding to the material type, thereby obtaining a materialized template corresponding to the material type. Here, the placeholder is a reserved position used to fill in the material file, supporting dynamic replacement; for example, the path information of the actual material file corresponding to the material type can be filled in. By using placeholders, different material files can be flexibly replaced without modifying the template structure.

[0191] In the above embodiments, the material template integrates the rendering parameters of the material type and their default parameter values ​​into a preset base template, ensuring the integrity of the rendering parameter information of the generated material template. At the same time, by associating multiple candidate parameter values ​​with variable parameters, the parameter values ​​of variable parameters are dynamically selected to flexibly adjust to correspond to the video style. The embedding of placeholders further decouples the parameter configuration of the initial video template from the material files, supporting the dynamic replacement of material files without modifying the template structure. Thus, it allows the assignment of corresponding candidate parameter values ​​to variable parameters according to application requirements to obtain freely combinable material templates. Ultimately, this significantly improves the flexibility and diversity of video template generation, thereby efficiently mass-producing video content with different styles and solving the problems of cumbersome parameter configuration, complex adaptation process, and serious homogenization of generated video content in traditional video production.

[0192] Based on the aforementioned multiple source templates, batch video generation or single video generation can be performed using these templates. For batch video generation scenarios, the multiple source templates provided in this application embodiment can generate multiple videos with diverse styles and differentiated content.

[0193] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0194] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 11, 12, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0195] Figure 3This is a schematic diagram of an electronic device structure provided for an exemplary embodiment of this application. For example... Figure 3 As shown, the device includes: a memory 34 and a processor 35.

[0196] Memory 34 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, a first video template, a first video, a target material template, etc.

[0197] The memory 34 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0198] The processor 35, coupled to the memory 34, is used to execute a computer program in the memory 34 for: downloading a first video template to a local terminal device, the first video template being a video template for rendering a set of video materials to generate a first video, the first video template including: a target material template adapted to the material type in the set of video materials from multiple material templates, the multiple material templates being obtained by varying the initial video template; each material template describing the rendering rules of a video material, the first video template describing the rendering rules and hierarchical relationship of a set of video materials, the hierarchical relationship being reflected in the positional relationship between the target material templates in the first video template; running video editing software on the terminal device, importing the first video template into the video editing software, opening the first video template in the editing area of ​​the video editing software, and displaying the first video generated based on the first video template in the preview area of ​​the video editing software; in response to a modification operation on the first video template, determining the modified rendering rules in any target material template and / or the modified positional relationship between any two target material templates, and editing the first video according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video.

[0199] In an optional embodiment, the processor 35, in response to a modification operation on the first video template, determines the modified rendering rules in any target material template, including: in response to a positioning operation on the first video template, displaying the target material template to be modified in the editing area; in response to a code editing operation in any target material template, obtaining the modified rendering rules; or in response to a trigger operation to modify any target material template in the first video template, displaying it in the first editing window; in response to a first editing operation in the first editing window, determining the name of any target material template and the modified rendering rules; and in response to a first edit submission operation, refreshing the first video template according to the name of any target material template and the modified rendering rules.

[0200] In an optional embodiment, the processor 35, in response to a modification operation on the first video template, determines the modified positional relationship between any two target material templates, including: in response to a drag operation on any target material template in the first video template, moving any target material template according to the drag trajectory of the drag operation, and determining another target material template according to the position when the drag operation ends; filling the structural position of the other target material template into the structural position of the first target material template, and filling the structural position of the other target material template into the structural position of the first target material template, to obtain the modified positional relationship between any two target material templates; or, in response to a trigger operation to adjust the positional relationship between any two target material templates in the first video template, displaying a second editing window; in response to a modification operation in the second editing window, determining the names of any two target material templates and the information of the structural positions after mutual swapping; in response to a modification submission operation, refreshing the first video template according to the names of any two target material templates and the information of the structural positions after mutual swapping, to obtain the modified positional relationship between any two target material templates; wherein, a structural position refers to a position describing the corresponding target material template in the first video template.

[0201] In an optional embodiment, the processor 35 downloads the first video template to the local terminal device, including: in response to an access operation to the batch video generation result page, displaying the batch video generation result page, which displays access links on the CDN network or server for each of the N video instance identifiers corresponding to the target video templates, access links on the CDN network or server for each of the N video instance identifiers corresponding to a set of video materials, and access links on the CDN network or server for each of the N video instance identifiers corresponding to the videos; where N is an integer ≥ 2; and in response to a trigger operation on the access link of any target video template on the CDN network or server, downloading any target video template as the first video template from the CDN network or server to the local terminal device.

[0202] In one optional embodiment, the processor 35, in response to a triggered operation of access links for multiple video materials on a CDN network or server, downloads multiple video materials from the CDN network or server to a local terminal device; wherein the multiple video materials are distributed in one or more groups of video materials; the processor runs video editing software on the terminal device, imports the multiple video materials into the video editing software, and displays the multiple video materials in the editing area of ​​the video editing software; in response to the editing operation of the multiple video materials, a third video is generated, and a second video template corresponding to the third video is generated based on the attribute information and hierarchical relationship of the edited multiple video materials. The second video template includes the rendering rules of the multiple video materials and the positional relationship of the multiple video materials in the second video template.

[0203] In an optional embodiment, the processor 35 responds to input operations on the video generation page regarding the number of videos and video categories, generating a batch video generation task, which includes the number of videos N and the video category; based on the batch video generation task, it generates N video instance identifiers and generates a set of video materials related to the video category for each video instance identifier; it obtains multiple material templates obtained by variable processing of the initial video template, the initial video template including rendering rules for multiple video materials required to generate videos and the hierarchical relationship between multiple video materials; for each video instance identifier, based on the material type in the set of video materials corresponding to the video instance identifier, it determines at least one target material template from the multiple material templates; based on the hierarchical relationship between the multiple video materials included in the initial video template, it combines the at least one target material template to obtain the target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rules and hierarchical relationship of a set of video materials; it performs video generation processing based on the target video templates corresponding to each of the N video instance identifiers and a set of video materials to obtain N videos under the video category; and it uploads the target video templates corresponding to each of the N video instance identifiers, a set of video materials, and the videos to a CDN network or server for storage.

[0204] In one optional embodiment, the processor 35 uses the material type as a splitting variable to parse the initial video template and obtain information fragments corresponding to multiple material types; it extracts rendering parameter information corresponding to each of the multiple material types from the information fragments corresponding to each of the multiple material types; and it templates the rendering parameter information corresponding to each of the multiple material types to obtain multiple material templates.

[0205] Furthermore, such as Figure 3 As shown, the electronic device also includes other components such as a communication component 36, a display 37, a power supply component 38, and an audio component 39. Figure 3The diagram only shows some components and does not mean that the computing platform includes only these components. Figure 3 The components shown. Additionally... Figure 3 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the work node. In this embodiment, the work node can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the work node in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 3 The components within the dashed box; if the working node in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may be omitted. Figure 3 The component within the dashed box.

[0206] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by an electronic device in the above method embodiments.

[0207] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0208] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.

[0209] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0210] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0211] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0212] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code.

[0213] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0214] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0215] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0216] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.

[0217] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0218] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0219] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0220] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A video editing method, characterized in that, include: Download the first video template to the local terminal device. The first video template is a video template used to render a set of video materials to generate the first video. The first video template includes: a target material template that matches the material type in the set of video materials from a variety of material templates. The variety of material templates are obtained by variableizing the initial video template. Each material template is used to describe the rendering rules of a video material. The first video template is used to describe the rendering rules and hierarchical relationship of the set of video materials. The hierarchical relationship is reflected in the positional relationship of the target material template among the first video templates. Run the video editing software on the terminal device, import the first video template into the video editing software, open the first video template in the editing area of ​​the video editing software, and display the first video generated based on the first video template in the preview area of ​​the video editing software; The first video generated based on the first video template is obtained by filling the target material template in the first video template with the set of video materials respectively, and rendering the set of video materials according to the rendering rules and hierarchical relationship of the set of video materials described in the first video template after filling. In response to the modification operation on the first video template, the modified rendering rules in any target material template and / or the modified positional relationship between any two target material templates are determined, and the first video is edited according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video. The second video has the same set of video materials as the first video, but the visual effect or hierarchical relationship of the set of video materials in the second video is different from that in the first video.

2. The method according to claim 1, characterized in that, In response to a modification operation on the first video template, determine the modified rendering rules in any target material template, including: In response to the positioning operation of the first video template, any target material template that needs to be modified is displayed in the editing area; in response to the code editing operation in the target material template, the modified rendering rules are obtained; or In response to a triggering operation that modifies any target material template in the first video template, the first editing window is displayed; in response to a first editing operation in the first editing window, the name of the target material template and the modified rendering rules are determined; in response to a first editing submission operation, the first video template is refreshed according to the name of the target material template and the modified rendering rules.

3. The method according to claim 1, characterized in that, In response to a modification operation on the first video template, determining the modified positional relationship between any two target material templates includes: In response to a drag operation on any target material template in the first video template, the target material template is moved according to the drag trajectory of the drag operation, and another target material template is determined according to the position when the drag operation ends; the target material template is filled into the structural position where the other target material template is located, and the other target material template is filled into the structural position of the target material template, so as to obtain the modified positional relationship between any two target material templates; or In response to a trigger operation that adjusts the positional relationship between any two target material templates in the first video template, a second editing window is displayed; in response to a modification operation in the second editing window, the names of the two target material templates and the information of their swapped structural positions are determined; in response to a modification submission operation, the first video template is refreshed based on the names of the two target material templates and the information of their swapped structural positions to obtain the modified positional relationship between the two target material templates. The structural bit refers to the position of the corresponding target material template in the first video template.

4. The method according to any one of claims 1-3, characterized in that, Download the first video template to your local terminal device, including: In response to an access operation to the batch video generation results page, the batch video generation results page is displayed. The batch video generation results page displays access links on the CDN network or server for each of the N video instance identifiers corresponding to the target video template, access links on the CDN network or server for each of the N video instance identifiers corresponding to a set of video materials, and access links on the CDN network or server for each of the N video instance identifiers corresponding to the video; N is an integer ≥ 2. In response to a triggered operation of accessing a link to any target video template on a CDN network or server, the target video template is used as the first video template and downloaded from the CDN network or server to the local terminal device.

5. The method according to claim 4, characterized in that, Also includes: In response to a triggered operation of accessing multiple video materials on a CDN network or server, the multiple video materials are downloaded from the CDN network or server to the local terminal device; wherein the multiple video materials are distributed in one or more groups of video materials; Run the video editing software on the terminal device, import the multiple video materials into the video editing software, and display the multiple video materials in the editing area of ​​the video editing software; In response to the editing operations on the multiple video materials, a third video is generated, and a second video template corresponding to the third video is generated based on the edited attribute information and hierarchical relationship of the multiple video materials. The second video template includes the rendering rules of the multiple video materials and the positional relationship of the multiple video materials in the second video template.

6. The method according to claim 4, characterized in that, Also includes: In response to input operations on the video generation page regarding the number of videos and video categories, a batch video generation task is generated, wherein the batch video generation task includes the number of videos N and the video categories; Based on the batch video generation task, N video instance identifiers are generated, and a set of video materials related to the video category are generated for each video instance identifier; Obtain multiple material templates by performing variable processing on an initial video template. The initial video template includes rendering rules for multiple video materials required to generate the video and the hierarchical relationship between the multiple video materials. For each video instance identifier, at least one target material template is determined from the multiple material templates based on the material type in a set of video materials corresponding to the video instance identifier. Based on the hierarchical relationship between the various video materials included in the initial video template, the at least one target material template is combined to obtain the target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationship of the group of video materials. Based on the target video template and a set of video materials corresponding to each of the N video instance identifiers, video generation processing is performed to obtain N videos under the video category; The target video template, a set of video materials, and the video corresponding to each of the N video instance identifiers are uploaded to the CDN network or server for storage.

7. The method according to any one of claims 1-3 or 6, characterized in that, Also includes: Using the material type as a splitting variable, the initial video template is parsed to obtain information fragments corresponding to various material types; Extract the rendering parameter information corresponding to each of the various material types from the information fragments corresponding to each of the various material types; The rendering parameter information corresponding to each of the various material types is templated to obtain the various material templates.

8. An electronic device, characterized in that, include: A processor and a memory, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1-7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-7.

10. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 1-7.