Video editing method and device, storage medium and program product
By modifying and editing video templates in video editing software, the complex and cost-effective video production in the existing technology is solved, and efficient and diversified video production is achieved.
Patent Information
- Application Number
- CN202510443784.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, video production needs to go through multiple complex links, relying on professional equipment and labor, resulting in high costs and low efficiency, and cannot meet the increasing video needs.
By downloading video templates used to describe the rendering rules and hierarchical relationships of video materials to the local terminal device, and importing and modifying the templates in the video editing software, efficient editing and generation of videos can be achieved.
It improves the efficiency and diversity of video production, reduces dependence on professional equipment and manual intervention, reduces costs, and enriches the types of video templates to meet diverse video needs.
Smart Images

Figure CN120186431A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video processing, and particularly to a video editing method, device, storage medium, and program product. Background Art
[0002] With the rapid development of digital media technology and the diversification of content consumption demands, video production plays a crucial role in many fields such as advertising production, film and television creation, online education, and short video platforms. For example, in advertising production, attractive video advertisements can be produced according to product features and target audiences, thereby enhancing brand awareness and product sales.
[0003] In the prior art, in order to meet flexible and diverse video requirements, the video production method needs to go through multiple complex links such as shooting, post - editing, and special effect adding. Each link requires professional equipment and personnel, as well as a large amount of time and effort to complete. The video production cost is high, and the production efficiency is low, unable to adapt to the increasing video requirements. Summary of the Invention
[0004] Embodiments of this application provide a video editing method, device, storage medium, and program product to improve the efficiency and diversity of video production, reduce the dependence on professional equipment and manual intervention, and meet the diverse video requirements of high efficiency and low cost.
[0005] An embodiment of this application provides a video editing method, including: downloading a first video template to a local terminal device, where the first video template is a video template for rendering a set of video materials to generate a first video. The first video template includes: a target materialized template adapted to the material type in the set of video materials among multiple materialized templates, and the multiple materialized templates are obtained by applying variation amounts to an initial video template; each materialized template is used to describe the rendering rule of a video material, and the first video template is used to describe the rendering rule and hierarchical relationship of a set of video materials, and the hierarchical relationship is reflected by the position relationship of the target materialized template in the first video template; running video editing software on the terminal device, importing the first video template into the video editing software, opening the first video template in the editing area of the video editing software, and displaying the first video generated based on the first video template in the preview area of the video editing software; in response to a modification operation on the first video template, determining the modified rendering rule in any target materialized template and / or the modified position relationship between any two target materialized templates, and editing the first video according to the modified rendering rule and / or the modified position relationship to obtain an edited second video.
[0006] An embodiment of the present application further provides an electronic device, including: a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the processor is enabled to implement each step in the video editing method provided by the embodiment of the present application.
[0007] An embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement each step in the video editing method provided by the embodiment of the present application.
[0008] An embodiment of the present application further provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the processor is enabled to implement each step in the video editing method provided by the embodiment of the present application.
[0009] In an embodiment of the present application, by downloading a first video template for describing the rendering rules and hierarchical relationships of a group of video materials to a local terminal device, and importing the first video template into a video editing software running on the terminal device; opening the first video template in the editing area of the video editing software, and displaying a first video generated based on the first video template in the preview area of the video editing software; further, by responding to a modification operation on the first video template, determining the modified rendering rules in any target materialized template and / or the modified positional relationship between any two target materialized templates, and editing the first video according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video. Based on the existing first video template, combining with the localized video editing software to perform secondary editing on the first video to obtain the edited second video can improve the efficiency of video production, reduce the dependence on professional equipment and manual intervention, thereby reducing costs; and since the first video template is obtained by personalized combination with the video materials for generating the video, the combined first video template has a corresponding relationship with the first video, enriching the types of video templates to a certain extent, and further meeting the video requirements for diverse video content generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0011] Figure 1 It is a schematic flowchart of a video editing method provided by an exemplary embodiment of the present application;
[0012] Figure 2 It is an interactive schematic diagram of a video batch generation method provided by an exemplary embodiment of the present application;
[0013] Figure 3 A schematic structural diagram of an electronic device provided for an exemplary embodiment of the present application. Detailed implementation manners
[0014] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0015] It should be noted that in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse. In addition, various models (including but not limited to language models or large models) involved in the present application comply with the relevant laws and standards.
[0016] In addition, it should be noted that in the case where the embodiments of the present application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of the present application include but are not limited to: interaction operations in various ways such as touch operations, gesture operations, voice operations, head movement operations, eye movement operations, etc.; among them, touch operations include but are not limited to: click operations, double-click operations, long-press operations, swipe operations, pinch operations or mouse hover operations, etc. Swipe operations include but are not limited to: linear swipes, curved swipes, etc.
[0017] Furthermore, it should be noted that in the case where the embodiments of the present application involve the jump between a first interface and a second interface, the jump methods involved in the embodiments of the present application include but are not limited to: directly jumping from the first interface to the second interface, first jumping from the first interface to a task interface and then jumping to the second interface when corresponding task operations are completed on the task interface; the completion of the corresponding task operations on the task interface includes but is not limited to: when the task interface is implemented as a game interface, completing game operations on the game interface; when the task interface is implemented as an identity authentication interface, completing identity authentication on the identity authentication interface; when the task interface is implemented as a recharge interface, completing recharge operations on the recharge interface; and so on.
[0018] In the prior art, in order to meet flexible and diverse video requirements, the video production method needs to go through multiple complex processes such as shooting, post-editing, and special effect adding. Each process requires professional equipment and personnel, as well as a large amount of time and effort to complete. The video production cost is high and the production efficiency is low, making it unable to adapt to the increasing video requirements.
[0019] To solve the problems existing in the above prior art, in the embodiments of the present application, a first video template for describing the rendering rules and hierarchical relationships of a group of video materials is downloaded to a local terminal device, and the first video template is imported into a video editing software running on the terminal device; the first video template is opened in the editing area of the video editing software, and a first video generated based on the first video template is displayed in the preview area of the video editing software; further, by responding to the modification operation on the first video template, the modified rendering rules in any target materialized template and / or the modified positional relationship between any two target materialized templates are determined, and the first video is edited according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video. Based on the existing first video template, the first video is secondarily edited by combining with the localized video editing software to obtain the edited second video, which can improve the efficiency and diversity of video production, reduce the dependence on professional equipment and manual intervention, thereby reducing costs; and since the first video template is obtained by personalized combination with the video materials for generating the video, the combined first video template has a corresponding relationship with the first video, enriching the types of video templates to a certain extent, and further meeting the video requirements for diverse video content generation. The technical solutions provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0020] Figure 1 The flowchart of a video editing method provided by an exemplary embodiment of the present application. As Figure 1 shown, the method includes:
[0021] S11. Download a first video template to a local terminal device. The first video template is a video template used to render a group of video materials to generate a first video; a target materialized template in multiple materialized templates that is adapted to the material type in a group of video materials. The multiple materialized templates are obtained by varying an initial video template; each materialized template is used to describe the rendering rules of a video material, and the first video template is used to describe the rendering rules and hierarchical relationships of a group of video materials. The hierarchical relationship is reflected in the positional relationship between the target materialized templates in the first video template;
[0022] S12. Run the video editing software on the terminal device, import the first video template into the video editing software, open the first video template in the editing area of the video editing software, and display the first video generated based on the first video template in the preview area of the video editing software.
[0023] S13. In response to the modification operation on the first video template, determine the modified rendering rules in any target materialized template and / or the modified positional relationship between any two target materialized templates, and edit the first video according to the modified rendering rules and / or the modified positional relationship to obtain the edited second video.
[0024] In the embodiments of the present application, the specific form or deployment method of the local terminal device is not limited. Any device with video processing and editing capabilities and capable of implementing the video editing method can be used as the local terminal device in the embodiments of the present application. For example, the local terminal device can be a smart phone or a tablet computer (such as an iOS / Android device) equipped with a mobile operating system, and the clipping operation is completed through touch interaction; it can also be a desktop computer or a workstation equipped with a professional GPU (Graphics Processing Unit), supporting 4K / 8K high-resolution video processing; it can also be an embedded device (such as a smart TV, an in-vehicle entertainment system, or a drone control terminal).
[0025] In the embodiments of the present application, the storage location of the first video template is not limited. Any cloud or edge computing node with storage resources can store the first video template in the embodiments of the present application. For example, the first video template can be stored in a CDN network (Content Delivery Network), and downloading the first video template from the CDN network can reduce the transmission delay; it can also be stored on the server side and downloaded from the server side; it can also be stored in an edge computing node, and the transmission quality of the first video template can be guaranteed by combining network slicing technology.
[0026] In the embodiments of the present application, the first video template is a video template used to render a group of video materials to generate the first video. Among them, the first video refers to the video obtained by rendering a group of video materials according to the rendering rules provided by the first video template. A group of video materials is video materials containing multiple material types. The material type refers to the type of different video elements (i.e., video materials) that make up the video content, including but not limited to: audio, video, background image, subtitle, digital human, and other material types. For example, a group of video materials can include a background video, an audio of a digital human explanation, a subtitle layer for display, and a background image.
[0027] In the embodiments of the present application, the first video template includes: a target materialized template among multiple materialized templates that is adapted to the material types in a set of video materials. Among them, the target materialized template refers to the materialized templates corresponding to multiple video materials combined according to the combination logic to obtain a materialized template that is adapted to the material types in a set of video materials. The combination logic includes a hierarchical relationship, which refers to the layer stacking order of multiple video materials during the video generation process and is used to determine the front-to-back coverage relationship and occlusion logic of multiple video materials during video generation. For example, when generating a video, the subtitles can be located in the relatively upper layer of the layer, the background image can be located in the relatively lower layer of the layer, the digital human can be superimposed on the upper layer of the background image, and under the layer of the subtitles.
[0028] In addition, the first video template further includes at least one access link to a video material. The access link refers to a network address or storage path used to locate and obtain the video material and is used to obtain the video material in the CDN network during the rendering process. For example, when the video material type is a digital human, the access link can point to the digital human stored in the CDN network, such as https: / / model-cdn.example.com / avatar_001.glb.
[0029] In the embodiments of the present application, multiple materialized templates refer to the templates obtained after variable processing of the initial video template. Among them, the initial video template refers to an existing video production framework. The video production framework can be preset or exported after video editing by an editing software. Regardless of which method, the initial video template includes the rendering rules for multiple video materials required for video generation and the hierarchical relationship between multiple video materials. Here, it should be noted that the initial video template does not include video materials, but includes rendering rules and hierarchical relationships. The rendering rules correspond to the types of video materials, and the hierarchical relationship describes the display order and occlusion logic existing between different video materials.
[0030] Optionally, the initial video template can be implemented as an MLT (Media Lovin' Toolkit) template. The MLT template is an open-source framework for multimedia processing, which is a file based on the XML format and records all parameters of video editing, such as video clips on the timeline, audio tracks, filter effects, transitions, etc., and can be used for video editing.
[0031] In the embodiments of the present application, part of the variational processing is reflected in taking the material type as the splitting dimension, structurally splitting the initial video template, and templatizing the parameter information to obtain the materialized template. Among them, there is a corresponding relationship between the materialized template and the material type. For example, the material types include material types such as audio, video, background image, subtitle, digital human, etc. Then the types of the corresponding materialized templates include but are not limited to: audio materialized template, video materialized template, background image materialized template, subtitle materialized template, digital human materialized template, and so on.
[0032] In the embodiments of the present application, each materialized template is used to describe the rendering rules of a video material, and the first video template is used to describe the rendering rules and hierarchical relationships of a group of video materials. The hierarchical relationship is reflected in the positional relationship between the target materialized template and the first video template. Among them, the rendering rules are used to control the presentation form of a video material in the generated video, or in other words, the rendering result. The rendering rules can define the presentation method of the video material by combining the rendering parameter information. The rendering parameter information refers to the specific parameters related to the rendering of the video material. The rendering parameter information can be used to describe at least one attribute of the material file of a material type. Each attribute can be used as a rendering parameter, and the attribute value of each attribute in the initial video template can be used as the default parameter value of the corresponding rendering parameter. The rendering parameters with default parameter values form a rendering parameter information corresponding to the material file of this material type. Among them, the "material file" refers to various resource files used in the materialized template, mainly including various video materials, such as pictures, videos, and audios. The rendering parameter information can be used to control the rendering logic of the material file so that the material file can obtain the presented rendering effect in the finally generated video. The rendering parameter information quantitatively describes the attributes of the material file, so that the attributes of the video material are converted into the attributes of the rendering parameter information, and the rendering effect of the video material in the finally generated video is controlled through the rendering parameter information. In other words, the attributes of the video material are described as the corresponding rendering parameters by the rendering parameter information, and the rendering parameters with default parameter values define the attribute values of a certain attribute of the material file.
[0033] Among them, the rendering parameter information can be extracted separately from the information segments corresponding to various material types. The rendering parameter information corresponding to each material type includes at least one rendering parameter and its corresponding default parameter value. Each rendering parameter is used to describe the attribute of the material file of a material type. In the embodiments of the present application, the specific attributes corresponding to the rendering parameters of various material types are not limited.
[0034] Among them, the hierarchical relationship is reflected in the positional relationship between the target materialized template and the first video templates. Specifically, it refers to the layer stacking order among different materialized templates during the video generation process. For example, in a video, the subtitle materialized template may be set to be on the upper layer, the background image materialized template is on the bottom layer, and the digital human materialized template may be in the middle layer. This positional relationship determines the display order and occlusion logic of multiple video materials in the final video, that is, the upper-layer material will cover the lower-layer material, and the middle-layer material may partially occlude the bottom-layer material while being occluded by the upper-layer material. In this way, the hierarchical relationship ensures the reasonable presentation of video content, enabling the various elements in the video to be superimposed and interact with each other as expected.
[0035] In the embodiments of the present application, the video editing software refers to an application program that can run on a terminal device and has a video editing area and a preview area. Among them, the video editing area refers to an interface where operations such as importing, editing, adjusting, and combining video materials or video templates can be performed. The first video template can be imported into the video editing software, the first video can be opened in the editing area of the video editing software, and modification operations can be performed on the first video template in the video editing area. For example, adjusting the order of video materials, modifying the rendering rules, changing the hierarchical relationship, etc. The preview area refers to a visual interface area that displays the results of the current editing operation in real time, and the first video generated based on the first video template can be displayed in the preview area. For example, after the user modifies the parameters in the editing area, the preview area will immediately display the effect of the second video generated based on the updated first video template, ensuring that the user can immediately verify the editing results. In the embodiments of the present application, the specific form of the video editing software is not limited. For example, the video editing software can be open-source free software, such as Shotcut that supports basic multi-platform editing; or it can be other custom-developed dedicated editing tools, such as enterprise-level video production platforms or industry-customized software.
[0036] In the embodiments of the present application, taking the video editing software implemented as Shotcut as an example, the specific implementation manner of editing the first video is described. Among them, the way of editing the first video template can modify the rendering rules involved in the target materialized template and / or the modified positional relationship between any two target materialized templates according to requirements, which will be described in detail through embodiments below.
[0037] Optional Embodiment 1: The way of editing the first video template can modify the rendering rules involved in the target materialized template according to requirements, and specifically can be implemented as customizing and adjusting the rendering parameter information of the video materials involved. For example, customizing and adjusting the rendering parameter information of the video materials can adjust the attribute values of the video materials.
[0038] In an alternative embodiment, customizing the rendering parameter information of video material can be performed by adjusting the attribute values of the specific attributes related to the modified video material types in the first video template, so as to optimize the picture quality, performance or adapt to specific requirements. Among them, Shotcut can provide multiple candidate parameter values for the corresponding attribute values of multiple video materials, and one of the candidate parameter values corresponding to the attribute value of the involved video material can be used to replace the attribute value of the specific attribute of the original video material type. For example, for a video material with a video material type, the attribute value of its attribute of resolution can be adjusted from "1080p" to "720p", or the attribute value of its attribute of brightness can be replaced from "50" to "60".
[0039] Alternative Embodiment 2: The method of editing the first video template can be based on the requirement to modify the positional relationship between any two target materialized templates involved in the target materialized template. Modifying the positional relationship between any two target materialized templates can be achieved by adjusting the modification hierarchy relationship for any two video materials, or can be achieved by changing the playback order of any two video materials.
[0040] In an alternative embodiment, the method of editing the first video template can also be to adjust the modification hierarchy relationship for any two video materials. The adjustment of the hierarchy relationship means that in a multi-track timeline, by changing the track position where the material is located or the overlay speed order, the covering relationship and presentation priority of different materials are controlled. In the embodiments of the present application, the method of adjusting the hierarchy relationship is not limited. For example, Shotcut can be used to adjust by means of track hierarchy, that is, each video track on the time axis (Time Axis, also known as the timeline, is a tool for connecting and displaying different video materials in chronological order) constitutes an independent hierarchy, and the track labels can be reordered by dragging with the mouse. The content of the top track always covers the lower layer. For example, moving the subtitle track above the video track to ensure that the subtitles are always visible; Shotcut can also be used to adjust by means of mask and composition. That is, in the time axis, the video material added later will default to cover the area of the previous video material. The "promote hierarchy" or "demote hierarchy" can be selected for the video material, or the material can be directly dragged above / below the target track to adjust the priority; Shotcut can also be used to adjust by means of keyframe animation hierarchy control. That is, for dynamic effects containing keyframes (such as zooming, rotating), adjusting its track position can ensure the interaction of the animation effect with other video materials during synthesis.
[0041] In another alternative embodiment, the first video template can also be edited by adjusting the playback order of video materials. The playback order of video materials can be adjusted for any two video materials. Adjusting the playback order of video materials means changing the arrangement position of video materials within the same track to reorganize the narrative logic or timeline structure of the video. For example, moving a piece of dialogue audio before the background audio, i.e., adjusting two video clips; or adjusting the playback order of multiple video materials.
[0042] In the embodiment of the present application, based on the modified rendering rules and / or the modified positional relationship of the first video template, the first video is edited to obtain the edited second video. According to the rendering rules and / or the modified positional relationship modified by the user for the first video, a video with different visual effects or structures is regenerated, which is called the second video. For example, in the original first video, the subtitles are located at the bottom layer and are blocked by the background image, resulting in unclear subtitles. By modifying the positional relationship, the subtitles are moved to the upper layer to make them more clearly visible, thus generating the second video.
[0043] In an alternative embodiment, in response to the modification operation on the first video template, one implementation method for determining the modified rendering rules in any target materialized template includes: in response to the positioning operation on the first video template, displaying any target materialized template to be modified in the editing area; in response to the code editing operation in any target materialized template, obtaining the modified rendering rules. Among them, the positioning operation is that when the user selects any material type to be modified through interface interaction, the system loads the corresponding any target materialized template into the editing area and visualizes the rendering rules of the current any target materialized template. After loading the corresponding any target materialized template into the editing area, the current rendering rules of the target materialized template can be seen in the editing area. For example, if the video material type is subtitles, the target materialized template can include the font size, color, transparency, etc. of the subtitles. The rendering rules can be adjusted by directly modifying the code or the attribute values of the corresponding attributes of the video material in the rendering parameter information. For example, if the target material template is a background image, the user can directly modify the "filter: none" of the target material template as the background image in the first video template code to "filter: blur(2px)", indicating that the background image changes from clear to slightly blurred to highlight the foreground content (such as a digital human or subtitles); or, the user changes the code value of the background image transparency from "0.5" to "0.8" in the editing area, which can make the background image clearer and is used to emphasize the background image content.
[0044] Optionally, in response to a modification operation on the first video template, determining another implementation manner of the modified rendering rule in any target materialized template includes: in response to a triggering operation for modifying any target materialized template in the first video template, presenting it in the first editing window; in response to a first editing operation in the first editing window, determining the name of any target materialized template and the modified rendering rule; in response to a first editing submission operation, refreshing the first video template according to the name of any target materialized template and the modified rendering rule. Among them, the triggering operation is an action for the user to start editing. For example, double-clicking on the background image on the timeline or right-clicking and selecting "Edit Parameters". The first editing window is an interface in the video editing software for modifying the target materialized template, which may include input boxes for entering the corresponding attributes or attribute values of the video material in the rendering parameter information, as well as sliders or drop-down menus, etc. For example, the editing window of the target materialized template with the video material type of background image can display a transparency adjustment slider. The first editing operation is a specific modification operation performed by the user on the target materialized template in the first editing window, so as to determine the name of the modified target materialized template and the modified rendering rule. For example, dragging the slider to change the transparency from 0.5 to 0.8, determining that the name of the target materialized template is the background image template, and the modified rendering rule is opacity: 0.8. The first editing submission operation is an operation for the user to confirm the modification after completing the editing. For example, the first editing submission operation can be clicking the "Save" or "Submit" button. Refreshing refers to the process of updating the rendering parameter information corresponding to the first video template according to the name of the target materialized template and the modified rendering rule, and refreshing the preview of the video effect after the first video editing in real time.
[0045] For example, a user is using video editing software to edit a first video template that includes a background image, subtitles, and a digital human. The user hopes to adjust the display effect of the subtitles to make them more prominent. The specific operations are as follows: The user selects the target materialized template of the subtitles in the video editing software, and the system detects this trigger operation; The system displays a first editing window that includes the name of the subtitle materialized template (such as "subtitle template 1") and the current rendering rules (such as the font color is white and the font size is 24px); The user modifies the rendering rules of the subtitles in the first editing window, changes the font color to yellow, the font size to 32px, and adds a shadow effect; The system detects the user's editing operation, determines that the name of the target materialized template is "subtitle template 1", and records the modified rendering rules (font color, size, and shadow effect); The user clicks the "Save" button to submit the modification, and the system refreshes the first video template according to the modified rendering rules, and the display effect of the subtitles is updated accordingly to obtain a second video. Through the above operations, the user successfully adjusts the display effect of the subtitles and generates a second video that meets the requirements. This process realizes efficient and accurate modification of the video template through real-time feedback for a single video materialized template and immediate refreshing after submission. Without manually modifying the underlying code, the user can intuitively adjust the rendering rules through a graphical interface, improving the efficiency and diversity of video production, reducing the dependence on professional equipment and manual intervention, and meeting the diverse video needs of high efficiency and low cost.
[0046] In an optional embodiment, an implementation manner for determining the modified positional relationship between any two target materialized templates in response to a modification operation on the first video template includes: In response to a dragging operation on any target materialized template in the first video template, move any target materialized template according to the dragging trajectory of the dragging operation, and determine another target materialized template according to the position when the dragging operation ends; Fill any materialized template into the structural position where another target materialized template is located, and fill another target materialized template into the structural position of any target materialized template to obtain the modified positional relationship between any two target materialized templates. Wherein, the dragging operation refers to the user dragging a target materialized template in the editing area. For example, the operation of the user dragging any target materialized template through a mouse or touch gesture is used to adjust its position. The system moves the visual position of the target materialized template in real time according to the dragging trajectory. When the user's dragging operation ends, the system can determine another selected target materialized template according to the final position when the dragging operation ends. For example, when the user drags the background image template in the editing area, when the user's dragging operation ends, another selected target materialized template, such as the subtitle template, is determined according to the final release position.
[0047] Among them, the structural position refers to the position of the corresponding target materialized template in the first video template, which is a placeholder preset in the target materialized template and used to identify the position where the target materialized template can be inserted. Each structural position corresponds to the filling of the target materialized template of a certain material type. The positional relationship between multiple structural positions can reflect the hierarchical relationship between multiple video materials. The information of the structural position can include the number or name of the structural position, and the information of the structural position determines the occlusion logic of the target materialized template. For example, the information of the structural position is that layer_1 (background layer) is at the bottom layer, the information of the structural position is that layer_2 (middle layer) covers it, and the information of the structural position is that layer_3 (foreground layer) is at the top layer.
[0048] After determining any two target materialized templates, any one of the materialized templates can be filled into the structural position where the other target materialized template is located, and the other target materialized template can be filled into the structural position of any one of the target materialized templates to achieve the structural position exchange of any two target materialized templates, so as to obtain the modified positional relationship between any two target materialized templates. Among them, the positional relationship is the hierarchical relationship or track order of any two target materialized templates in the video, which can determine the occlusion logic of the two target materialized templates in the video. Finally, the hierarchy or track position of any two target materialized templates is reversed, thereby changing the occlusion logic of any two target materialized templates in the finally generated second video. For example, the occlusion logic can be reflected as one target materialized template covering another target materialized template.
[0049] For example, a user is using video editing software to edit a first video template that includes a background image, subtitles, and a digital human. The user wishes to place the background image above the subtitles. The specific operations are as follows: The user can perform a drag operation to move a selected target materialized template, which is templatized as a "background image template", upward; when the user drags the background image template to the position where the "subtitle template" is located and releases the mouse, the system determines that another target materialized template is the "subtitle template"; it is determined that the structure bit information of the background image template is layer 1 and the structure bit information of the subtitle template is layer 2, and the structure bit information of the background image template being layer 1 is swapped with the structure bit information of the subtitle template being layer 2, obtaining the modified positional relationship between the two target materialized templates, that is, the new structure bit information of the background image template is "layer 2" and the new structure bit of the subtitle template is "layer 1". Finally, in the generated second video, the background image will cover the subtitles, resulting in the subtitles being blocked. This process allows the user to intuitively exchange the structure bits of the target materialized templates through visual drag-and-drop operations, thereby modifying their positional relationships. This interaction method simplifies the steps of complex hierarchical adjustments, enabling the user to quickly reverse the layer order without manually editing code, improving the efficiency of video production, reducing the dependence on professional equipment and manual intervention, and meeting the requirements for high-efficiency and low-cost videos.
[0050] Optionally, another implementation method for determining the modified positional relationship between any two target materialized templates in response to a modification operation on the first video template includes: in response to a trigger operation for adjusting the positional relationship between any two target materialized templates in the first video template, a second editing window is displayed; in response to a modification operation in the second editing window, the names of any two target materialized templates and the structure bit information after mutual swapping are determined; in response to a modification submission operation, the first video template is refreshed according to the names of any two target materialized templates and the structure bit information after mutual swapping to obtain the modified positional relationship between any two target materialized templates. Among them, the names of all target materialized templates and the current structure bit information can be displayed in the second editing window. And, in the second editing window, there is a dropdown menu or checkbox for the user to select all the target materialized templates that can be swapped. In addition, there is an input box in the second editing window for the user to add new target materialized templates.
[0051] Among them, for the modification operation in the second editing window, the specific modification operation method for determining the names of any two target materialized templates and the information of the structural positions after mutual swapping is not limited. For example, a menu or checkbox that can be pulled down in the second editing window and contains all target materialized templates that can be swapped for the user to select can be used to perform the modification operation in visual mode, select any two target materialized templates, so as to determine the names of any two target materialized templates and the information of the structural positions after mutual swapping; alternatively, an input box that contains an input box for the user to add new target materialized templates in the second editing window can be used to perform the modification method in manual mode, input the names of any two target materialized templates and the information of the corresponding new structural positions, so as to determine the names of any two target materialized templates and the information of the structural positions after mutual swapping. In the embodiments of the present application, regarding the response to the modification submission operation, refreshing the first video template according to the names of any two target materialized templates and the information of the structural positions after mutual swapping to obtain the modified positional relationship between any two target materialized templates has been described in detail in the foregoing embodiments and will not be elaborated herein.
[0052] For example, the user hopes to swap the subtitle template from layer 2 to layer 1 to cover the background image. The user can select the "background image template" and "subtitle template" in the editing interface in visual mode and click the "swap positions" button. The following shows the names of all target materialized templates and the information of the current structural positions in the second editing window:
[0053] Target materialization template Information of the current structure bit Background map template Layer 1 Subtitle template Layer 2 Digital human template Layer 3
[0054] The user selects the "subtitle template" and "background image template" and selects the "swap positions" function in the drop-down menu.
[0055] Optionally, the user can also manually input the name of target materialized template A: subtitle template → new structural position: layer 1, and the name of target materialized template B: background image template → new structural position: layer 2. The system determines that the modified target materialized templates are the subtitle template and the background image template, and their new structural positions are subtitle → layer 1 and background image → layer 2.
[0056] After the user clicks "submit", the hierarchical relationship in the first video template is updated to generate the second video, that is, the subtitle template is located in layer 1 (top layer) and covers the background image template (layer 2). Through the interaction design of the second editing window, this process not only supports the selection in visual mode but also supports manual input, allowing the user to flexibly adjust the structural positions of the two target materialized templates, thereby changing their positional relationship. This design takes into account the intuitiveness and flexibility of the operation, is applicable to complex hierarchical management scenarios, can accurately update the first video template, and ensures that the modified hierarchical relationship takes effect in real time.
[0057] In an optional embodiment, downloading the first video template to a local terminal device includes: in response to an access operation to a batch video generation result page, displaying the batch video generation result page, on which access links of target video templates corresponding to N video instance identifiers respectively on a CDN network or a server, access links of a set of video materials corresponding to N video instance identifiers respectively on the CDN network or the server, and access links of videos corresponding to N video instance identifiers respectively on the CDN network or the server are displayed; N is an integer greater than or equal to 2; in response to a trigger operation on the access link of any target video template on the CDN network or the server, taking any target video template as the first video template and downloading it from the CDN network or the server to the local terminal device. Wherein, the batch video generation result page is a page that centrally displays batch-generated videos and their related information.
[0058] In the embodiment of the present application, the N video instance identifiers are unique identifiers generated according to a batch generation task, and the N video instance identifiers are different from each other. Since each video instance identifier can be used to represent a video to be generated, and each video instance identifier has a corresponding target video template. Therefore, for the N video instance identifiers, access links of target video templates corresponding to the N video instance identifiers respectively on the CDN network or the server can be displayed on the batch video generation result page.
[0059] Among them, a video to be generated has a related set of video materials, and this set of video materials is used to generate the video corresponding to the video instance identifier. And each video instance identifier is convenient for tracking the generation process of the video to be generated. That is to say, this video instance identifier can also be used as the unique identity identifier of the video to be generated during the generation process and corresponds to the generation of a video. Therefore, for the N video instance identifiers, access links of a set of video materials corresponding to the N video instance identifiers respectively on the CDN network or the server and access links of videos corresponding to the N video instance identifiers respectively on the CDN network or the server can be displayed on the batch video generation result page.
[0060] Among them, the method for generating the N video instance identifiers is not limited. For example, it includes but is not limited to numbers and strings, etc. For example, it can be an increasing sequence starting from any integer with a fixed step size to obtain N integers as the N video instance identifiers; or it can also be N preset strings, etc.
[0061] In an optional embodiment, for any video instance identifier, when the multiple video materials corresponding to the video instance identifier include a target subtitle, a target audio, a green screen video of a target digital human, and a target background image, on the batch video generation result page, the access links of the target subtitle, the target audio, the green screen video of the target digital human, and the target background image on the CDN network or the server, the video corresponding to the video instance identifier, and the access link of the target video template corresponding to the video instance identifier on the CDN network or the server are respectively displayed.
[0062] For example, the user clicks the "View Generation Results" button to enter the batch video generation result page. The page displays three columns of information for 3 video instances:
[0063]
[0064] In an optional embodiment, as shown above, the batch video generation result page also includes the access links of the video materials. Based on this, in addition to performing local secondary editing on the first video, multiple video materials can also be downloaded from the CDN network or the server to the local terminal device for local synthesis of a new video.
[0065] In an optional embodiment, in response to a trigger operation on the access links of multiple video materials on the CDN network or the server, the multiple video materials are downloaded from the CDN network or the server to the local terminal device; wherein, the multiple video materials are distributed in one or more groups of video materials; the video editing software on the running terminal device is used to import the multiple video materials into the video editing software, and the multiple video materials are displayed in the editing area of the video editing software; in response to the editing operation on the multiple video materials, a third video is generated, and according to the attribute information and hierarchical relationship of the multiple video materials after editing, a second video template corresponding to the third video is generated, and the second video template includes the rendering rules of the multiple video materials and the positional relationship of the multiple video materials in the second video template.
[0066] Among them, the user can download multiple video materials from the CDN network or the server to the local terminal device through a trigger operation. The downloaded multiple video materials can be distributed in one or more groups of video materials, and the user can select the video materials in the same group or different groups for download.
[0067] For example, the first group of video materials includes:
[0068]
[0069] The second group of video materials includes:
[0070]
[0071] The background image in the first group of video materials can be selected: Provide background pictures with a wedding theme (such as a church, a garden, a starry sky, etc.) and subtitle templates: Preset wedding subtitle styles (such as the dynamic text effect of "Forever&Always") as multiple video materials corresponding to the third video and download them from the CDN network; alternatively, the background image in the first group of video materials can be selected: Provide background pictures with a wedding theme (such as a church, a garden, a starry sky, etc.) and the audio materials in the second group of video materials: Wedding background music (such as piano music, symphony, pop song clips) and subtitle templates: Romantic-style subtitle animations (such as the subtitle effect of petals falling) as multiple video materials corresponding to the third video and download them from the CDN network. In the embodiments of the present application, the user can freely select different groups of video materials for combination according to the creation requirements, without being restricted by the grouping of video materials, improving the flexibility of video generation; moreover, rich types of video materials (such as videos, audios, background pictures, subtitles, etc.) are provided in the CDN network or the server, and the user can freely match them to create unique video content. Through this flexible video material download and combination function, the required video materials can be efficiently obtained from the CDN network or the server and personalized creation can be carried out on the local terminal device to meet the diverse video production requirements.
[0072] After the download is completed, run the video editing software on the terminal device. The selected multiple video materials can be imported using the video editing software on the terminal device and displayed and edited in the editing area. Among them, after the user performs editing operations on these materials, a new video can be generated, that is, the final video file generated according to the multiple video materials after the editing operation is the third video. At the same time, according to the edited material attribute information and hierarchical relationship, the system will generate a corresponding second video template. The second video template contains the rendering rules of multiple video materials and their positional relationships in the second video template, so that similar videos can be quickly generated based on this second video template in the future.
[0073] In the embodiments of the present application, regarding generating the third video in response to the editing operation on multiple video materials, the specific implementation principle is the same as that of editing the first video in response to the modification operation on the first video template to obtain the edited second video in the foregoing embodiments, and will not be elaborated here.
[0074] Among them, according to the attribute information and hierarchical relationship of the multiple video materials after editing, a second video template corresponding to the third video is generated. The system automatically records the rendering rules and positional relationships regarding the third video after editing to form a reusable second video template for quickly generating similar videos in the future.
[0075] Scenario example: The user creates multiple product advertising videos, and each video requires a different background and subtitles, but shares a unified animation effect.
[0076] First, the user downloads the following video material types from the CDN network:
[0077] Background group: Background 1.mp4 (beach), Background 2.mp4 (city) Subtitle group: Product A subtitle.srt, Product B subtitle.srt Animation group: Product display animation.mp4 (the same animation template)
[0078] The user imports three video materials into the video editing software, drags "background1.mp4" to the bottom layer of the timeline, adds "Product A subtitle" to the top layer, and sets the animation transparency to 70%. Finally, an advertising video containing a beach background, Product A subtitle, and semi-transparent animation is generated. The system saves the edited settings (such as background transparency, subtitle position) to form a new template, namely the second video template. This process combines distributed material acquisition and local editing to generate video templates, achieving efficient and flexible video production: the user can flexibly call different types of video materials from the CDN network, make personalized combinations to obtain the second video template, enhancing the diversity of video generation; in addition, new videos can be quickly generated through local editing and saved as video templates, improving the efficiency of video production, reducing the dependence on professional equipment and manual intervention, and meeting the requirements for efficient and low-cost videos.
[0079] Next, a detailed description is given of an implementation manner for batch video generation based on multiple materialized templates provided in the embodiments of the present application.
[0080] In an optional embodiment, in response to input operations on the video generation page for the number of videos and video categories, a batch video generation task is generated. The batch video generation task includes the number of videos N and the video category; according to the batch video generation task, N video instance identifiers are generated, and a set of video materials related to the video category is generated for each video instance identifier; multiple materialized templates obtained by variable processing of the initial video template are acquired. The initial video template includes the rendering rules of multiple video materials required for generating the video and the hierarchical relationship between multiple video materials; for each video instance identifier, at least one target materialized template is determined from the multiple materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target materialized template is combined to obtain the target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationship of a set of video materials; video generation processing is performed according to the target video templates and a set of video materials corresponding to the N video instance identifiers respectively to obtain N videos under the video category.
[0081] In this embodiment, the execution subject of the above video batch generation method is not limited. For example, this method can be implemented as a service product, which can adopt a client-server architecture. In the case of adopting a client-server architecture, on the one hand, a video generation page for batch video generation is provided to the user on the client side to receive the user's input operations for the number of videos and video categories, and then a batch video generation task is initiated to the server. On the other hand, in response to the client initiating a batch video generation task through the video generation service page on the server side, the server batch generates videos through resources such as computing resources, network bandwidth, and storage resources of the server, which is beneficial to improving the speed of video generation.
[0082] For another example, as the processing power of the corresponding hardware device of the client increases, the above method can also be executed by the client. The client can provide a video generation page to the user and, in response to the user's input operations for the number of videos and video categories on the video generation page, generate a batch video generation task, and then batch generate videos through resources such as the client's own computing resources and storage resources. Among them, when the batch generation task is deployed to be executed on the client side, there is no need to transmit data to the server, which can save network latency.
[0083] In this embodiment, the batch video generation task includes the number of videos N and video categories, where N is an integer greater than or equal to 2. The video category is used to describe the expression theme of the video content to be batch generated, and the specific implementation of the video category is not limited. For example, it includes but is not limited to: product introduction, teaching explanation, knowledge popularization, beauty and skincare, etc. Optionally, the video category can be implemented as a single-level video category. For example, it can be implemented as business registration or legal consultation, etc. Optionally, the video category can also be implemented as a multi-level video category. For example, the first-level video category can be business registration; the second-level video categories of this first-level video category can be cleaning or food business, etc.
[0084] In this embodiment, according to the batch generation task, N video instance identifiers are generated, and the N video instance identifiers are all different. Each video instance identifier can be used to uniquely represent a video to be generated, so as to facilitate tracking the required video materials and video templates of the video to be generated. That is to say, this video instance identifier can also be used as the unique identity identifier of the relevant content (such as video materials) of the video to be generated.
[0085] Among them, the method for generating N video instance identifiers is not limited. For example, it includes but is not limited to numbers and strings, etc. For example, it can be an increasing sequence starting from any integer with a fixed step size to obtain N integers as the N video instance identifiers; or, it can also be N preset strings, etc.
[0086] In this embodiment, a set of video materials related to the video category is generated for each video instance identifier, and this set of video materials is used to generate the video corresponding to the video instance identifier. Among them, each set of video materials includes video materials of at least one material type. In some embodiments of the present application, the video materials of one material type are simply referred to as one type of video material.
[0087] Among them, the material type refers to the media resource type that constitutes the video, including but not limited to: material types such as audio, video, background image, subtitle, digital human, etc.
[0088] In this embodiment, the implementation manner of generating a set of video materials related to the video category for each video instance identifier is not limited.
[0089] In an alternative implementation manner, for any video instance identifier, video materials can be randomly extracted from multiple material types stored in the basic material library, and at least one video material extracted is used as a set of video materials for the video instance identifier. Among them, the basic material library stores multiple video materials under multiple video categories.
[0090] In another alternative implementation manner, for any video instance identifier, the semantic similarity between at least one video material in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the video materials that meet the similarity conditions in the multiple similarity information are used as a set of video materials for the video instance identifier.
[0091] Furthermore, in this embodiment, multiple materialized templates obtained by variable processing of the initial video template are acquired.
[0092] In the embodiments of the present application, the initial video template can be an existing video production framework. Among them, the video production framework can be preset or exported after video editing by an editing software. In the embodiments of the present application, the specific manner of obtaining the initial video template is not limited. For example, the initial video template can be a preset general video template provided by the system by default, or a custom video template created by the user according to needs, or a video template imported from other external sources.
[0093] Among them, the initial video template includes the rendering rules of various video materials required for generating a video, as well as the hierarchical relationship between various video materials. In the embodiments of the present application, the rendering rule of each video material is used to describe the rendering logic followed by this video material during the video rendering process, so as to achieve the expected rendering effect of this video material in the generated video. Among them, the hierarchical relationship formed between various video materials in the initial video template refers to the layer stacking order of various video materials during the video generation process in the initial video template, which is used to determine the front-back covering relationship and occlusion logic of various video materials during video generation.
[0094] In this embodiment, part of the variable processing of the initial video template is reflected in taking the material type as the splitting variable, structurally splitting the initial video template and templatizing the rendering parameter information to obtain multiple materialized templates. Among them, there is a corresponding relationship between the materialized template and the material type. For example, the material types include materials such as audio, video, background image, subtitle, digital human, etc. Then the types of the corresponding materialized templates include but are not limited to: audio materialized template, video materialized template, background image materialized template, subtitle materialized template, digital human materialized template, etc. In other words, one materialized template corresponds to each material type, and each materialized template is used to describe the rendering rule of this video material.
[0095] In this embodiment, there is no limitation on the timing of the variable processing. For example, the initial video template can be variablized in advance. Another example is that the initial video template can also be variablized dynamically. For the detailed content on how to perform the variable processing, reference can be made to the subsequent embodiments.
[0096] In this embodiment, the materialized template after variable processing can be modularly reorganized. For each video instance identifier, at least one target materialized template is determined from multiple materialized templates according to the material types in a group of video materials corresponding to this video instance identifier. There is a corresponding relationship between the material types in this group of video materials and the materialized templates.
[0097] For example, in the case where the material types of the audio, video, background image, and subtitle included in this group of video materials, the target materialized templates can include the target materialized templates corresponding to the audio, video, background image, and subtitle respectively; another example is that in the case where the material types of the audio, video, background image, subtitle, and digital human included in this group of video materials, the target materialized templates can include the target materialized templates corresponding to the material types of the audio, video, background image, subtitle, and digital human respectively.
[0098] Further, in the case of obtaining at least one target materialized template corresponding to a set of video materials, based on the hierarchical relationship among multiple video materials included in the initial video template, combine at least one target materialized template to obtain the target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationship of a set of video materials.
[0099] Among them, the hierarchical relationship among multiple video materials included in the initial video template can represent the layer stacking order of multiple materialized templates and can be used to organize and integrate the target materialized templates, so as to form a target video template with a clear hierarchical relationship and corresponding rendering rules for video materials. The target video template describes the hierarchical relationship of a set of video materials and is consistent with the hierarchical relationship of this set of video materials in the initial video template.
[0100] Among them, in this embodiment, the N video instance identifiers respectively correspond to their own target video templates, that is to say, each set of video materials has its own corresponding target video template, enriching the types of target video templates used for batch video generation. Based on the rendering rules described in the N target video templates, perform video generation processing on the grouped video materials corresponding to the N video instance identifiers, so that the style of each video generated in batch matches the video template used respectively, improving the diversification degree of the video content generated in batch.
[0101] In the case of obtaining the target video template corresponding to the video instance identifier, perform video generation processing according to the target video templates respectively corresponding to the N video instance identifiers and a set of video materials to obtain N videos under the video category. Video generation processing refers to the process of filling and rendering video materials for the target video template based on the target video template corresponding to each video instance identifier and a set of video materials to obtain the video corresponding to the video instance identifier. For example, for the N video instance identifiers, a set of video materials respectively corresponding to the N video instance identifiers can be filled into the N target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain the videos under the video category.
[0102] In an optional embodiment, batch video generation processing can be performed on the video generation corresponding to the N video instance identifiers. Each batch performs parallel video generation processing on M videos, where M is less than N and M is an integer. After the M videos are generated, continue to perform video generation processing on the subsequent M videos until all the videos corresponding to the N video instance identifiers are processed. Through batch video generation processing, the resource utilization rate of the server is improved, and at the same time, the server pressure overload caused by high concurrency is avoided.
[0103] In the embodiments of the present application, in batch video generation, materialized templates of various material types obtained by variable processing of an initial video template are used for recombination of the materialized templates to obtain a video template required for generating a video. Further, a video instance identifier is generated for each video, and the video instance identifier is bound to a corresponding set of video materials. According to the set of video materials corresponding to the video instance identifier, a target materialized template for recombination is determined from various materialized templates, and in combination with the hierarchical relationship between the materialized templates provided by the initial template, the target materialized template is organized and integrated to obtain a video template corresponding to each video instance identifier for video generation processing of a corresponding set of video materials for each video instance identifier. Since the video template can be obtained by personalized recombination in combination with the video materials for generating the video, the flexibility of video generation is improved. In addition, the recombined video template has a corresponding relationship with the video instance identifier, which enriches the types of video templates to a certain extent, so that the style of each video generated in batch matches the adopted video template respectively, and the diversification degree of the batch-generated video content is improved.
[0104] In an alternative embodiment, the target video templates, a set of video materials, and videos corresponding to N video instance identifiers can be uploaded to a CDN network or a server for storage. In this way, it can be ensured that the target video templates, a set of video materials, and videos corresponding to N video instance identifiers can be quickly accessed and distributed, while reducing the storage pressure on local terminal devices. The CDN network can efficiently distribute the target video templates, a set of video materials, and videos corresponding to N video instance identifiers to nodes near users, improving the access speed and stability. Server storage can centrally manage the target video templates, a set of video materials, and videos corresponding to N video instance identifiers, supporting larger-scale batch video generation requirements. By storing the target video templates, video materials, and videos in the CDN network or the server, users can obtain these resources at any time through an access link for further editing or viewing, thereby improving the efficiency and flexibility of video generation.
[0105] In the embodiments of the present application, the method of variable processing will be introduced in detail in subsequent embodiments.
[0106] In the embodiments of the present application, the initial video template is subjected to variable processing to obtain various materialized templates. During batch video generation, for each video instance identifier, at least one target materialized template is determined from various materialized templates according to the material type in the set of video materials corresponding to the video instance identifier.
[0107] In an optional embodiment, when determining at least one target materialization template from multiple materialization templates according to the material types in a group of video materials corresponding to a video instance identifier, it includes: identifying at least one material type included in the group of video materials corresponding to the video instance identifier; selecting at least one initial materialization template from multiple materialization templates according to at least one material type included in the group of video materials, with each material type corresponding to one initial materialization template; and adjusting the parameters of at least some of the selected at least one initial materialization templates to obtain at least one target materialization template.
[0108] In this embodiment, the materialization templates included in the multiple materialization templates obtained by variable processing of the initial video template are called initial materialization templates. In the case of determining at least one initial materialization template corresponding to a certain grouping from multiple initial materialization templates, the rendering rules described by the initial materialization templates can be called initial rendering rules. Furthermore, the parameters of at least some of the at least one initial materialization templates are adjusted to obtain at least one target materialization template. Each initial materialization template obtains a corresponding target materialization template after parameter adjustment. The target materialization template after parameter adjustment includes target rendering rules, and the target rendering rules are different from the initial rendering rules described by the initial materialization templates.
[0109] In this embodiment, since the initial materialization template is obtained by variable processing, the variable processing can convert the fixed rendering parameters in the initial video template into variable parameters that can be dynamically assigned values to achieve flexible configuration of the template content. If the values of the optional parameters are different, the rendering rules will be different. Optionally, each initial materialization template includes at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values and default parameter values.
[0110] In this embodiment, the parameter adjustment of the initial materialization template is to obtain multiple target materialization templates with different candidate parameter values by assigning different candidate parameter values to the variable parameters of the initial materialization template. Among them, the target materialization template for a certain material type can be used for rendering video materials in different groupings of video instance identifiers. If the parameter values of the variable parameters are different, it means that the rendering rules described by the target materialization templates of the same material type are different, so that the rendering results of the batch-generated videos are different, and differences are formed between different videos, improving the richness of video content. Hereinafter, how to adjust the parameters of the initial materialization template will be introduced.
[0111] In an optional embodiment, when adjusting parameters of at least part of at least one initial materialized template to obtain at least one target materialized template, the following steps are included: determining the number of templates with parameters to be adjusted, where the number of templates is less than or equal to the number of at least one initial materialized template; selecting the initial materialized templates to be adjusted from at least one initial materialized template according to the number of templates to be adjusted; determining the variable parameters to be adjusted from the initial materialized templates to be adjusted; randomly determining target parameter values from multiple candidate parameter values associated with the variable parameters to be adjusted; and assigning the target parameter values to the variable parameters to be adjusted to obtain the target materialized template.
[0112] In this embodiment, there is no limitation on the method for determining the number of templates with parameters to be adjusted. For example, if parameter adjustment is performed on each initialized video template, the number of templates to be adjusted is the number of at least one initial materialized template. Another example is that, according to a random number generation algorithm, a random integer within a preset range is generated as the number of templates to be adjusted, and the preset range means less than or equal to the number of at least one initial materialized template. There is no limitation on the random number generation algorithm, including but not limited to: the Linear congruential generator (LCG) and the Mersenne Twister, etc.
[0113] Further, select the initial materialized templates to be adjusted from at least one initial materialized template according to the number of templates to be adjusted. In an optional embodiment, select the initial materialized templates to be adjusted according to the priority of the material types and the number of templates to be adjusted. The priority of the material type refers to the importance degree of the material type to the presentation effect of the generated video. For example, the subtitle material type generally has a lower importance degree to the presentation effect, so the priority of the subtitle material type can be set to a lower priority; relatively, the priority of the background image can be higher than that of the subtitle, so it can be set to a medium priority; and, the digital human has a higher importance degree to the presentation effect, so it can be set to a higher priority. Then, according to the high or low priority, preferentially select the templates with parameters to be adjusted from the initial material templates with a higher priority of the material type. If the number of these initial material templates is less than the previously determined number of templates to be adjusted, then it can be further selected from the initial materialized templates with a lower priority of the material type. The finally selected number of templates to be adjusted is less than or equal to the number of at least one initial materialized template.
[0114] Further, determine the variable parameters to be adjusted from the initial materialized template to be adjusted. In an alternative embodiment, all variable parameters in the initial materialized template to be adjusted can be used as the variable parameters to be adjusted. In another alternative embodiment, the target variable parameters in the initial materialized template to be adjusted are used as the optional parameters to be adjusted, and the target variable parameters are pre-selected optional parameters.
[0115] Furthermore, randomly determine the target parameter values from multiple candidate parameter values associated with the variable parameters to be adjusted. In some embodiments, among the at least one initial materialized template corresponding to each of the N video instance identifiers, there are the same variable parameters to be adjusted, and the multiple candidate parameter values associated with the variable parameters to be adjusted can be randomly used as the target parameter values for each of the N video instance identifiers, so that the optional parameter values of the N video instance identifiers are as different as possible.
[0116] Further, assign the target parameter values to the variable parameters to be adjusted to obtain the target materialized template.
[0117] In the case of obtaining the target materialized template, combine the target materialized templates to obtain the target video template. The embodiment does not limit the combination method. Two combination methods are provided below, but are not limited thereto.
[0118] In an alternative embodiment, according to the hierarchical relationship between multiple video materials included in the initial video template, generate a basic video template. This basic video template serves as the framework of the target video template and includes multiple blank structure positions corresponding to multiple video materials. Among them, the blank structure position is a placeholder for the target materialized template preset in the basic video template, used to identify the position where the target materialized template can be inserted, and each blank structure position corresponds to the filling of the target materialized template of a certain material type. The positional relationship between multiple blank structure positions reflects the hierarchical relationship between multiple video materials. As described in the above embodiment, the hierarchical relationship represents the front-to-back stacking order of multiple materials in the generated video. In some embodiments, this hierarchical relationship is extracted from the initial video template; or, it can also be preset, preset based on at least one target materialized template, that is to say, the hierarchical relationship can be set as needed. For example, the structure position corresponding to the subtitle is located in the upper layer, the structure position corresponding to the digital human is located in the middle layer, and the structure position corresponding to the background image is located in the lower layer. Further, insert at least one target materialized template into the corresponding blank structure position in the basic video template to obtain the target video template corresponding to the video instance identifier.
[0119] In another alternative embodiment, according to at least one target materialization template, the structural positions where the rendering rules of video materials of the same material type in the initial video template are overwritten to obtain the target video template corresponding to the video instance identifier. Among them, the difference between the structural position and the above-mentioned blank structural position is that the rendering rule of each video material in the initial video template occupies a structural position, and the blank structural position is empty. Among them, the positional relationship between the structural positions reflects the hierarchical relationship between multiple video materials.
[0120] In the case of obtaining the target video template, video generation processing is performed according to the target video template corresponding to each video instance identifier and a set of video materials. In an alternative embodiment, when performing video generation processing according to the target video templates corresponding to N video instance identifiers and a set of video materials to obtain N videos under the video category, it includes: for each video instance identifier, filling the set of video materials corresponding to the video instance identifier into the target materialization template in the target video template corresponding to the video instance identifier; for the filled target video template, rendering the set of video materials according to the rendering rules and hierarchical relationship of the set of video materials described in the filled target video template to obtain a video under the video category.
[0121] Among them, the target materialization template includes at least one placeholder corresponding to the material type for filling video materials of the material type.
[0122] In an alternative embodiment, generating a set of video materials related to the video category for each video instance identifier includes: obtaining a set of video material description information related to the video category for each video instance identifier, and each set of video material description information includes description information of multiple video materials; for each video instance identifier, according to the set of video material description information corresponding to the video instance identifier, calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously uploading the multiple video materials corresponding to the video instance identifier to the content distribution network; correspondingly, before performing video generation processing according to the target video templates corresponding to N video instance identifiers and a set of video materials to obtain N videos under the video category, it also includes: in response to a batch video generation trigger event, respectively obtaining multiple sets of video materials corresponding to N video instance identifiers from the content distribution network. Hosting video materials through the content distribution network reduces the storage pressure on the server side, enabling the server to efficiently render a large number of videos and improving the user experience.
[0123] In this embodiment, the description information of each group of video materials is used to generate video materials corresponding to the video instance identifier thereof. Among them, each group of video materials includes video materials of multiple material types. The material type refers to the type of different video elements that make up the video content, including but not limited to: audio, video, background image, subtitle, digital human and other material types.
[0124] In this embodiment, the implementation manner of generating the description information of a group of video materials related to the video category for each video instance identifier is not limited.
[0125] In an alternative embodiment, for any video instance identifier, the description information of video materials can be randomly extracted from multiple material types stored in the basic material library, and the extracted description information of multiple video materials is used as the description information of a group of video materials for any video instance identifier. Among them, the basic material library stores the description information of multiple video materials under multiple video categories.
[0126] In another alternative embodiment, for any video instance identifier, the semantic similarity between the description information of multiple video materials in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the description information of the video materials that meet the similarity condition among the multiple similarity information is used as the description information of a group of video materials for this video instance identifier.
[0127] In yet another alternative embodiment, for any video instance identifier, according to the video category, a material description information generation model is called, and this model is used to generate the description information of multiple video materials related to this video category. Among them, this material description information generation model is trained by combining the description information of sample video materials of a large number of different sample video categories and different material types. By learning the semantic correlation between the description information of sample video materials of different sample video categories and different material types, this model can specifically combine different video categories and generate the description information of multiple video materials related thereto.
[0128] Furthermore, for each video instance identifier, according to the description information of a group of video materials corresponding to this video instance identifier, multiple material generation models based on artificial intelligence are called to generate multiple video materials corresponding to this video instance identifier, and the multiple video materials corresponding to this video instance identifier are synchronously uploaded to the content distribution network.
[0129] Among them, one material generation model can generate at least some of the multiple video materials. For example, one material generation model can generate one video material. The following also takes this as an example for illustration, but is not limited thereto.
[0130] In this embodiment, in the case of generating video materials with video instance identifiers, according to multiple video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined. The target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to the video instance identifier. The determination method of the target video template corresponding to each video instance identifier can refer to the above embodiment and will not be elaborated here.
[0131] In this embodiment, in response to a batch video generation trigger event, according to N video instance identifiers, multiple video materials corresponding to the N video instance identifiers are respectively obtained from the content delivery network; video generation is performed according to the multiple video materials and the target video templates respectively corresponding to the N video instance identifiers to obtain N videos under the video category.
[0132] In this embodiment, the specific implementation of the batch video generation trigger event is not limited and can be flexibly configured according to actual application requirements. For example, it can be in the case where all multiple materials corresponding to the N video instance identifiers are generated; or, it can also be in the case where multiple materials corresponding to each video instance identifier are generated. In this case, video generation will be performed on the multiple video materials corresponding to the video instance identifier; or the batch video generation trigger event can also be a preset trigger time. For example, it can be after a period of time in response to an input operation on the video quantity and video category on the video generation page. For example, it can be 2 hours, 1 day, or 1 week, etc. The time span is not limited.
[0133] It should be noted that each video instance identifier corresponds to the generation of one video, and N video instance identifiers can correspond to the generation of N videos. When performing video generation on the multiple video materials respectively corresponding to the N video instance identifiers, the generation process of the video corresponding to each video instance identifier is asynchronous, and the video generation corresponding to each video instance identifier does not affect each other, so as to improve the generation efficiency of the N videos.
[0134] Further optionally, the description information of each group of video materials includes but is not limited to: subtitle description information, audio type description information, digital human description information, and background image description information. Such as Figure 2As shown in the figure, a set of video material description information corresponding to the video instance identifier is used to call multiple material generation models based on artificial intelligence to generate various video materials corresponding to the video instance identifier, including: calling a generative language model according to the subtitle description information to generate text information to obtain a target subtitle; calling a text-to-speech model according to the target subtitle and the audio type description information to convert the target subtitle into a target audio adapted to the audio type description information; calling a multi-modal model according to the target audio and the digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video of the target digital human based on the target audio to obtain a green screen video of the target digital human; calling an image generation model according to the background image description information to generate a background image to obtain a target background image.
[0135] In this embodiment, the APIs (Application Programming Interfaces) of multiple material generation models are associated with endpoints. In this embodiment, the video generation page is the presentation of the front-end code of the endpoint. Among them, the endpoint is used for the generation of video materials, the management of batch video generation tasks, and the control of the video generation process.
[0136] Among them, the endpoint includes front-end code and back-end code, that is, the client-server structure is adopted as described in the above embodiment. The front-end code refers to a video generation page built based on a front-end framework. This video generation page runs on the client and is used to interact with the user, receive the user's input operations, and initiate a batch video generation task to the server. The back-end code of the endpoint runs on the server. The back-end code is obtained by API-ifying the code of an existing video editing software, such as shortcut, etc. Among them, in the existing video editing software, the front-end UI code and the video rendering function code are highly coupled. In this embodiment, it is split according to functions to decouple the front-end UI of the video editing software from the video rendering function code, and then encapsulate the video rendering function code of the existing video editing software into an independent API, so that the external can call the video rendering function through the standard API method to achieve automated video generation.
[0137] As Figure 2 For the endpoint in the figure, in one example, the front-end code of the endpoint can be a video generation page built based on Astro, which is responsible for receiving callback notifications. For example, when the material generation model finishes processing, it can send a notification to the endpoint to notify that the subsequent process of video generation can continue. The back-end code of the endpoint can be obtained by API-ifying the video rendering code of the video editing software and exposed to the outside through the API interface for external calls.
[0138] In an optional embodiment, when calling a generative language model to generate text information according to subtitle description information to obtain a target subtitle, the steps include: matching corresponding keywords according to the video category, calling a pre-designed prompt template, and filling the keywords of the video category into the prompt template to obtain the prompt for the video category; inputting the prompt for the video category into the generative language model to generate text information and obtain a target subtitle related to the video category.
[0139] Among them, the keywords are used to describe the theme content expressed by the video category. Different video categories can correspond to different keywords, and the prompt templates corresponding to different video types can also be different. Among them, the target subtitle includes all the text content required for each video. By calling the generative language model, the target subtitle can be automatically generated based on the video category without manual intervention, improving the generation efficiency.
[0140] Further optionally, according to the target subtitle and audio type description information, call a text-to-speech model to convert the target subtitle into a target audio adapted to the audio type description information. Among them, the target subtitle is used to provide text content, and the audio type description information is used to specify the audio type of the generated target audio. The audio type includes but is not limited to: audio format, audio language, audio tone, audio quality, etc. Any audio type that can be used to specify the sound effect of the target audio is applicable to this embodiment. Based on the target subtitle and audio description information, call the text-to-speech model to generate the corresponding target audio and upload it to the content delivery network to improve the automation degree of audio generation.
[0141] Further, call a multimodal model to generate a green screen video of a virtual human. A virtual human refers to a virtual character generated based on AI technology, which can synchronously simulate the appearance, voice, lip shape and other behaviors of a real person, and can be used in scenarios such as intelligent customer service, short video production, virtual anchors, etc., but is not limited to this. In this embodiment, by combining various video materials corresponding to each video instance identifier, a video with the effect of a real person speaking can be generated.
[0142] In an optional embodiment, the virtual human can be a pre-recorded real person video or picture. In subsequent embodiments, the real person video and picture are collectively referred to as video frames, and the number of video frames can be one or more. In this case, the description information for different virtual humans can be implemented as identification information, and the identification information serves as the unique identity identifier of the virtual human and is used to obtain the video frames of the virtual human corresponding to the identification information. In another optional embodiment, the video frames of the virtual human can be dynamically generated based on a multimodal model. In this case, the description information of the virtual human can be the prompt for generating the virtual human, and the description information of the virtual human corresponding to each video instance identifier can be different.
[0143] Further optionally, when calling a multimodal model according to the target audio and the digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain the green screen video of the target digital human, the following steps are included: obtaining corresponding target digital human video frames based on the digital human description information; calling a multimodal model according to the video frames of the target digital human and the target audio, and performing multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features; wherein, the multi-dimensional speech features include but are not limited to: speech content features and speech emotion features; determining lip movement control parameters of the target digital human according to the speech content features; determining facial expression control parameters of the target digital human according to the speech emotion features; determining limb movement control parameters of the target digital human according to the speech content features and the speech emotion features; generating a green screen video of the target digital human based on the lip movement control parameters, expression control parameters and limb movement control parameters of the target digital human; the lip movement, facial expression and limb movement of the green screen video of the target digital human match the target audio.
[0144] In this embodiment, the speech content features are used to reflect the semantic information in the target audio, such as lexical content, grammatical structure, speech intention and speech rhythm, and are mainly used to drive the lip movement of the digital human to be synchronized with the semantics of the target audio. The speech emotion features are used to reflect the emotional state of the target audio, including but not limited to the intensity of tone, the speed of speech and the change of intonation, etc., and are mainly used to drive the facial expression and limb movement of the digital human.
[0145] Further, determine the lip movement control parameters of the target digital human according to the speech content features; determine the facial expression control parameters of the target digital human according to the speech emotion features; determine the limb movement control parameters of the target digital human according to the speech content features and the speech emotion features.
[0146] In the case of obtaining the lip movement control parameters, facial expression control parameters and limb movement control parameters, drive the target digital human to perform action rendering based on the lip movement control parameters, expression control parameters and limb movement control parameters of the target digital human, and generate the corresponding green screen video of the target digital human, which is convenient for subsequent flexible replacement with the background image. Since the green screen video of the target digital human is generated by controlling the multi-dimensional speech features extracted from the target audio content, it is ensured that the dynamic performance of the lip movement, facial expression and limb movement of the target digital human in the green screen video matches the target audio in terms of semantics, timing and emotion, ensuring the accurate alignment of the lip movement and improving the realism and visual experience of the generated video.
[0147] Further, in this embodiment, the target audio and the target subtitle are aligned, for example, the timestamp of the target subtitle is inferred and aligned and marked through the timestamp of the speech content in the target audio, so as to ensure that the subsequent display of the target subtitle is completely synchronized with the target audio.
[0148] In this embodiment, there are two ways to generate the background image. One is to call the text-to-image model according to the background image description information to generate the background image and obtain the target background image. The other is to directly obtain the pre-generated or captured background image, which is not limited herein.
[0149] In an optional embodiment, a corresponding set of video material generation states is maintained for each of the N video instance identifiers. The generation state of any video material of any material type in each set of video materials includes: ready-to-generate state, generating state, generation success state, and generation failure state. For example, for any video instance identifier, taking the multiple video materials that need to be generated for this video instance identifier, including: target subtitles, target audio, green screen video of the target digital human, and target background image as an example. Then, the generation states of the target subtitles, target audio, green screen video of the target digital human, and target background image can be maintained respectively.
[0150] Among them, for each video instance identifier, mark the multiple video materials corresponding to this video instance identifier as the ready-to-generate state; when calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to this video instance identifier, if any material generation model returns a generating response message, update the video material that should be generated by this material generation model to the generating state; if any material generation model returns a generation success message, update the video material that should be generated by this material generation model to the generation success state; if any material generation model returns a generation failure message, update the video material that should be generated by this material generation model to the generation failure state. As Figure 2 shown, a subscription service is provided. This subscription service can receive notifications from each material generation model to notify the subscription service of the generation state of the video material of this material model, for example, it can be the generation success state. Furthermore, the subscription service will return a generation success message to the server to inform the server that the corresponding video material has been successfully generated. Figure 2 Only the notification process of the target digital human is taken as an example herein, but it is not limited thereto.
[0151] Further optionally, if the generation state of the video material is updated to the generation failure state, a failure reminder message is output to the user who initiated the input operation. The failure reminder message includes the material type of the video material updated to the generation failure state and its corresponding video instance identifier. If a re-generation operation triggered by the user is received, according to the video instance identifier corresponding to the video material, obtain the description information of the corresponding video material, and call the corresponding material generation model based on artificial intelligence to re-generate the corresponding video material.
[0152] In this embodiment, when generating multiple video materials corresponding to each video instance identifier, the multiple video materials corresponding to each video instance identifier can be synchronously uploaded to the content delivery network, and the access links of the video materials corresponding to each video instance identifier in the content delivery network can be obtained. Continuing with the above example, for any video instance identifier, when the multiple video materials corresponding to the video instance identifier include a target subtitle, a target audio, a green screen video of a target digital human, and a target background image, the target subtitle, the target audio, the green screen video of the target digital human, and the target background image are respectively uploaded to the content delivery network, and the access links of the target subtitle, the target audio, the green screen video of the target digital human, and the target background image are obtained.
[0153] Further, in response to a batch video generation trigger event, according to the access links of the multiple video materials corresponding to each of the N video instance identifiers in the content delivery network, the multiple video materials corresponding to the N video instance identifiers are respectively obtained; and, video generation is performed according to the multiple video materials corresponding to each of the N video instance identifiers and a target video template to obtain N videos under the video category.
[0154] Continuing with the above example, in response to a batch video generation trigger event, according to the access links of the target subtitle, the target audio, the green screen video of the target digital human, and the target background image corresponding to each of the N video instance identifiers in the content delivery network, the target subtitle, the target audio, the green screen video of the target digital human, and the target background image corresponding to the N video instance identifiers are respectively obtained; and, video generation is performed according to the target subtitle, the target audio, the green screen video of the target digital human, the target background image, and the target video template corresponding to each of the N video instance identifiers to obtain N videos under the video category.
[0155] Further optionally, the videos corresponding to the N video instance identifiers and the target video template are uploaded to the content delivery network, and the access links of the videos corresponding to the N video instance identifiers and the target video template in the content delivery network are obtained, and the access links are added to the video generation result page; in response to a viewing operation of the video result page, the video generation result page is displayed, and the video generation result page includes the access links of a set of video materials, videos, and the target video template corresponding to at least one video instance identifier among the N video instance identifiers in the content delivery network.
[0156] In this embodiment, the video materials, videos, and target video templates corresponding to each video instance identifier are stored in the content delivery network. The server does not need to store a large number of files, reducing the disk I / O load and ensuring service stability. Further, for the video materials, videos, and target video templates corresponding to each video instance identifier, the access links stored in the content delivery network can avoid data expansion, reduce query pressure, and at the same time support larger-scale data management.
[0157] In this alternative embodiment, in response to a trigger operation for access links of a set of video materials, videos, and / or target video templates corresponding to at least one video instance identifier in a content delivery network, a set of video materials, videos, and / or target video templates corresponding to at least one video instance identifier information is accessed. By hosting videos and video templates through the CDN, the video generation result page can directly load videos from the content delivery network. Compared with pulling videos from the server, it occupies less bandwidth and has a faster loading speed, which can significantly improve performance and is applicable to large-scale video browsing scenarios.
[0158] The following introduces the method of variable processing.
[0159] In an alternative embodiment, using the material type as a splitting variable, the initial video template is parsed to obtain information segments corresponding to multiple material types; from the information segments corresponding to each of the multiple material types, the rendering parameter information corresponding to each of the multiple material types is respectively extracted; and the rendering parameter information corresponding to each of the multiple material types is templated to obtain multiple materialized templates.
[0160] Furthermore, after obtaining the initial video template, using the material type as a splitting variable, the initial video template is parsed. Herein, the material type refers to different types of video materials, which is a classification standard for differentiating different video materials, and the splitting variable refers to the independent information segments obtained by dividing the initial video template according to the material type when parsing the initial video template. For example, if the initial video template contains two types of video materials, namely images and audio, the initial video template will be split into independent information segments of images and audio according to the material type as the classification standard, and subsequent processing will be performed separately.
[0161] Among them, the information segment refers to the information segment extracted from the initial video template and corresponding to the material type, and the information segment may include the rendering parameter information of the corresponding material type. For example, an information segment may include rendering parameter information such as path information, playing time, and special effect application for the material type.
[0162] In the embodiments of the present application, the rendering parameter information is used to describe at least one attribute of a material file of a certain material type. Each attribute can be used as a rendering parameter, and the attribute value of each attribute in the initial video template can be used as the default parameter value of the corresponding rendering parameter. The rendering parameters with default parameter values form a rendering parameter information corresponding to the material file of this material type. Among them, the "material file" refers to various resource files used in the materialized template, mainly including various video materials, such as pictures, videos, and audios, etc. The rendering parameter information can be used to control the rendering logic of the material file, so that the rendering effect presented by the material file in the finally generated video can be obtained. The rendering parameter information quantitatively describes the attributes of the material file, so that the attributes of the video material are converted into the attributes of the rendering parameter information, and the rendering effect of the video material in the finally generated video is controlled through the rendering parameter information. In other words, the attributes of the video material are described as corresponding rendering parameters by the rendering parameter information, and the rendering parameters with default parameter values define the attribute values of a certain attribute of the material file.
[0163] Among them, the rendering parameter information can be extracted separately from the information segments corresponding to various material types. The rendering parameter information corresponding to each material type includes at least one rendering parameter and its corresponding default parameter value, and each rendering parameter is used to describe the attributes of a material file of a certain material type. In the embodiments of the present application, the specific attributes corresponding to the rendering parameters of various material types are not limited.
[0164] In the embodiments of the present application, templatization can convert the rendering parameter information corresponding to various material types from the rendering parameters with fixed configuration default parameter values into a variable parameter set with dynamically replaceable parameter values. Among them, by templatizing the rendering parameter information corresponding to various material types, various materialized templates corresponding to various material types can be obtained. Each materialized template is used to describe the rendering rules of a certain material type. The rendering rules include at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values to facilitate controlling the rendering effect of video generation.
[0165] In the embodiments of the present application, a materialized template is the result of templatizing the rendering parameter information corresponding to various material types. It includes a placeholder for the path information of the material file, as well as variable parameters and multiple candidate parameter values associated therewith. The materialized template describes the rendering rules of the material file corresponding to the material type in the generated video. The rendering rules are used to describe the rendering logic followed by this type of video material during the video rendering process, so as to control the expected rendering effect of this material file in the generated video. Among them, by converting the rendering parameter information from the rendering parameters with fixed configurations and default parameter values into a set of dynamically replaceable variable parameters to form a materialized template, different candidate parameter values can be assigned to the variable parameters to flexibly adjust the rendering effect of the material file in the generated video in different scenarios. Among them, the rendering parameter information is templatized to form a variety of materialized templates that can be modularly recombined. For the materialized templates corresponding to different material types, different target video templates can be recombined, and then diverse videos can be batch-generated without separately designing and adjusting the video templates for each video, which not only saves time and human resources, but also improves the flexibility and content diversity of video template generation, and thus realizes the efficient batch generation of videos with diverse styles.
[0166] For example, for the same materialized template, by assigning different candidate parameter values to the variable parameters, the attribute values corresponding to various attributes of the material type corresponding to this materialized template can be flexibly set, so as to achieve refined control of the rendering rules of the material file. The materialized template improves the flexibility of video generation, can significantly improve the efficiency and quality of video generation, and at the same time meets diverse creative needs.
[0167] In the embodiments of the present application, the specific content of the candidate parameter values associated with the variable parameters obtained after templatizing the rendering parameter information corresponding to various material types is not limited.
[0168] Among them, by associating multiple candidate parameter values with the variable parameters, the variable parameters can be flexibly adjusted according to requirements, improving the flexibility and diversity of video template generation, and thus realizing the efficient batch generation of videos with diverse styles.
[0169] In an alternative embodiment, the initial video template is parsed using the material type as the splitting variable to obtain information segments corresponding to multiple material types, including: loading the XML document corresponding to the initial video template, where the XML document includes a root element and multiple non-root elements connected to the root element, and multiple specific elements are included in the multiple non-root elements, with each specific element used to describe the rendering rules of a material type; starting from the root element, traversing the non-root elements in the XML document to identify multiple specific elements; and extracting the multiple information segments where the multiple specific elements are located as the information segments corresponding to multiple material types. Among them, the XML document of the initial video template can be loaded into memory to form a parsable tree structure. In the embodiments of the present application, the XML document can be loaded into memory through a DOM (Document Object Model) parser to form a tree structure parsing method.
[0170] Among them, the tree structure includes a root element and non-root elements. The root element is the top-level node of the XML document, and the non-root elements are all directly or indirectly nested under the root element. An XML document has one and only one root element, and the root element is the first element of the XML document and can be used as the starting point of the XML document.
[0171] Among them, the non-root elements are child elements directly or indirectly nested under the root element and are used to divide different information segments of the initial video template according to the material type. Multiple specific elements are included in the non-root elements. Starting from the root element, traversing the non-root elements in the XML document can identify multiple specific elements. A specific element is an element that can directly describe the rendering rules of a certain type of material, and each specific element corresponds to a material type.
[0172] By extracting these information segments, the system can quickly identify the attributes of the material file and perform rendering according to their attribute values. This design of extracting information segments significantly improves the flexibility and diversity of video generation, meets the diverse creation needs, and provides a technical basis for efficient batch video generation.
[0173] In an optional embodiment, starting from the root element, non-root elements in the XML document are traversed to identify multiple specific elements, including: S1. Starting from the root element, traverse non-root elements in the XML document; S2. For the currently traversed non-root element, obtain the element tag included in the currently traversed non-root element; S3. If the element tag is a specific tag, determine whether the currently traversed non-root element contains sub-elements; S4. If the currently traversed non-root element contains sub-elements, use the sub-elements as the currently traversed non-root element and return to execute step S2; S5. If the element tag is a non-specific tag, continue to the next non-root element and return to execute step S2; S6. If the currently traversed non-root element does not contain sub-elements, use the currently traversed non-root element as a specific element.
[0174] Among them, in step S1, starting from the root element of the XML document, non-root elements are accessed one by one for traversal. In the embodiments of the present application, the specific implementation strategy of traversal is not limited. For example, the traversal can be implemented using a depth-first search algorithm or a breadth-first search algorithm. In the embodiments of the present application, taking the depth-first search algorithm as an example, the process of traversing non-root elements in the XML document to identify multiple specific elements is described in detail.
[0175] Next, execute step S2. For the currently traversed non-root element, obtain the element tag included in the currently traversed non-root element. Among them, in the XML document, the element tag is the identifier within angle brackets (<>), which is used to mark the type and semantic meaning of the element. Different material types are distinguished by the name of the element tag; the specific position of the material is located through the hierarchical relationship of the element tags.
[0176] After obtaining the element tag included in the currently traversed non-root element, determine whether the element tag is a specific tag. Among them, the specific tag can be a predefined set of key tags, which represent the element tags that need to be specially processed in the XML document and are used to identify the material types to be extracted.
[0177] Next, execute step S3 or S5. If step S5 is executed, that is, the element tag is a non-specific tag, continue to the next non-root element and return to execute step S2.
[0178] If step S3 is executed, that is, the element tag is a specific tag, continue to determine whether the currently traversed non-root element contains sub-elements. Among them, the sub-elements can be other elements nested inside the currently traversed non-root element.
[0179] Next, execute step S4 or S6. If step S4 is executed, that is, the currently traversed non-root element contains sub-elements, then take the sub-elements as the currently traversed non-root element, and return to execute step S2. If step S6 is executed, that is, the currently traversed non-root element does not contain sub-elements, then take the currently traversed non-root element as a specific element. Among them, step S4 is the recursive processing when there are sub-elements. Set the sub-elements of the currently traversed non-root element as the new traversal starting point, and re-execute steps S2 - S6 for each sub-element. After the recursion ends, continue to traverse other non-root elements. Step S6 is the end collection when there are no sub-elements, and take the currently traversed non-root element as a specific element.
[0180] In an optional embodiment, render parameter information corresponding to each of multiple material types is respectively extracted from information segments corresponding to the multiple material types, including: for each information segment, extracting path information of a material file and at least one attribute value of the material file from the information segment, where the attribute value is used to render the material file; reading the material file according to the path information, and determining the material type described by the information segment according to the extension of the material file; taking at least one attribute to which the at least one attribute value belongs as at least one render parameter, and taking the at least one attribute value as the default parameter value of the at least one render parameter, so as to obtain the render parameter information corresponding to the material type described by the information segment.
[0181] Among them, the information segment is a part related to the material type extracted from the XML document, and contains the path information of the material file and at least one attribute value of the material file. Among them, the material file can be used to construct various media files for the finally generated video. In the embodiments of the present application, the material file may include but is not limited to: files in formats such as MP4 and AVI, which are video files containing dynamic images and audio and are used to display a series of continuous pictures; files in formats such as WAV and MP3, which are audio files providing background music, narration, or special effect sounds, etc.
[0182] Among them, at least one attribute value of the material file extracted from the information segment is the specific value corresponding to the attribute, so as to describe the specific state of the attribute. The attribute can describe the rendering effect of the material file. For example, for a material file that is a video file, the attribute can be video duration, video speed, filter effect, etc. If the attribute is video duration, the attribute value can be 2 minutes, 1 hour, 1 day, etc.
[0183] In an alternative embodiment, the rendering parameter information corresponding to the material type described by the information segment includes at least one rendering parameter and the default parameter values of at least one rendering parameter. The default parameter value can be the attribute value of the attribute corresponding to the rendering parameter before quantization. Among them, at least one attribute for extracting at least one attribute value of the material file in the information segment can be used as at least one rendering parameter, and at least one attribute value can be used as the default parameter value of at least one rendering parameter, so as to obtain the rendering parameter information. That is to say, a key-value pair set is composed of the rendering parameter and the default parameter value, which is used to control the rendering logic of the material file in the video to achieve the expected rendering effect.
[0184] In an alternative embodiment, the rendering parameter information corresponding to various material types is templated to obtain various material templates, including: for each material type, at least one variable parameter is selected from the rendering parameter information corresponding to the material type; multiple candidate parameter values are associated with at least one variable parameter; according to the multiple candidate parameter values associated with at least one variable parameter, a material template corresponding to the material type is generated. Among them, the variable parameter is obtained by variable processing of the rendering parameter with a default parameter value. Each variable parameter is associated with multiple candidate parameter values, and different candidate parameter values correspond to different rendering logics, so as to produce different rendering effects in the generated video, and can be used to generate diverse material templates. For each material type, at least one variable parameter is selected from its corresponding rendering parameter information, and these variable parameters can vary within the adjustment range to obtain different material templates. For example, when the material file is a video material, the playback speed and filter effects can be selected as variable parameters. The variable parameter is associated with multiple candidate optional values, which are used to limit the adjustment range of the variable parameter, ensure that the material file of the generated video meets the design requirements, and avoid invalid settings.
[0185] In an alternative embodiment, when selecting at least one variable parameter from the rendering parameter information corresponding to the material type, two selection methods are provided: one is to use all rendering parameters as variable parameters, and the other is to select some rendering parameters as variable parameters according to the weight value. Among them, when choosing to use all rendering parameters as variable parameters, there is no need to compare the weight values, and all rendering parameters can be adjusted as variable parameters.
[0186] In an optional embodiment, the manner of selecting some rendering parameters as variable parameters according to weight values may be to select at least one variable parameter from the rendering parameter information corresponding to the material type, including: configuring the weight values of each rendering parameter for the material type in advance; parsing each rendering parameter from the rendering parameter information corresponding to the material type, and selecting at least one rendering parameter with a weight value greater than the set weight threshold from each rendering parameter as at least one variable parameter. Among them, the weight value can be an importance score assigned to each rendering parameter, used to quantify the influence degree of the rendering parameter on the rendering effect of the material file in video generation, and at least one rendering parameter that has a greater impact on user perception or application goals can be selected as at least one variable parameter. The weight threshold is a predefined critical value. By screening the rendering parameters corresponding to the weight values higher than the weight threshold, the user can select the parameters with weight values higher than the weight threshold as variable parameters, avoiding taking too many irrelevant rendering parameters as variable parameters to prevent problems such as configuration conflicts or rendering logic confusion.
[0187] In the embodiments of the present application, the manner of pre-configuring the weight values of each rendering parameter is not limited. For example, the weight values of each rendering parameter can be pre-configured according to empirical annotation or automatically generated through user behavior data analysis.
[0188] Optionally, selecting at least one variable parameter from the rendering parameter information corresponding to the material type can also be randomly selected according to the set number of variable parameters. In an optional embodiment, generating a materialized template corresponding to the material type according to multiple candidate parameter values associated with at least one variable parameter includes: adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to a preset template file, and adding multiple candidate parameter values associated with at least one variable parameter to the preset template file; and adding a placeholder for carrying the material file corresponding to the material type in the preset template file to obtain a materialized template corresponding to the material type. Among them, the preset template file is a basic video template containing basic configurations, including the basic structure of the video template and some preset rendering parameters. Among them, adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to the preset template file can ensure that the finally obtained materialized template has complete rendering parameter information.
[0189] In this embodiment, on the basis of adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to the preset template file, multiple candidate parameter values associated with at least one variable parameter are added to the preset template file. Among them, the candidate parameter values provide multiple choices, allowing the user or the system to select different candidate parameter values according to needs.
[0190] In an alternative embodiment, based on adding multiple candidate parameter values associated with at least one variable parameter to a preset template file, a placeholder for carrying a material file corresponding to a material type can be added to the preset template file to obtain a materialized template corresponding to the material type. Herein, the placeholder is a reserved position for filling the material file and supports dynamic replacement. For example, the path information of the actual material file corresponding to the material type can be filled. Through the placeholder, different material files can be flexibly replaced without modifying the template structure.
[0191] In the above embodiment, by integrating the rendering parameters of the material type and their default parameter values into a preset basic template, the integrity of the rendering parameter information of the generated materialized template is ensured. Meanwhile, by associating multiple candidate parameter values with the variable parameter, dynamic selection of the parameter values of the variable parameter is realized to flexibly adjust the video style. The embedding of the placeholder further decouples the parameter configuration of the initial video template from the material file and supports dynamic replacement of the material file without modifying the template structure. Thus, it is allowed to assign corresponding candidate parameter values to the variable parameter according to the application requirements to obtain a freely combined materialized template, ultimately significantly improving the flexibility and diversity of video template generation, and then efficiently batch-producing video content with various styles, solving the problems such as cumbersome parameter configuration of video templates, complex adaptation process, and serious homogenization of generated video content in traditional video production.
[0192] Based on obtaining the above multiple materialized templates, batch video generation or single video generation can be performed based on the multiple materialized templates. For the batch video generation scenario, multiple videos with diverse styles and different contents can be generated by using the multiple materialized templates provided in the embodiments of the present application.
[0193] The detailed implementation manners and beneficial effects of each step in the method of this embodiment have been described in detail in the foregoing embodiments and will not be elaborated herein.
[0194] In addition, in some processes described in the above embodiments and the accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The operation numbers such as 11 and 12 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0195] Figure 3A schematic structural diagram of an electronic device provided for an exemplary embodiment of the present application. As Figure 3 shown, the device includes: a memory 34 and a processor 35.
[0196] The memory 34 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method for operating on the electronic device, a first video template, a first video, a target materialization template, etc.
[0197] The memory 34 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0198] The processor 35, coupled to the memory 34, is configured to execute the computer programs in the memory 34 for: downloading the first video template to the local terminal device, where the first video template is a video template for rendering a set of video materials to generate a first video, and the first video template includes: a target materialization template adapted to the material type in the set of video materials among multiple materialization templates, and the multiple materialization templates are obtained by varying an initial video template; each materialization template is used to describe the rendering rules of a video material, and the first video template is used to describe the rendering rules and hierarchical relationship of a set of video materials, and the hierarchical relationship is reflected by the position relationship of the target materialization template in the first video template; running video editing software on the terminal device, importing the first video template into the video editing software, opening the first video template in the editing area of the video editing software, and displaying the first video generated based on the first video template in the preview area of the video editing software; in response to a modification operation on the first video template, determining the modified rendering rules in any target materialization template and / or the modified position relationship between any two target materialization templates, and editing the first video according to the modified rendering rules and / or the modified position relationship to obtain an edited second video.
[0199] In an alternative embodiment, the processor 35 determines the modified rendering rules in any target materialized template in response to a modification operation on the first video template, including: in response to a positioning operation on the first video template, displaying any target materialized template to be modified in the editing area; in response to a code editing operation in any target materialized template, obtaining the modified rendering rules; or in response to a trigger operation for modifying any target materialized template in the first video template, displaying it in the first editing window; in response to a first editing operation in the first editing window, determining the name of any target materialized template and the modified rendering rules; in response to a first editing submission operation, refreshing the first video template according to the name of any target materialized template and the modified rendering rules.
[0200] In an alternative embodiment, the processor 35 determines the modified positional relationship between any two target materialized templates in response to a modification operation on the first video template, including: in response to a dragging operation on any target materialized template in the first video template, moving any target materialized template according to the dragging trajectory of the dragging operation, and determining another target materialized template according to the position when the dragging operation ends; filling any materialized template into the structural position where another target materialized template is located, and filling another target materialized template into the structural position of any target materialized template to obtain the modified positional relationship between any two target materialized templates; or in response to a trigger operation for adjusting the positional relationship between any two target materialized templates in the first video template, displaying a second editing window; in response to a modification operation in the second editing window, determining the names of any two target materialized templates and the information of the structural positions after mutual swapping; in response to a modification submission operation, refreshing the first video template according to the names of any two target materialized templates and the information of the structural positions after mutual swapping to obtain the modified positional relationship between any two target materialized templates; wherein, the structural position refers to the position of the corresponding target materialized template in the first video template.
[0201] In an alternative embodiment, the processor 35 downloads the first video template to the local terminal device, including: in response to an access operation on the batch video generation result page, displaying the batch video generation result page, on which the access links of the target video templates corresponding to N video instance identifiers on the CDN network or the server, the access links of a group of video materials corresponding to N video instance identifiers on the CDN network or the server, and the access links of the videos corresponding to N video instance identifiers on the CDN network or the server are displayed; N is an integer greater than or equal to 2; in response to a trigger operation on the access link of any target video template on the CDN network or the server, using any target video template as the first video template and downloading it from the CDN network or the server to the local terminal device.
[0202] In an alternative embodiment, the processor 35, in response to a triggering operation on access links of multiple video materials on a CDN network or a server, downloads the multiple video materials from the CDN network or the server to a local terminal device; wherein, the multiple video materials are distributed in one or more groups of video materials; runs video editing software on the terminal device, imports the multiple video materials into the video editing software, and displays the multiple video materials in an editing area of the video editing software; in response to an editing operation on the multiple video materials, generates a third video, and generates a second video template corresponding to the third video according to the attribute information and hierarchical relationship of the multiple video materials after editing, where the second video template includes rendering rules of the multiple video materials and the positional relationship of the multiple video materials in the second video template.
[0203] In an alternative embodiment, the processor 35, in response to an input operation on the number of videos and video categories on a video generation page, generates a batch video generation task, where the batch video generation task includes the number of videos N and video categories; according to the batch video generation task, generates N video instance identifiers, and generates a group of video materials related to the video categories for each video instance identifier; obtains multiple materialized templates obtained by variable processing of an initial video template, where the initial video template includes rendering rules of multiple video materials required for generating a video and the hierarchical relationship between the multiple video materials; for each video instance identifier, determines at least one target materialized template from the multiple materialized templates according to the material type in the group of video materials corresponding to the video instance identifier; based on the hierarchical relationship between the multiple video materials included in the initial video template, combines the at least one target materialized template to obtain a target video template corresponding to the video instance identifier, where the target video template is used to describe the rendering rules and hierarchical relationship of a group of video materials; performs video generation processing according to the target video templates and a group of video materials respectively corresponding to the N video instance identifiers to obtain N videos under the video category; uploads the target video templates, a group of video materials, and the videos respectively corresponding to the N video instance identifiers to a CDN network or a server for storage.
[0204] In an alternative embodiment, the processor 35 uses the material type as a splitting variable to parse the initial video template, and obtains information segments corresponding to multiple material types; respectively extracts rendering parameter information corresponding to multiple material types from the information segments corresponding to multiple material types; templates the rendering parameter information corresponding to multiple material types to obtain multiple materialized templates.
[0205] Further, as Figure 3 shown, the electronic device further includes: other components such as a communication component 36, a display 37, a power supply component 38, and an audio component 39. Figure 3Only some components are schematically shown, which does not mean that the computing platform only includes Figure 3 the components shown. Additionally, Figure 3 the components within the dashed box in Figure 3 are optional components rather than mandatory components, and can be determined according to the product form of the working node. The working node in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the working node in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include Figure 3 the components within the dashed box; if the working node in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include
[0206] the components within the dashed box.
[0207] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0208] The above communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a communication standard-based wireless network, such as a WiFi, 2G, 3G, 4G / LTE, 5G, or other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0209] The above display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations.
[0210] The above power supply component provides power to various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.
[0211] The above audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0212] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, Compact Disc Read-Only Memory (CD-ROM), optical memory, etc.) that contain computer-usable program code.
[0213] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0214] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0215] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0216] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and a memory.
[0217] Memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0218] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0219] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0220] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A video editing method, characterized in that: include: Downloading a first video template to a local terminal device, the first video template is a video template used to render a group of video materials to generate a first video, the first video template comprising: a target material template adapted to a material type in the group of video materials among multiple material templates, the multiple material templates being obtained by varying an initial video template; each material template being used to describe a rendering rule for a type of video material, the first video template being used to describe a rendering rule and a hierarchical relationship for the group of video materials, the hierarchical relationship being embodied as a positional relationship between the target material template and the first video template; Running the video editing software on the terminal device, importing the first video template into the video editing software, opening the first video template in the editing area of the video editing software, and displaying the first video generated based on the first video template in the preview area of the video editing software; In response to a modification operation on the first video template, the modified rendering rules in any target materialization template and / or the modified positional relationship between any two target materialization templates are determined, and the first video is edited according to the modified rendering rules and / or the modified positional relationship to obtain an edited second video.
2. The method according to claim 1, characterized in that: In response to the modification operation on the first video template, determining the modified rendering rule in any target materialization template includes: In response to the positioning operation on the first video template, any target materialization template that needs to be modified is displayed in the editing area; in response to the code editing operation in any target materialization template, a modified rendering rule is obtained; or In response to a trigger operation to modify any target materialization template in the first video template, it is displayed in a first editing window; in response to a first editing operation in the first editing window, the name of any target materialization template and the modified rendering rule are determined; in response to the first editing submission operation, the first video template is refreshed according to the name of any target materialization template and the modified rendering rule.
3. The method according to claim 1, characterized in that In response to the modification operation on the first video template, determining the modified positional relationship between any two target materialized templates includes: In response to a drag operation on any target materialization template in the first video template, the target materialization template is moved according to a dragging track of the drag operation, and another target materialization template is determined according to a position when the drag operation is terminated; the target materialization template is filled into a structure position where the other target materialization template is located, and the other target materialization template is filled into a structure position of the target materialization template, so as to obtain a modified positional relationship between any two target materialization templates; or In response to a trigger operation of adjusting the positional relationship between any two target materialization templates in the first video template, a second editing window is displayed; in response to a modification operation in the second editing window, the names of the any two target materialization templates and information of the swapped structural positions are determined; in response to a modification submission operation, the first video template is refreshed according to the names of the any two target materialization templates and information of the swapped structural positions, so as to obtain a modified positional relationship between the any two target materialization templates; The structure position refers to a position describing the corresponding target material template in the first video template.
4. The method according to any one of claims 1 to 3, characterized in that: Download the first video template to the local terminal device, including: In response to an access operation to a batch video generation result page, a batch video generation result page is displayed, wherein the batch video generation result page displays access links of target video templates corresponding to each of the N video instance identifiers on the CDN network or the server, access links of a group of video materials corresponding to each of the N video instance identifiers on the CDN network or the server, and access links of videos corresponding to each of the N video instance identifiers on the CDN network or the server; N is an integer ≥ 2; In response to a triggering operation of an access link of any target video template on a CDN network or a server, the any target video template is used as the first video template and downloaded from the CDN network or the server to a local terminal device.
5. The method according to claim 4, characterized in that Also includes: In response to a triggering operation of accessing links of multiple video materials on a CDN network or a server, downloading the multiple video materials from the CDN network or the server to a local terminal device; wherein the multiple video materials are distributed in one or more groups of video materials; Running the video editing software on the terminal device, importing the multiple video materials into the video editing software, and displaying the multiple video materials in the editing area of the video editing software; In response to the editing operation on the multiple video materials, a third video is generated, and based on the edited attribute information and hierarchical relationship of the multiple video materials, a second video template corresponding to the third video is generated, and the second video template includes the rendering rules of the multiple video materials and the positional relationship of the multiple video materials in the second video template.
6. The method according to claim 4, characterized in that Also includes: In response to an input operation on the video generation page for the number of videos and the video category, a batch video generation task is generated, wherein the batch video generation task includes the number of videos N and the video category; Generate N video instance identifiers according to the batch video generation task, and generate a group of video materials related to the video category for each video instance identifier; Acquire multiple material templates obtained by performing variable processing on an initial video template, wherein the initial video template includes rendering rules of multiple video materials required for generating a video and hierarchical relationships between the multiple video materials; For each video instance identifier, determining at least one target materialization template from the multiple materialization templates according to a material type in a group of video materials corresponding to the video instance identifier; Based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target material template is combined to obtain a target video template corresponding to the video instance identifier, wherein the target video template is used to describe the rendering rules and hierarchical relationship of the group of video materials; Performing video generation processing according to target video templates and a group of video materials corresponding to the N video instance identifiers, so as to obtain N videos under the video category; The target video templates, a group of video materials and videos corresponding to the N video instance identifiers are uploaded to the CDN network or the server for storage.
7. The method according to any one of claims 1 to 3 or 6, characterized in that: Also includes: Taking the material type as a splitting variable, parsing the initial video template to obtain information fragments corresponding to multiple material types; Extracting rendering parameter information corresponding to each of the multiple material types from the information fragments corresponding to each of the multiple material types; The rendering parameter information corresponding to each of the multiple material types is templated to obtain the multiple material templates.
8. An electronic device, characterized in that: include: A processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the processor is enabled to implement the steps in the method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps in the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, enables the processor to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Visual programming cloud intelligent editing system and method thereof
CN112040271A
Method and device for adjusting video combination material
CN113315883A
Video synthesis method and electronic equipment
CN117793271A
Video generation method, device and equipment and computer readable storage medium
CN117939255A
Video editing method and apparatus, and device and medium
WO2025020416A1
Cited By
Video material processing method and device, electronic equipment, storage medium and program product
CN120786118A
Video processing method, video processor, storage medium and program product
CN121194031A
Video modification method and device, equipment, storage medium and product
CN121865066A