Video template language definition method and device based on XML (Extensible Markup Language)

Through the XML-based video template language definition method, the problem of cross-platform sharing and reuse of video templates is solved, the standardized description and dynamic generation of video content are realized, and the efficiency and flexibility of video production are improved.

CN120751214APending Publication Date: 2025-10-03SHENZHEN THINKIVE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511013689.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing video production tools lack unified standards, making it difficult to achieve cross-platform sharing and reuse of video templates. The templates are not dynamic and programmable enough, the separation of editing and rendering leads to complex secondary editing, and it is difficult to build consistency across terminals.

Method used

Adopting the XML-based video template language definition method, by creating a VTL video template, using the template element and the root element for definition, nesting the scene element and the content element, and introducing the logic element, dynamic logic processing and external data binding are realized in conjunction with dynamic calculation expressions.

Benefits of technology

It achieves standardized description of video content, high dynamism, easy secondary editing and unified construction across terminals, and improves the efficiency and flexibility of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751214A_ABST
    Figure CN120751214A_ABST
Patent Text Reader

Abstract

The invention discloses an XML-based video template language definition method and device, and the method comprises the steps: creating a VTL video template, and setting a plurality of XML attributes for a template element for definition; embedding one or more scene elements in the template element through a parent-child relationship of XML (Extensible Markup Language); various content elements are nested in each scene element; and introducing a plurality of logic elements with the same level as the template element, wrapping the template element and the plurality of logic elements by using a root element, and realizing dynamic logic processing and data binding in cooperation with a dynamic calculation expression so as to generate the VTL video template capable of being analyzed and rendered. Through the method provided by the invention, the VTL completely, accurately and flexibly defines a complex video item as a standardized text file by utilizing the structuralization, extensibility and self-description characteristics of the XML, so that comprehensive description from static layout to dynamic content generation is realized; standardized description of the video content is achieved, high dynamic performance is achieved, secondary editing and cross-terminal unified construction are easy, and the efficiency and flexibility of video production are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video templates, and in particular to an XML-based video template language definition method and device. Background Art

[0002] With the rapid development of digital media technology, video content plays an increasingly important role in information dissemination. However, the current video production field faces multiple challenges, particularly in achieving automated video content generation, efficient editing, and cross-platform compatibility. Traditional video production processes often rely on specific software tools and professional skills, resulting in a lack of unified standards for the definition and reuse of video templates. This makes it difficult to adapt to rapidly changing market demands and diverse application scenarios, and also suffers from the following drawbacks:

[0003] 1. Lack of a unified standard for video description: Existing video production tools and platforms often use proprietary video file formats or project file structures to describe video content, lacking an open, universal, and easily parsable standardized video description language. This results in poor interoperability of video templates between different systems, making it difficult to share and reuse templates across platforms.

[0004] 2. Insufficient template dynamism and programmability: Traditional video templates are mostly static presets, making it difficult to dynamically generate content based on external data or complex logic. Personalizing or automatically updating video content often requires complex programming interfaces or secondary development, lacking the ability to directly bind data and process logic at the template level.

[0005] 3. Separation of editing and rendering, making secondary editing complex: In existing video production workflows, template editing and final rendering are typically separated. Once a video is generated, modifying or adjusting it often requires restarting the complex rendering process. Modifying the original template can also be unintuitive, resulting in inefficient secondary editing and difficulty meeting the demands of rapid iteration.

[0006] 4. Difficulty in building consistency across terminals: With the widespread adoption of mobile internet and web technologies, video content must maintain consistent presentation across multiple terminals, including PCs, mobile devices, and web browsers. However, due to the lack of a unified video description standard, building and rendering videos on different terminals often requires the development and maintenance of multiple sets of logic, increasing development costs and maintenance difficulties. Summary of the Invention

[0007] To this end, the present invention aims to at least to some extent address the deficiencies in the prior art, thereby proposing a method and apparatus for defining a video template language based on XML.

[0008] In a first aspect, the present invention provides a method for defining a video template language based on XML, the method comprising:

[0009] Creating a VTL video template, setting multiple XML attributes for a template element to define it, wherein the VTL video template includes the template element and a root element;

[0010] Nest one or more scene elements within the template element using XML parent-child relationships;

[0011] Further nest various content elements within each of the scene elements;

[0012] Introduce multiple logical elements at the same level as the template element, use the root element to wrap the template element and the multiple logical elements at the same level, and use dynamic calculation expressions to implement dynamic logic processing and external data binding based on the multiple logical elements to generate the VTL video template that can be parsed and rendered.

[0013] In a second aspect, the present invention provides an XML-based video template language definition device, the device comprising:

[0014] A definition module is used to create a VTL video template and set multiple XML attributes for the template element for definition, wherein the VTL video template includes the template element and the root element;

[0015] Partitioning module: used to nest one or more scene elements within the template element through XML parent-child relationships;

[0016] Filling module: used to further nest various content elements within each of the scene elements;

[0017] Dynamic module: used to introduce multiple logical elements at the same level as the template element, use the root element to wrap the template element and the multiple logical elements at the same level, and cooperate with dynamic calculation expressions to realize dynamic logic processing and external data binding according to the multiple logical elements to generate the VTL video template that can be parsed and rendered.

[0018] In the third aspect, the present invention also provides an XML-based video template language definition device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps in the XML-based video template language definition method as described in the first aspect.

[0019] In a fourth aspect, the present invention further provides a storage medium storing a computer program, which, when executed, implements the various steps of the XML-based video template language definition method as described in the first aspect.

[0020] The present invention provides an XML-based video template language definition method and device, which includes: creating a VTL video template, setting multiple XML attributes for the template element for definition, wherein the VTL video template includes the template element and the root element; nesting one or more scene elements in the template element through the XML parent-child relationship; further nesting various content elements in each scene element; introducing multiple logical elements at the same level as the template element, using the root element to wrap the template element and the multiple logical elements at the same level together, and coordinating dynamic calculation expressions to realize dynamic logic processing and external data binding according to the multiple logical elements to generate the VTL video template that can be parsed and rendered. Through the method provided by the present invention, VTL adopts XML as its basic data structure, abstracts video content into a series of nested elements and attributes, and accurately describes the various components of the video and their relationships through the XML structure. The VTL language can utilize the structured, extensible and self-describing characteristics of XML to completely, accurately and flexibly define a complex video project as a standardized text file, realizing a comprehensive description from static layout to dynamic content generation; and realizes the structuring, programmability and visual editing of video content, providing a set of standardized, extensible video description specifications that support dynamic content generation and multi-terminal unified construction, so as to realize standardized description of video content, high dynamism, easy secondary editing and unified construction across terminals, thereby improving the efficiency and flexibility of video production. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0022] Figure 1 A schematic diagram of the flow of the XML-based video template language definition method of the present invention;

[0023] Figure 2Schematic diagram of sub-processes in the XML-based video template language definition method of the present invention;

[0024] Figure 3 Another flowchart of the XML-based video template language definition method of the present invention is shown;

[0025] Figure 4 This is a schematic diagram of the program modules of the XML-based video template language definition device of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0027] Please refer to Figure 1 , Figure 1 : is a flow chart of an XML-based video template language definition method according to an embodiment of the present application. In this embodiment, the XML-based video template language definition method includes:

[0028] Step 101: Create a VTL video template and set multiple XML attributes for the template element to define it. The VTL video template includes the template element and the root element.

[0029] In this embodiment, a Video Template Language (VTL) is first created. The VTL language uses XML as its underlying data structure. Because the existing video production field lacks an open and universal video content description standard, VTL uses XML as its foundation to abstract video content into a series of nested elements and attributes. This provides a unified, semantic description paradigm for the video's structure, elements, attributes, timeline, and interaction logic. This enables different systems and tools to create, parse, edit, and render video templates based on the same set of standards, greatly improving the interoperability and portability of video content. This standardization avoids the compatibility issues and data silos associated with traditional proprietary formats, enabling cross-platform interoperability and portability of VTL video templates. Furthermore, as an XML language decoupled from the rendering engine, VTL can be uniformly parsed and rendered in different terminal environments (Web, desktop, server), ensuring consistent presentation of video content across multiple terminals and reducing the complexity and maintenance costs of multi-terminal development. Specifically, the core root element of VTL, template, is used as a starting point. The definition is completed by setting a series of XML attributes for the template element. The key XML attributes include: the `width` and `height` attributes define the video resolution (e.g., width="1920" height="1080"); the `fps` attribute defines the video frame rate (e.g., fps="30"); the `backgroundColor` attribute defines the global background color of the video; and the `videoCodec` and `audioCodec` attributes preset the video and audio encoder standards. This step establishes the top-level container and basic technical specifications for the entire video. Much like preparing a canvas with a fixed size and background for painting, the global context of the video is uniquely determined, providing a benchmark for the placement and rendering of all subsequent elements.

[0030] The template element, as the core root node of VTL, defines the global metadata and properties of the entire video, such as the video's width (width), height (height), frame rate (fps), background color (backgroundColor), video codec (videoCodec), and audio codec (audioCodec). All specific video content elements are nested within this node.

[0031] Furthermore, the creating of the VTL video template includes:

[0032] The VTL video template is combined with the VTM function library. Specifically, the structured design of the VTL video template is combined with the VTM (Video Template Model) function library to support the front-end visual editor to intuitively edit the XML template and save it directly. The VTM function library also provides the ability to parse VTL into an operable Template object instance. This provides a solid foundation for the front-end visual editor (such as ave-editor-view), allowing users to intuitively operate the XML structure on the graphical interface to achieve a what-you-see-is-what-you-get editing experience. At the same time, the editor's modifications can be directly reflected in the XML file, ensuring the consistency and traceability of the template. This design allows non-professional users to easily perform secondary editing and modification of the video template, greatly reducing the threshold for video production, achieving a close combination of editing and rendering, greatly improving the efficiency and convenience of secondary editing, and reducing modification costs. Moreover, VTL, as a plain text XML description language, is decoupled from the specific rendering engine. This means that as long as there is a VTM function library and a corresponding rendering engine, the VTL template can be parsed and rendered in different terminal environments (such as web browsers, desktop applications, and server-side rendering services), achieving "one-time definition, multi-terminal construction, and unified presentation" of video content. This solves the problem of poor cross-platform compatibility in traditional video production and reduces the complexity and maintenance costs of multi-terminal development. The VTM function library contains rendering components (such as text rendering, animation engines), logic tools (such as data formatting, mathematical calculations), and resource management modules (such as font loading and audio playback). The template element in the VTL video template can call the content in the VTM function library for rendering, calculation, and other functions.

[0033] Step 102: nest one or more scene elements within the template element through XML parent-child relationships.

[0034] In this embodiment, the hierarchical relationship of XML ( <template> <scene / > < / template> ) clearly expresses the subordinate relationship of "template containing scenes." Each scene element defines an independent video segment through its attributes. This step macro-segments the video along the time dimension, constructing a narrative structure similar to a movie "storyboard." It also establishes the overall timing and segmentation structure of the video. The timeline of all subsequently added content elements will be calculated relative to the starting point of the scene element to which they belong, achieving modular management of the time dimension.

[0035] Furthermore, each of the scene elements is defined as an independent video segment through its attributes, including duration attributes and transition-related attributes.

[0036] In this embodiment, the scene element represents an independent segment or "storyboard" in the video. Each scene can have its own attributes such as duration, background color, and transition effect; various video content elements can be nested inside the scene, and the time points of the internal elements are relative to the starting time point of the scene; the modular design of the scene makes the organization and editing of video content clearer. Among them, the `duration` attribute accurately defines the duration of the scene (in milliseconds). The total duration of the video is the sum of the `duration` of all scenes. `transition` related attributes (such as `transition-type` and `transition-duration`) define the transition effect between the current scene and the previous scene, achieving a smooth connection between scenes.

[0037] Step 103: Further nest various content elements in each of the scene elements.

[0038] In this embodiment, various content elements, such as text, image, video, and audio, are nested within each scene element, forming a deeply nested XML structure of template -> scene -> content element. This step populates the divided scenes with specific visual and auditory content, accurately describing its temporal and spatial properties, style, and animation.

[0039] The design of VTL elements follows the principle of modularization. Each content element (such as <text> 、 、 <scene>) have clear responsibilities and configurable properties, and these content elements can be called in the VTM function library. This modular design not only improves the readability and maintainability of the template, but also facilitates the future expansion of the VTL language itself. Developers can easily define new VTL elements or expand the properties of existing elements according to new requirements without modifying the core parsing logic, ensuring the vitality and adaptability of the language. Among them, the text element ( <text>): Used to display text content in the video, supporting rich style attributes such as font, font size, color, weight, alignment, line height, word spacing, as well as entry / exit effects (enterEffect, exitEffect) and path animation (pathEffect). It also supports rich text and special character escape. Image element ( ): Used to display static or dynamic images (such as WebP, GIF, APNG) in the video, supporting attributes such as position, size, transparency, rotation, scaling, and entry / exit effects. Can be used as a background image. Audio element ( <audio>): Used to add background music or sound effects to the video, supporting attributes such as volume, loop playback, fade in and fade out. Video element ( <video>): Used to embed other video clips in the video, supporting properties such as playback control, volume, and looping.

[0040] Furthermore, various content elements are further nested within each of the scene elements, including:

[0041] Each of the content elements is comprehensively defined using its rich XML attribute set, wherein the XML attribute set includes spatial relationship attributes, temporal relationship attributes, visual style attributes and dynamic behavior attributes.

[0042] In this embodiment, spatial relationship attributes: `x` and `y` attributes define the precise coordinates of the element on the scene's 2D canvas; `width` and `height` define its size; `zIndex` defines its stacking order on the Z axis to resolve the occlusion relationship between elements. Time relationship attributes: `startTime` and `endTime` attributes define the life cycle of the element within the timeline of the scene to which it belongs, that is, when it appears and when it disappears. Visual style attributes: provide specific style attributes for different elements, such as <text>`fontFamily` (font), `fontColor` (color), `fontSize` (font size), etc., or `opacity` (transparency), `borderRadius` (rounded corners), etc. Dynamic behavior properties: Through compound properties such as `enterEffect` (entry effect), `exitEffect` (exit effect), and `pathEffect` (path animation), you can precisely define complex dynamic behaviors of elements, such as fade-in, fly-in effects, or movement along a preset path. These effects can also be called from the VTM function library.

[0043] Step 104: Introduce multiple logical elements at the same level as the template element, use the root element to wrap the template element and the multiple logical elements at the same level, and use dynamic calculation expressions to implement dynamic logic processing and external data binding based on the multiple logical elements to generate the VTL video template that can be parsed and rendered.

[0044] In this embodiment, the VTL also includes a root element, which serves as the top-level container of the entire VTL document and is mainly used to process <template>When an element coexists with other peer logical elements, specifically, the root element wraps the template element and multiple peer logical elements to ensure the legality of the XML format. After the root (root) element wraps the template element and multiple peer logical elements, a dynamic calculation expression is used for data binding, so that the final video template can be dynamically generated. This step can transform the description of the static VTL video template into an "intelligent template" that can respond to data and execute logic, that is, generate a parsable and renderable VTL video template to achieve the automatic generation of video content.

[0045] Specifically, if only the template element exists, the root element is omitted.

[0046] The following is an example code of XML:

[0047] <template id="aca794b35e714e5e8e2319c77368fd1e" name="Test Template" width="1920" height="1080" aspectRatio="16:9">

[0048] <scene id="22oe8RJLPcxiiLce" width="1920" height="1080" aspectRatio="16:9" duration="5000">

[0049] <text id="20e9Ksna7pYa3t" x="623.863" y="420.6015" width="677" height="238.922" enterffect-type="wordBounceUp" enterffect-duration="1000" exitEffect-type="fade" exitEffect-duration="500" backgroundColor="transparent" startTime="0" endTime="5000" fontFamily="Youshe_title_Black" fontSize="172" fontColor="#ededed" lineHeight="1.189" textAlign="center">Test Text< / template> < / text>

[0050] <image id="22oeB30S0KJAlrLe"name="as.jpeg"width="1920"height="1080" enterEffect-type="enlarge" enterEffect-duration="5000"exiteffect-type="fade" exitEffect-duration="500" isBackground="true"startTime="0”src=" / ave / upload / 3 / 115 / image / 20220730 / 10af9087d56740daa903565d0e521b1f.jpeg" / >

[0051] < / video> < / audio> < / text> < / scene>

[0052] <template>

[0053] In summary, the VTL language uses XML as its basic data structure, abstracting video content into a series of nested elements and attributes. Leveraging XML's structured, extensible, and self-describing characteristics, the VTL language can completely, accurately, and flexibly define a complex video project as a standardized text file, enabling comprehensive descriptions from static layout to dynamic content generation. This choice is based on the following advantages of XML:

[0054] 1. Structural and hierarchical: XML's tree structure is naturally suitable for describing the hierarchical and inclusion relationships between scenes and elements (such as text, images, and audio) in a video.

[0055] 2. Readability and maintainability: The semantic nature of XML tags makes VTL templates highly readable and easy for humans to understand and maintain.

[0056] 3. Scalability: XML allows custom tags and attributes, providing a flexible mechanism for future functional expansion of VTL.

[0057] 4. Universal parsing: There are a large number of mature XML parsers and tools that facilitate VTL parsing and processing in different platforms and language environments.

[0058] An embodiment of the present application provides an XML-based video template language definition method, which includes: creating a VTL video template, setting multiple XML attributes for the template element for definition, wherein the VTL video template includes the template element and the root element; nesting one or more scene elements in the template element through XML parent-child relationships; further nesting various content elements in each of the scene elements; introducing multiple logical elements at the same level as the template element, using the root element to wrap the template element and the multiple logical elements at the same level together, and coordinating dynamic calculation expressions to implement dynamic logic processing and external data binding based on the multiple logical elements to generate a parsable and renderable VTL video template. Through the method provided by the present invention, VTL adopts XML as its basic data structure, abstracts video content into a series of nested elements and attributes, and accurately describes the various components of the video and their relationships through the XML structure. The VTL language can utilize the structured, extensible and self-describing characteristics of XML to completely, accurately and flexibly define a complex video project as a standardized text file, realizing a comprehensive description from static layout to dynamic content generation; and realizes the structuring, programmability and visual editing of video content, providing a set of standardized, extensible video description specifications that support dynamic content generation and multi-terminal unified construction, so as to realize standardized description of video content, high dynamism, easy secondary editing and unified construction across terminals, thereby improving the efficiency and flexibility of video production.

[0059] Further, please refer to Figure 2 , Figure 2 This is a sub-flow diagram of the XML-based video template language definition method in an embodiment of the present application. In this embodiment, the logical elements at the same level as the template element include the data element, the source element, and the script element. The introduction of multiple logical elements at the same level as the template element, and the use of the root element to wrap the template element and the multiple logical elements at the same level together, and the use of dynamic calculation expressions to implement dynamic logic processing and external data binding based on the multiple logical elements include:

[0060] Step 201: Wrap the template element, data element, source element, and script element with the root element;

[0061] Step 202: At the same time, JavaScript Lambda expressions are embedded in the attribute values ​​or text content of the template element to implement dynamic logic processing and binding of external data.

[0062] In this embodiment, specifically, in the context of the same level as the template element, the data element, source element, and script element are introduced, and the root element is used to wrap them together with the template element. Among them, the data element (data declaration element): allows constants or static data sets to be declared in the template, and these constants can be referenced through JavaScript Lambda expressions at any location in the VTL template; this provides the template with the ability to inject static data. The source element (external data source element): supports the definition of one or more external data sources (such as RESTful API) in the template, which is used to dynamically pull data and inject it into the template; this enables the VTL template to be seamlessly integrated with external data systems to achieve data-driven video content generation; it also supports alternative sources and competitive loading mechanisms, which improves the robustness of data acquisition. The script element (script extension element): allows JavaScript code snippets to be embedded in the template to define reusable functions or complex logic; these functions can be called in the Lambda expression of the VTL template, greatly enhancing the flexibility and expressiveness of the template. Among them, the logical elements also include form elements (form elements), which are used to generate user interaction forms on the front-end interface, allowing users to enter data or make selections, and inject these data into the VTL video template, which facilitates manual secondary editing and parameterized generation.

[0063] At the same time, in the attribute value or text content of the template element, use the `{{}}` syntax to embed JavaScript Lambda expressions. JavaScript Lambda expressions give VTL video templates the ability to perform dynamic logic processing and external data binding, thereby converting the static VTL video template description into an "intelligent template" that can respond to data and execute logic, thereby realizing the automatic generation of video content.

[0064] Further, please refer to Figure 3 , Figure 3 This is another sub-flow diagram of the method for defining an XML video template language in an embodiment of the present application. In this embodiment, the {{}} syntax is used to embed JavaScript Lambda expressions in the attribute value or text content of the template element to implement dynamic logic processing and external data binding, including:

[0065] Step 301: Set the "compile="true" attribute in the template element embedded with the JavaScript Lambda expression, which can trigger the compilation engine;

[0066] Step 302: Based on the compilation engine, dynamic logic processing and calculation of external data are implemented according to the logic elements through the JavaScript Lambda expression, and the obtained calculation result is directly replaced at the location of the JavaScript Lambda expression.

[0067] In this embodiment, when a VTL video template is processed, elements with the `compile="true"` attribute set will trigger the compilation engine, such as whether to compile the template at runtime or mark a block as requiring dynamic compilation. The `compile` attribute is used to mark a portion of the template as requiring dynamic parsing and execution, rather than static rendering. This allows for dynamic logic, such as conditional judgments and loops. JavaScript Lambda integration, for example, allows Lambda functions to be defined in a VTL video template as callbacks for processing data. When external data is passed in, the compilation engine calls these Lambda functions to process the data and then renders the results into the template, achieving dynamic data binding. For example, in a VTL video template, it is necessary to dynamically generate play items based on an external video list. The `compile` attribute is used to mark the loop area, and a Lambda function is used internally to process the title and duration of each video item. This achieves both dynamic logic (looping) and binding to external data (video list). The `compile` attribute enables a dynamic compilation context, allowing JavaScript Lambda expressions to be executed within this context, accessing external data and dynamic logic within the template, thereby simultaneously achieving dynamic logic processing and data binding. In this embodiment, the compilation engine will execute these JavaScript Lambda expressions - whether referencing variables declared in the data element, calling functions defined in the script element, or processing data pulled from the source element, the JavaScript Lambda expressions will perform dynamic logical processing and calculations based on these logical elements, and the calculation results will directly replace the location of the JavaScript Lambda expression, thereby dynamically generating the final video content. For example: <text> {{stockData.name}}< / text> `The company name can be dynamically displayed based on the stock data obtained from the API.

[0068] Traditional video templates are usually static and difficult to adapt to the needs of dynamic content generation. VTL video templates implement dynamic logic processing and external data binding within the template by introducing the compile attribute and supporting JavaScript Lambda expressions, making VTL video templates highly programmable and able to dynamically generate video content based on external data and complex logic, far exceeding traditional static templates. This means that VTL video templates are not just static layout descriptions, but also programmable "smart templates" that can dynamically generate content based on external data and complex logic. For example, through Lambda expressions, you can perform calculations, conditional judgments, loops, and other operations directly in XML attributes or content. Combined with <data>and <source> Elements enable data-driven video content generation, which is difficult to achieve directly with existing template technologies.

[0069] Furthermore, the XML-based video template language definition method proposed in this application implements the VTM function library, which serves as the core of VTL language parsing, verification, construction, and compilation. It has undergone rigorous unit testing and integration testing to ensure the correctness of the VTL language specification and the stability of the parsing process. Specific verification includes:

[0070] XML structure validity verification: Verify whether the VTL XML structure complies with the W3C XML specification, and whether the definitions of custom tags and attributes are clear and unambiguous, and can accurately express video content and logic.

[0071] Verification of parsing and building functions: Tests the VTM function library's ability to parse various complex VTL templates, including nested elements, attribute inheritance, special character processing, etc., to ensure that XML documents can be correctly converted into the Template object model in memory.

[0072] Compilation and data binding verification: This focuses on verifying the compilation mechanism of JavaScript Lambda expressions in VTL templates, including the correct execution of logic such as constant references, external data source loading, function calls, conditional judgments, and loops, to ensure that data can be accurately injected into the template and generate expected results.

[0073] Extensibility verification: By trying to extend the VTL language, such as adding new video element types or custom attributes, verify whether the VTL definition method can flexibly adapt to future functional requirements without large-scale modifications to the core parsing logic.

[0074] Multi-terminal compatibility verification: Verify the consistency of parsing and processing of VTL templates in different environments (such as Web front-end editors and Node.js back-end rendering services) to ensure the implementation of "one-time definition, multi-terminal construction".

[0075] Experimental results show that the VTL language definition method proposed in this invention is robust, flexible and efficient, and can effectively support the standardized description, dynamic generation and visual editing of video content, providing a solid technical foundation for automated video production; the VTL language specification, as the underlying standard model, has been applied in short video content creation products such as the Xiaocaizi Financial Intelligence Platform, Xunjian Mini Program, and Xiaosi Miaochuang.

[0076] The XML-based video template language definition method proposed in this application, as a general video content description specification, can be widely used in the following fields:

[0077] Automated video generation platform: Provides core template definition capabilities for various automated video generation systems, enabling them to dynamically generate short videos of various content based on data; such as daily market reports, hot news, and other video content.

[0078] Online video editing tool: As the underlying template language of web or desktop video editors, it supports users to create, modify and manage video templates through a visual interface, achieving efficient secondary editing.

[0079] Cross-terminal video creation: enables video creation to achieve unified template description and rendering on different terminals (web, mobile, desktop, server, etc.), supporting cross-terminal synchronization and creation.

[0080] AI intelligent video creation: Using VTL as the standard video description format file, combined with AI technologies such as TTS / ASR, digital human technology, and multimodal large model technology, a more controllable AI video creation platform is built.

[0081] Furthermore, the embodiment of the present application also provides an XML-based video template language definition device 400. Figure 4 Schematic diagram of program modules of an XML-based video template language definition apparatus in an embodiment of the present application. In this embodiment, the XML-based video template language definition apparatus 400 includes:

[0082] Definition module 401: used to create a VTL video template, and set multiple XML attributes for the template element to define it, wherein the VTL video template includes the template element and the root element;

[0083] The division module 402 is used to nest one or more scene elements within the template element through the parent-child relationship of XML;

[0084] Filling module 403: used to further nest various content elements in each scene element;

[0085] Dynamic module 404: used to introduce multiple logical elements at the same level as the template element, use the root element to wrap the template element and the multiple logical elements at the same level, and cooperate with dynamic calculation expressions to realize dynamic logic processing and external data binding according to the multiple logical elements to generate the VTL video template that can be parsed and rendered.

[0086] An embodiment of the present application provides an XML-based video template language definition device 400, which can achieve: creating a VTL video template, setting multiple XML attributes for the template element for definition, wherein the VTL video template includes the template element and the root element; nesting one or more scene elements in the template element through the XML parent-child relationship; further nesting various content elements in each of the scene elements; introducing multiple logical elements at the same level as the template element, using the root element to wrap the template element and the multiple logical elements at the same level together, and cooperating with dynamic calculation expressions to realize dynamic logic processing and external data binding according to the multiple logical elements to generate the VTL video template that can be parsed and rendered. Through the method provided by the present invention, VTL adopts XML as its basic data structure, abstracts video content into a series of nested elements and attributes, and accurately describes the various components of the video and their relationships through the XML structure. The VTL language can utilize the structured, extensible and self-describing characteristics of XML to completely, accurately and flexibly define a complex video project as a standardized text file, realizing a comprehensive description from static layout to dynamic content generation; and realizes the structuring, programmability and visual editing of video content, providing a set of standardized, extensible video description specifications that support dynamic content generation and multi-terminal unified construction, so as to realize standardized description of video content, high dynamism, easy secondary editing and unified construction across terminals, thereby improving the efficiency and flexibility of video production.

[0087] Furthermore, the present application also provides an XML-based video template language definition device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the various steps in the above-mentioned XML-based video template language definition method.

[0088] Furthermore, the present application also provides a storage medium on which a computer program is stored. When the computer is executed by a processor, the computer implements the various steps in the above-mentioned XML-based video template language definition method.

[0089] The functional modules in various embodiments of the present invention may be integrated into a single processing module, each module may exist physically separately, or two or more modules may be integrated into a single module. The integrated modules may be implemented in the form of hardware or software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may be stored in a computer-readable storage medium.

[0090] Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for causing a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0091] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0092] For those skilled in the art, according to the concept of the embodiments of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.< / data> < / template> < / text>

Claims

1. A method for defining a video template language based on XML, characterized in that: The method comprises: Creating a VTL video template, setting multiple XML attributes for a template element to define it, wherein the VTL video template includes the template element and a root element; Nest one or more scene elements within the template element using XML parent-child relationships; Further nest various content elements within each of the scene elements; Introduce multiple logical elements at the same level as the template element, use the root element to wrap the template element and the multiple logical elements at the same level, and use dynamic calculation expressions to implement dynamic logic processing and external data binding based on the multiple logical elements to generate the VTL video template that can be parsed and rendered.

2. The method according to claim 1, characterized in that The step of creating a VTL video template includes: The VTL video template is combined with the VTM function library.

3. The method according to claim 1, characterized in that Each of the scene elements is defined as an independent video segment through its attributes, including duration attributes and transition-related attributes.

4. The method according to claim 1, wherein Various content elements are further nested within each scene element, including: Each of the content elements is comprehensively defined using its rich XML attribute set, wherein the XML attribute set includes spatial relationship attributes, temporal relationship attributes, visual style attributes, and dynamic behavior attributes.

5. The method according to claim 1, wherein The logical elements at the same level as the template element include data elements, source elements, and script elements; the introduction of multiple logical elements at the same level as the template element, and the use of the root element to wrap the template element and the multiple logical elements at the same level together, and the use of dynamic calculation expressions to implement dynamic logic processing and external data binding based on the multiple logical elements, including: Wrap the template element and the data element, source element, and script element through the root element; At the same time, JavaScript Lambda expressions are embedded in the attribute values ​​or text contents of the template element to implement dynamic logic processing and binding of external data.

6. The method according to claim 5, characterized in that Embedding JavaScript Lambda expressions in the attribute value or text content of the template element to implement dynamic logic processing and external data binding includes: Setting the "compile="true"" attribute in the template element embedded with the JavaScript Lambda expression to trigger the compilation engine; Based on the compilation engine, dynamic logic processing and calculation of external data are implemented according to the logic elements through the JavaScript Lambda expression, and the obtained calculation result directly replaces the location of the JavaScript Lambda expression.

7. The method according to claim 5, characterized in that The step of introducing multiple logical elements at the same level as the template element and using the root element to wrap the template element and the multiple logical elements at the same level includes: If only the template element exists, the root element is omitted.

8. An XML-based video template language definition device, characterized in that: The device comprises: A definition module is used to create a VTL video template and set multiple XML attributes for the template element for definition, wherein the VTL video template includes the template element and the root element; Partitioning module: used to nest one or more scene elements within the template element through XML parent-child relationships; Filling module: used to further nest various content elements within each of the scene elements; Dynamic module: used to introduce multiple logical elements at the same level as the template element, use the root element to wrap the template element and the multiple logical elements at the same level, and cooperate with dynamic calculation expressions to realize dynamic logic processing and external data binding according to the multiple logical elements to generate the VTL video template that can be parsed and rendered.

9. An XML-based video template language definition device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the processor implements the steps of the XML-based video template language definition method according to any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the XML-based video template language definition method according to any one of claims 1 to 7 is implemented.