Video template variable method, device, storage medium and program product
By parsing the material types and template-izing the rendering parameters of video templates, the problem of high learning barriers and monotonous generation styles of video editing software is solved, enabling efficient batch generation of diverse videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video editing software suffers from high learning curves, high labor costs, and monotonous, homogenized video styles generated based on preset templates, resulting in poor flexibility.
By obtaining an initial video template, using the material type as a splitting variable for parsing, rendering parameter information of various material types is extracted and templated to generate multiple material templates, thereby improving the flexibility and diversity of video generation.
It enables efficient batch generation of videos with diverse styles, improving the flexibility and content diversity of video generation while reducing the learning threshold and labor costs.
Smart Images

Figure CN120186432B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to a video template variable method, device, storage medium and program product. BACKGROUND
[0002] In the field of video generation, video editing software can be used to manually edit materials to obtain a video that meets individual needs. However, video editing software has a certain learning threshold, and the use of artificial cost is high, which is difficult to realize batch video generation.
[0003] In order to reduce the learning threshold and improve the efficiency of video generation, in some video editing software, a preset template is provided, and the text or image in the preset template is allowed to be replaced. The user's material is replaced with the text or image in the preset template to generate the video required by the user. Among them, the way of generating a new video by directly replacing the text or image in the preset template simplifies the process of video generation, and is suitable for batch generation of videos. However, the video generated based on the preset template is monotonous in style and highly homogenized, and the flexibility of video generation is poor. SUMMARY
[0004] Embodiments of the present application provide a video template variable method, device, storage medium and program product to improve the flexibility and content diversity of video generation, and to realize efficient batch generation of diversified style videos.
[0005] The embodiment of the present application provides a video template variable method, which comprises: obtaining an initial video template, the initial video template comprising rendering rules of a plurality of video materials required for generating a video and a hierarchical relationship between the plurality of video materials; analyzing the initial video template with a material type as a split variable to obtain information segments corresponding to a plurality of material types; extracting rendering parameter information corresponding to the plurality of material types from the information segments corresponding to the plurality of material types, respectively; and templateizing the rendering parameter information corresponding to the plurality of material types to obtain a plurality of material templates, each material template being used to describe the rendering rules of a video material.
[0006] The embodiment of the present application also provides an electronic device comprising a processor and a memory, the memory storing a computer program, when the computer program is executed by the processor, the processor can realize each step of the video template variable method provided by the embodiment of the present application.
[0007] The embodiment of the present application also provides a computer readable storage medium storing a computer program, when the computer program is executed by the processor, the processor can realize each step of the video template variable method provided by the embodiment of the present application.
[0008] The embodiment of the present application further provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, enable the processor to implement each step in the video template variable method provided by the embodiment of the present application.
[0009] In the embodiment of the present application, by acquiring an initial video template, and taking the material type as a split variable, the initial video template is parsed to obtain information segments corresponding to multiple material types; further, rendering parameter information corresponding to each of the multiple material types is extracted from the information segments corresponding to each of the multiple material types respectively; the rendering parameter information corresponding to each of the multiple material types is templated to obtain multiple materialized templates, and each materialized template is used to describe the rendering rule of a video material. The multiple materialized templates are applied to video generation, so as to improve the flexibility and content diversity of video generation, and further to realize efficient batch generation of diversified style videos. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, illustrate the exemplary embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:
[0011] Figure 1 A flowchart of a video template variable method provided for the exemplary embodiments of the present application;
[0012] Figure 2 A flowchart of a video batch generation method provided for the exemplary embodiments of the present application;
[0013] Figure 3 An interaction diagram of another video batch generation method provided for the exemplary embodiments of the present application;
[0014] Figure 4 A structural diagram of an electronic device provided for the exemplary embodiments of the present application. DETAILED DESCRIPTION
[0015] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in conjunction with the specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0016] It should be noted that in the case of the user information involved in the embodiments of the present application, the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region, and provide corresponding operation portal for the user to choose authorization or refusal. In addition, the various models (including but not limited to language models or large models) involved in the present application are in line with the relevant legal and standard regulations.
[0017] To solve the problem of monotonous and serious homogeneity of video style generated based on preset templates and poor flexibility of video generation in the prior art. In the embodiments of the present application, an initial video template is obtained, and the initial video template is parsed with the material type as a split variable to obtain information segments corresponding to a plurality of material types; further, rendering parameter information corresponding to each of the plurality of material types is extracted from the information segments corresponding to each of the plurality of material types respectively; the rendering parameter information corresponding to each of the plurality of material types is templated to obtain a plurality of materialization templates, and each materialization template is used to describe the rendering rule of a video material. The plurality of materialization templates are applied to video generation to improve the flexibility and content diversity of video generation, thereby realizing efficient batch generation of diversified style videos.
[0018] The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0019] Figure 1 A flowchart of a video template variable method provided by an exemplary embodiment of the present application is shown in FIG. Figure 1 As shown in the figure, the method comprises:
[0020] S11, obtaining an initial video template, the initial video template including rendering rules of a plurality of video materials required for generating a video and a hierarchical relationship between the plurality of video materials;
[0021] S12, parsing the initial video template with the material type as a split variable to obtain information segments corresponding to a plurality of material types;
[0022] S13, extracting rendering parameter information corresponding to each of the plurality of material types from the information segments corresponding to each of the plurality of material types respectively;
[0023] S14, templating the rendering parameter information corresponding to each of the plurality of material types to obtain a plurality of materialization templates, and each materialization template is used to describe the rendering rule of a video material.
[0024] In the embodiment of the present application, the initial video template can be an existing video production framework. The video production framework can be pre-set or exported after video editing by an editing software. Regardless of the manner, the initial video template includes rendering rules of various video materials required for generating a video and hierarchical relationships between the various video materials. In the embodiment of the present application, the specific manner of obtaining the initial video template is not limited. For example, the initial video template can be a pre-set common video template provided by default by the system, a custom video template created by a user according to requirements, or a video template imported from other external sources.
[0025] The various video materials required for generating a video refer to various multimedia elements constituting video content. The various video materials together can provide rich content for a video. In the embodiment of the present application, the specific type of video material is not limited. For example, the video material can be a video, an image, text, or audio. It is specified herein that the initial video template does not include video materials, but includes rendering rules and hierarchical relationships. The rendering rules correspond to the type of video material, and the hierarchical relationships describe the display order and occlusion logic between different video materials.
[0026] In the embodiment of the present application, the rendering rules of the various video materials can be specific parameters and operation modes followed when the various video materials are synthesized into a video, so as to achieve the presentation effect of the various video materials in the video. In the embodiment of the present application, the specific type of rendering rule is not limited. For example, the rendering rule can include special effect rules for describing, but not limited to, visual effects such as blur, shadow, and halo in a video. The rendering rule can also include color adjustment for describing, but not limited to, brightness, contrast, and saturation in a video. The rendering rule can also include filter effects for describing, but not limited to, retro, black and white, and color filter effects in a video. In addition, the rendering rule can also be animation effects such as fade-in and fade-out, scaling, and rotation. In the embodiment of the present application, the rendering effect represented by the specific parameters in the rendering rule is not limited.
[0027] The rendering rule of each video material type is independently defined. For example, the rendering rule of a video material that is a piece of background music can be "set the volume to 50%, start playing at the 5th second of the video and continue to the end"; and the rendering rule of a video material that is a piece of text content can be "white sans-serif font, displayed in the upper right corner of the screen with a 0.5 second fade-in effect".
[0028] In the embodiments of the present application, the hierarchical relationship between the multiple video materials describes the display order and occlusion logic of different material types in video synthesis. In the embodiments of the present application, the specific type of the hierarchical relationship between different material types is not limited. For example, the hierarchical relationship between different material types can be a picture-in-picture effect, in which one material type is located at a higher level and covers another material type at a lower level; it can also be a special effect superposition, in which multiple special effect layers are stacked together to form a complex visual effect; it can also be text content covering, in which text content is located at a higher level and covers a video or image at a lower level. The hierarchical relationship determines the front and back positions of different material types in the picture, affects the occlusion and display effect between different material types, and different visual effects can be achieved by adjusting the hierarchical relationship between different material types, making the video more consistent with the creative intent. For example, the material type of video material can be located at a lower level as a background layer to provide basic visual content for the entire video; the material type of image material can be located at an intermediate layer to cover the video material for highlighting specific information or decoration; and the material type of text material can be located at the highest level as a foreground layer to ensure clear and readable text information. This hierarchical relationship arrangement enables text to effectively convey information, while images and videos provide rich visual backgrounds. In addition, the hierarchical relationship between different material types can be flexibly adjusted according to creative needs. For example, in order to highlight a certain image material, its level can be raised to cover the text material. Or, in order to create a picture-in-picture effect, a video material can be placed at a higher level to be displayed above another video material. In this way, the creator can precisely control the visual hierarchy of each element in the video, thereby achieving more rich and diverse video effects.
[0029] Further, after the initial video template is acquired, the initial video template is parsed with the material type as a split variable. The material type refers to different categories of video materials, and is used to distinguish the classification criteria of different video materials. The split variable refers to an independent information segment obtained by dividing the initial video template according to the material type when the initial video template is parsed. For example, if the initial video template contains two kinds of video materials, namely, images and audios, the initial video template is split into independent information segments of images and audios as classification criteria, and subsequent processing is performed respectively.
[0030] The information segment refers to an information segment corresponding to the material type extracted from the initial video template. The information segment can contain rendering parameter information of the corresponding material type. For example, an information segment can contain rendering parameter information such as path information, playing time, and special effect application of the material type.
[0031] In the embodiments of the present application, the rendering parameter information is used to describe at least one attribute of a material file of a material type. Each attribute can be regarded as a rendering parameter, and the attribute value of each attribute in the initial video template can be regarded as the default parameter value of the corresponding rendering parameter. The rendering parameters with the default parameter values form a rendering parameter information corresponding to the material file of the material type. The "material file" refers to various resource files used in the materialized template, mainly including various video materials such as pictures, videos, and audios. The rendering parameter information can be used to control the rendering logic of the material file, so that the material file can obtain the rendering effect in the finally generated video. The rendering parameter information quantitatively describes the attributes of the material file, so that the attributes of the video material are converted into the attributes of the rendering parameter information. The rendering effect of the video material in the finally generated video is controlled by the rendering parameter information. In other words, the attributes of the video material are described as corresponding rendering parameters by the rendering parameter information, and the rendering parameters with the default parameter values define the attribute values of certain attributes of the material file.
[0032] The rendering parameter information can be obtained by extracting from the information segments corresponding to the plurality of material types respectively. The rendering parameter information corresponding to each material type contains at least one rendering parameter and the corresponding default parameter value. Each rendering parameter is used to describe the attributes of the material file of a material type. In the embodiments of the present application, the specific attributes corresponding to the rendering parameters of the plurality of material types are not limited.
[0033] In the embodiments of the present application, the templating can convert the rendering parameter information corresponding to each of the plurality of material types from the rendering parameter with the fixedly configured default parameter value to the variable parameter set with the dynamically replaceable parameter value. The templating of the rendering parameter information corresponding to each of the plurality of material types can obtain a plurality of materialized templates corresponding to each of the plurality of material types, and each materialized template is used to describe the rendering rule of a material type, and the rendering rule includes at least one variable parameter, and each variable parameter is associated with a plurality of candidate parameter values, so as to control the rendering effect of the generated video.
[0034] In the embodiments of the present application, the materialized template is a result of the templating of the rendering parameter information corresponding to each of the plurality of material types. It includes the path information placeholder of the material file and the variable parameter and the plurality of candidate parameter values associated therewith. The materialized template describes the rendering rule of the material file of the corresponding material type in the generated video, and the rendering rule is used to describe the rendering logic followed by the video material in the video rendering process, so as to control the expected rendering effect of the presentation of the material file in the generated video. The materialized template formed by converting the rendering parameter information from the rendering parameter with the fixedly configured default parameter value to the variable parameter set with the dynamically replaceable parameter value can assign different candidate parameter values to the variable parameter, so as to flexibly adjust the rendering effect of the material file in the generated video in different scenarios. The templating of the rendering parameter information forms the plurality of materialized templates which can be modularly recombined. The materialized templates for different material types can be recombined to obtain different target video templates, and then a plurality of diversified videos can be batch generated without designing and adjusting the video template for each video. This not only saves time and human resources, but also improves the flexibility, content diversity of the video template generation, and then realizes the efficient batch generation of diversified style videos.
[0035] For example, a plurality of materialized templates can be flexibly selected and combined. Different materialized templates can be selected according to specific needs, and different materialized templates can be combined together to generate more complex and diversified video content. For example, one materialized template can define the background effect of the video, and another materialized template can define the animation effect of the text. By combining the two materialized templates, a video with rich visual effects can be quickly generated.
[0036] For example, for the same materialized template, various properties and property values of the material type corresponding to the materialized template can also be flexibly set. For example, the transparency of the image material, the playback speed of the video material or the font size of the text material can be adjusted, so as to realize the fine control of the video content. The materialized template improves the flexibility of the video generation, and can significantly improve the efficiency and quality of the video generation, while meeting the diversified creation needs.
[0037] In the embodiments of the present application, the specific content of the candidate parameter value of the variable parameter associated with the templateized rendering parameter information of each of the plurality of material types is not limited. For example, for a materialization template of the video material type of video, the configured optional parameter values can include "playback speed" (optional values: 0.5 times speed, 1 times speed, 1.5 times speed), "filter type" (optional values: black and white, retro, bright), "transparency" (optional range: 0% to 100%), and the like; for a materialization template of the video material type of image, the configured optional parameter values can include "scaling ratio" (optional values: 50%, 100%, 200%), "rotation angle" (optional range: 0° to 360°), "layer position" (optional values: center, top left corner, bottom right corner), and the like.
[0038] By associating a plurality of candidate parameter values with the variable parameter, the variable parameter can be flexibly adjusted according to requirements, improving the flexibility and diversity of video template generation, and thus achieving efficient batch generation of diversified style videos.
[0039] In an optional embodiment, the initial video template can be implemented as an MLT (Media Lovin' Toolkit) template and stored in an XML (eXtensible Markup Language) structured data format, i.e., an XML document.
[0040] In an optional embodiment, the initial video template is parsed with the material type as the split variable to obtain information segments corresponding to the plurality of material types, including: loading an XML document corresponding to the initial video template, the XML document including a root element and a plurality of non-root elements connected to the root element, the plurality of non-root elements including a plurality of specific elements, each specific element being used to describe the rendering rule of a material type; starting from the root element, traversing the non-root elements in the XML document to identify the plurality of specific elements; and extracting the plurality of information segments in which the plurality of specific elements are located as the information segments corresponding to the plurality of material types. In the embodiments of the present application, the XML document of the initial video template can be loaded into the memory to form a tree structure that can be parsed. In the embodiments of the present application, the XML document can be loaded into the memory by a DOM (Document Object Model) parser to form a parsing mode of a tree structure.
[0041] The tree structure includes a root element and non-root elements. The root element is the top node of the XML document, and the non-root elements are directly or indirectly nested under the root element. An XML document has only one root element, which is the first element of the XML document and can be used as the starting point of the XML document. For example, <template> <metadata> 、 <images>) are sub-elements of the root element, and are non-root elements.
[0042] The non-root elements are sub-elements directly or indirectly nested under the root element, and are used to divide different information segments of the initial video template according to the material types. A plurality of specific elements are included in the non-root elements. Starting from the root element, the non-root elements in the XML document are traversed, and the plurality of specific elements can be identified. The specific elements are elements that can directly describe the rendering rules of a certain type of material, and each specific element corresponds to a material type. For example, if the material type is an image, then is a specific element, and the <path> 、 <position>The sub-nodes can define the path, position of the image.
[0043] For example, the XML document corresponding to the initial video template is as follows:
[0044]
[0045] The XML document corresponding to the initial video template is loaded, and the root element is identified <template>, self-rooted element <template>Start on non-root element in XML document <metadata>and <media>traverse, non-root element <media>The following identifies specific elements <video>and <audio>, extracting specific elements <video>Video material information piece where the user is located:
[0046]
[0047] Extracting specific elements <audio>Audio material information segments where the audio material is located:
[0048]
[0049] By extracting these information segments, the system can quickly identify the properties of the material file and render it according to its property values. This design of extracting information segments significantly improves the flexibility and diversity of video generation, meets the diversified creation needs, and provides a technical foundation for efficient batch video generation.
[0050] In an optional embodiment, starting from the root element, the non-root elements in the XML document are traversed to identify a plurality of specific elements, including: S1, starting from the root element, traversing the non-root elements in the XML document; S2, for the currently traversed non-root element, obtaining the element tag contained by the currently traversed non-root element; S3, if the element tag is a specific tag, determining whether the currently traversed non-root element contains a child element; S4, if the currently traversed non-root element contains a child element, taking the child element as the currently traversed non-root element and returning to step S2; S5, if the element tag is a non-specific tag, continuing with the next non-root element and returning to step S2; S6, if the currently traversed non-root element does not contain a child element, taking the currently traversed non-root element as a specific element.
[0051] In step S1, starting from the root element of the XML document, the non-root elements are accessed one by one for traversal. In the embodiments of the present application, the specific implementation strategy of traversal is not limited. For example, the traversal can be implemented using a depth-first search algorithm, or a breadth-first search algorithm. In the embodiments of the present application, the process of traversing the non-root elements in the XML document to identify a plurality of specific elements is described in detail taking the depth-first search algorithm as an example.
[0052] Next, step S2 is performed, for the currently traversed non-root element, obtaining the element tag contained by the currently traversed non-root element. In the XML document, the element tag is an identifier within angle brackets (<>) and is used to mark the type and semantic meaning of the element. Different material types are distinguished by the name of the element tag, such as representing image material; the specific position of the material is located by the hierarchical relationship of the element tag, such as <images>under the container element.
[0053] After obtaining the element tag contained in the non-root element currently traversed, it is judged whether the element tag is a specific tag. The specific tag can be a set of key tags defined in advance, representing an element tag that needs special processing in the XML document, and is used to identify the material type that needs to be extracted. For example, <text>, and <video>The specific tags can be specific tags because they describe rendering rules for different material types. The non-specific tags are tags that do not belong to a material type. For example, the non-specific tags can be, for example, <images>or <texts>A container tag for an organizational structure, can also be as follows <metadata>or <author>The metadata tag for describing the overall information of the initial video template can also be, for example, <settings>or <global-effects>An auxiliary tag defining a global parameter.
[0054] Then, step S3 or S5 is executed. If step S5 is executed, i.e. the element tag is a non-specific tag, the next non-root element is continued, and the process returns to step S2.
[0055] If step S3 is executed, i.e. the element tag is a specific tag, it is continued to judge whether the current traversed non-root element contains a sub-element. The sub-element can be other elements nested inside the current traversed non-root element. For example, An element under the element for defining a file path <path>Element.
[0056] Then, step S4 or S6 is executed. If step S4 is executed, i.e. the current traversed non-root element contains sub-elements, the sub-elements are taken as the current traversed non-root element and the execution returns to step S2. If step S6 is executed, i.e. the current traversed non-root element does not contain sub-elements, the current traversed non-root element is taken as a specific element. Step S4 is a recursive process for the case of containing sub-elements, the sub-elements of the current traversed non-root element are taken as the new starting point of the traversal and steps S2-S6 are re-executed for each sub-element. After the recursion is finished, the other non-root elements are traversed. Step S6 is a terminal collection for the case of no sub-elements, the current traversed non-root element is taken as a specific element.
[0057] For example, the XML document is:
[0058]
[0059] The traversal process is: first, step S1 is executed from the root element <videotemplate>Initially, the non-root elements are visited one by one for traversal; the step S2 is performed for the non-root element currently being traversed <textelement>, get non-root element <textelement>The element label TextElement is included; it is judged whether the element label TextElement is a specific label, and if so, the step S3 of judging a non-root element is continued <textelement>whether to include sub-elements; determine non-root elements <textelement>contains sub-elements <position>and <fontsize>If yes, then step S4 is performed, the sub-element <position>Return to perform step S2 for the currently iterated non-root element <position>, get non-root element <position>The element label Position is included; it is determined that the element label Position is a non-specific label, and then step S5 is executed, and the next non-root element is continued <fontsize>, return to perform step S2; performing step S2 for the non-root element currently being iterated over <fontsize>, get non-root element <fontsize>The element tag FontSize is included. It is determined that the element tag FontSize is a non-specific tag. When the non-root element <textelement>all sub-elements of the element are processed, the non-root element <textelement>The included element tag TextElement itself is a specific tag and the child elements have been traversed, and finally the step S6 executes the non-root element <textelement>Marking specific elements.
[0060] In an optional embodiment, the rendering parameter information corresponding to each of the plurality of material types is extracted from the information segment corresponding to each of the plurality of material types, including: for each information segment, extracting path information of a material file and at least one attribute value of the material file from the information segment, the attribute value being used for rendering the material file; reading the material file according to the path information, and determining the material type described by the information segment according to the extension of the material file; taking at least one attribute to which the at least one attribute value belongs as at least one rendering parameter, and taking the at least one attribute value as a default parameter value of the at least one rendering parameter, to obtain the rendering parameter information corresponding to the material type described by the information segment.
[0061] The information segment is a part related to the material type extracted from the XML document, and contains path information of a material file and at least one attribute value of the material file. The material file can be used to construct various media files of the finally generated video. In the embodiments of the present application, the material file can include but is not limited to: a file in a format such as MP4 or AVI, containing dynamic images and audio, used to show a series of continuous video files; a file in a format such as WAV or MP3, providing audio files of background music, narration or special effect sound; a picture in a format such as JPEG or PNG, which can be used as a background picture, an icon or a visual element in a specific scene; a text file that does not directly serve as display content, but can contain subtitle information or other text content that needs to be superimposed on the video. The path information of the material file is a string of characters representing the storage location of the material file, which can include the file name and the extension, and the material file can be located and read according to the path information. For example, / videos / intro.mp4 is a path information pointing to the intro.mp4 video file stored in the videos directory.
[0062] In an optional embodiment, the extension of the material file is an identifier at the end of the path information of the material file, used to determine the material type described by the information segment, and different extensions represent different material types. For example,.mp4,.mov,.avi represent the material type as video;.mp3,.wav represent the material type as audio;.png,.jpg represent the material type as image;.srt represents the material type as text.
[0063] The at least one attribute value of the material file extracted from the information segment is a specific value corresponding to the attribute, to describe the specific state of the attribute. The attribute can describe the rendering effect of the material file. For example, for the material file being a video file, the attribute can be video duration, video speed, filter effect, etc., and if the attribute is video duration, the attribute value can be 2 minutes, 1 hour, 1 day, etc.
[0064] In an optional embodiment, the rendering parameter information corresponding to the material type described by the information segment includes at least one rendering parameter and a default parameter value of the at least one rendering parameter. The default parameter value can be an attribute value of an attribute corresponding to the rendering parameter before quantization. Wherein, at least one attribute of at least one attribute value of the extracted material file in the information segment can be taken as at least one rendering parameter, and the at least one attribute value can be taken as a default parameter value of the at least one rendering parameter to obtain the rendering parameter information. That is, a key-value pair set composed of the rendering parameter and the default parameter value is used to control the rendering logic of the material file in the video to achieve the expected rendering effect.
[0065] For example, the XML document of an information segment is as follows:
[0066]
[0067] Wherein, the path information of the material file is / images / logo.png, the material file is read according to the path information to obtain the extension of the material file as.png; it is determined that the material type described by the information segment is an image according to the extension.png; the attributes of the material file are width, height and format, and the attribute values are 100 (width), 100 (height) and PNG (format). The attributes width, height and format are taken as rendering parameters, and the attribute values 100 (width), 100 (height) and PNG (format) are taken as default parameter values of the rendering parameters to obtain the rendering parameter information corresponding to the material type described by the information segment as "rendering parameter width, default value 100; rendering parameter height, default value 100; rendering parameter format, default value PNG". Through the above steps, the corresponding rendering parameter information of each material type can be generated, and the rendering parameter information can be used in the subsequent process of generating and rendering the video to ensure that the material can be correctly displayed according to the predetermined rendering rule.
[0068] In an optional embodiment, the rendering parameter information corresponding to each material type is templated to obtain a plurality of materialization templates, including: for each material type, selecting at least one variable parameter from the rendering parameter information corresponding to the material type; associating a plurality of candidate parameter values with the at least one variable parameter; and generating a materialization template corresponding to the material type according to the plurality of candidate parameter values associated with the at least one variable parameter. The variable parameter is obtained by variable processing on the rendering parameter with a default parameter value, and each variable parameter is associated with a plurality of candidate parameter values. Different candidate parameter values correspond to different rendering logics to produce different rendering effects in the generated video, which can be used to generate diversified materialization templates. For each material type, at least one variable parameter is selected from the rendering parameter information corresponding to the material type. These variable parameters can be adjusted within a range to obtain different materialization templates. For example, for video material files, the playback speed and filter effects can be selected as variable parameters. The variable parameter is associated with a plurality of candidate values to limit the adjustment range of the variable parameter, ensuring that the material file of the generated video meets the design requirements and avoids invalid settings. For example, for the playback speed variable parameter, the candidate parameter values can include 0.5x, 1x, and 1.5x; and for the filter effect, the candidate parameter values can include grayscale and retro.
[0069] In an optional embodiment, when selecting at least one variable parameter from the rendering parameter information corresponding to the material type, two selection methods are provided: one is to select all rendering parameters as variable parameters, and the other is to select part of the rendering parameters as variable parameters according to the weight values. When all rendering parameters are selected as variable parameters, there is no need to compare the weight values, and all rendering parameters can be adjusted as variable parameters.
[0070] In an optional embodiment, the method of selecting part of the rendering parameters as variable parameters according to the weight values can include: pre-configuring the weight values of each rendering parameter for the material type; and parsing each rendering parameter from the rendering parameter information corresponding to the material type, and selecting at least one rendering parameter with a weight value greater than a set weight threshold as at least one variable parameter. The weight value can be used as an importance score assigned to each rendering parameter to quantify the influence of the rendering parameter on the rendering effect of the material file in video generation. At least one rendering parameter with a greater influence on user perception or application target can be selected as at least one variable parameter. The weight threshold is a pre-defined critical value. The user can select the parameter with a weight value higher than the weight threshold as a variable parameter to avoid selecting too many irrelevant rendering parameters as variable parameters to prevent configuration conflicts or rendering logic confusion.
[0071] In the embodiments of the present application, the manner of pre-configuring the weight values of the various rendering parameters is not limited. For example, the weight values of the various rendering parameters can be empirically labeled or automatically generated through user behavior data analysis.
[0072] For example, the weight value configuration of the rendering parameters of the material type of image can be:
[0073] rendering parameters weight values default parameter values brightness 0.7 100% contrast 0.5 90% saturation 0.9 80%
[0074] According to the weight threshold setting of 0.6 to filter the rendering parameters, according to 0.7>0.6, 0.9>0.6, and 0.5<0.6, the variable parameters are selected as brightness and saturation, and the rendering parameter contrast is excluded.
[0075] Optionally, the selection of at least one variable parameter from the rendering parameter information corresponding to the material type can also be randomly selected according to the set number of variable parameters. In an optional embodiment, according to the plurality of candidate parameter values associated with the at least one variable parameter, a materialized template corresponding to the material type is generated, including: adding the various rendering parameters corresponding to the material type and the default parameter values of the various rendering parameters to a preset template file, and adding the plurality of candidate parameter values associated with the at least one variable parameter in the preset template file; and adding a placeholder for carrying a material file corresponding to the material type in the preset template file to obtain a materialized template corresponding to the material type. Wherein, the preset template file is a basic video template containing basic configuration, containing the basic structure of the video template and some preset rendering parameters. Wherein, adding the various rendering parameters corresponding to the material type and the default parameter values of the various rendering parameters to the preset template file can ensure that the finally obtained materialized template has complete rendering parameter information.
[0076] In the embodiments of the present application, the manner of pre-configuring the weight values of the various rendering parameters is not limited. For example, the weight values of the various rendering parameters can be empirically labeled or automatically generated through user behavior data analysis.
[0077] In an optional embodiment, based on adding the plurality of candidate parameter values associated with the at least one variable parameter in the preset template file, a placeholder for carrying the material file corresponding to the material type can be added in the preset template file to obtain the materialized template corresponding to the material type. The placeholder is a reserved position for filling the material file, supports dynamic replacement, and can fill the path information of the actual material file corresponding to the material type, for example. Different material files can be flexibly replaced without modifying the template structure through the placeholder.
[0078] Embodiments of scenarios:
[0079] Taking the video as an example, the rendering parameter information of the material type includes the rendering parameters and the default parameter values as follows:
[0080]
[0081] The preset weight values of the respective rendering parameters are as follows: duration (video duration): 0.8, language: 0.6, resolution: 0.4; the weight threshold is set to 0.5, and duration and language are selected as the variable parameters.
[0082] The plurality of candidate parameter values associated with the at least one variable parameter are as follows: duration: 1 minute, 2 minutes, 3 minutes, language: Chinese, English, Japanese.
[0083] By adding the rendering parameters and the default parameter values, the plurality of candidate parameter values associated with the at least one variable parameter, and the placeholder [VIDEO_PATH] for carrying the actual path of the video file and the default value of / videos / intro.mp4 in the preset template file, the materialized template is obtained as follows:
[0084]
[0085] In the above embodiment, the materialized template ensures the completeness of the rendering parameter information of the generated materialized template by integrating the rendering parameters of the material type and their default parameter values into the preset base template, and realizes dynamic selection of the parameter values of the variable parameters by associating multiple candidate parameter values for the variable parameters, so as to correspond to the flexible adjustment of the video style. The embedding of the placeholder further decouples the parameter configuration of the initial video template and the material file, supports dynamic replacement of the material file without modifying the template structure, so that the corresponding candidate parameter values can be assigned to the variable parameters according to the application requirements to obtain the freely combined materialized template, and finally the flexibility and diversity of the video template generation are significantly improved, thereby efficiently mass producing video contents with different styles, and solving the problems of complicated parameter configuration, complex adaptation process, and serious homogenization of generated video contents in traditional video production.
[0086] On the basis of the above plurality of materialized templates, batch video generation or single video generation can be performed based on the plurality of materialized templates. For the batch video generation scenario, the plurality of materialized templates provided by the embodiments of the present application can generate a plurality of videos with diverse styles and differentiated contents. The following describes an embodiment of batch video generation based on the plurality of materialized templates provided by the embodiments of the present application.
[0087] Figure 2 A flowchart of a video batch generation method provided in the embodiments of the present application is shown in FIG. 6. As shown in FIG. 6, the method comprises the following steps. Figure 2
[0088] S21, in response to input operations on the video generation page for the number of videos and the video category, generating a batch video generation task, the batch video generation task comprising the number of videos N and the video category, N being an integer greater than or equal to 2;
[0089] S22, generating N video instance identifiers according to the batch video generation task, and generating a set of video materials related to the video category for each video instance identifier;
[0090] S23, for each video instance identifier, determining at least one target materialized template from a plurality of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier;
[0091] S24, combining the at least one target materialized template based on the hierarchical relationship between the plurality of video materials included in the initial video template to obtain a target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rules and hierarchical relationship of the set of video materials;
[0092] S25, generating videos according to the N video instance identifiers, the respective target video template and the set of video materials, to obtain N videos under the video category.
[0093] In this embodiment, the execution subject of the above-mentioned video batch generation method is not limited. For example, the method can be implemented as a service product, which can adopt a client-server architecture. On the one hand, the video generation page for batch video generation is provided to the user on the client side, to receive the user's input operation for the video quantity and the video category, and then initiate a batch video generation task to the server. On the other hand, the server responds to the batch video generation task initiated by the client through the video generation service page, and generates videos in batch through the computing resources, network bandwidth and storage resources of the server, etc., which is conducive to improving the speed of video generation.
[0094] For another example, as the processing capability of the corresponding hardware device of the client side is enhanced, the above-mentioned method can also be executed by the client side. The client side can provide a video generation page to the user, and generate a batch video generation task in response to the user's input operation for the video quantity and the video category on the video generation page, and then generate videos in batch through the computing resources and storage resources of the client side, etc. Wherein, in the case that the batch generation task is deployed to be executed on the client side, there is no need to transmit data to the server, which can save network delay.
[0095] In this embodiment, the batch video generation task includes the video quantity N and the video category, and N is an integer ≥ 2. The video category is used to describe the expression theme of the video content of the batch video generation, and the specific implementation of the video category is not limited. For example, it includes but is not limited to product introduction, teaching explanation, knowledge popularization and beauty and skin care, etc. Optionally, the video category can be implemented as a single-level video category, for example, it can be implemented as business registration or legal consultation, etc. Optionally, the video category can also be implemented as a multi-level video category. For example, the first-level video category can be business registration; the second-level video category of the first-level video category can be cleaning or food operation, etc.
[0096] In this embodiment, according to the batch generation task, N video instance identifiers are generated, and the N video instance identifiers are all different. Each video instance identifier can be used to uniquely represent a video to be generated, so as to track the required video materials and video templates of the video to be generated, that is, the video instance identifier can also be used as the unique identity of the related content (such as video materials) of the video to be generated.
[0097] Wherein, the way of generating the N video instance identifiers is not limited. For example, including but not limited to numbers and strings, etc. For example, it can be an increasing sequence starting from an arbitrary integer and formed by a fixed step to obtain N integers as N video instance identifiers; or it can also be a preset N strings, etc.
[0098] In the embodiment, a set of video materials related to the video category is generated for each video instance identifier, and the set of video materials is used to generate a video corresponding to the video instance identifier. Wherein, each set of video materials contains video materials of at least one material type. In some embodiments of the application, the video materials of one material type are referred to as one video material.
[0099] Wherein, the material type refers to the type of media resources that constitute a video, including but not limited to: audio, video, background picture, subtitle, digital person, etc.
[0100] In the embodiment, the implementation of generating a set of video materials related to the video category for each video instance identifier is not limited.
[0101] In an optional implementation, for any video instance identifier, video materials can be randomly extracted from the multiple material types stored in the basic material library, and at least one video material is extracted as a set of video materials for the video instance identifier. Wherein, the basic material library stores multiple video materials under multiple video categories.
[0102] In another optional implementation, for any video instance identifier, the semantic similarity between at least one video material in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the video material that meets the similarity condition in the multiple similarity information is taken as a set of video materials for the video instance identifier.
[0103] Further, in the embodiment, multiple materialized templates obtained by variable processing of the initial video template are obtained. Optionally, the initial video template can be implemented as an MLT (Media Lovin'Toolkit) template, which is an open source framework for multimedia processing and is based on an XML format file that records all video editing parameters such as video clips on the timeline, audio tracks, filter effects, transitions, etc. and can be used for video editing.
[0104] In the embodiment of the present application, the initial video template can be an existing video production framework. The video production framework can be pre-set or exported after video editing by an editing software. In the embodiment of the present application, the specific manner of obtaining the initial video template is not limited. For example, the initial video template can be a pre-set general video template provided by the system by default, a custom video template created by the user according to the demand, or a video template imported from other external sources.
[0105] The initial video template includes rendering rules of various video materials required for generating a video, and a hierarchical relationship between the various video materials. In the embodiment of the present application, the rendering rule of each video material is used to describe the rendering logic followed by the video material in the video rendering process, so as to achieve the expected rendering effect of the video material in the generated video. The hierarchical relationship between the various video materials formed in the initial video template refers to the layer superposition order of the various video materials in the video generation process in the initial video template, which is used to determine the front-back coverage relationship and occlusion logic of the various video materials in the video generation. For example, in generating a video, the subtitle can be located in the relatively upper layer, the background picture can be located in the relatively bottom layer, the digital person can be superimposed on the upper layer of the background picture, and the layer of the subtitle.
[0106] In the embodiment, the variable processing of the initial video template is partly reflected in that the initial video template is structurally split and the rendering parameter information is templated by taking the material type as a split variable, so as to obtain a plurality of materialized templates. The materialized template is corresponding to the material type. For example, the material type includes audio, video, background picture, subtitle, digital person, and the like. The corresponding materialized template types include, but are not limited to, audio materialized template, video materialized template, background picture materialized template, subtitle materialized template, digital person materialized template, and the like. In other words, each material type corresponds to a materialized template, and each materialized template is used to describe the rendering rule of the video material.
[0107] In the embodiment, the timing of the variable processing is not limited. For example, the initial video template can be pre-processed. For another example, the initial video template can be dynamically processed. The details of how to perform the variable processing can be referred to the subsequent embodiments.
[0108] In the embodiment, the materialized template subjected to the variable processing can be modularly recombined. For each video instance identifier, at least one target materialized template is determined from the plurality of materialized templates according to the material type in the group of video materials corresponding to the video instance identifier. The material type in the group of video materials is corresponding to the materialized template.
[0109] For example, if the set of video materials includes audio, video, background images, and subtitles, then the target material template can include the target material templates corresponding to each of the audio, video, background images, and subtitles. Similarly, if the set of video materials includes audio, video, background images, subtitles, and digital human materials, then the target material template can include the target material templates corresponding to each of the audio, video, background images, subtitles, and digital human materials.
[0110] Furthermore, given at least one target material template corresponding to a set of video materials, based on the hierarchical relationship between the various video materials included in the initial video template, the at least one target material template is combined to obtain the target video template corresponding to the video instance identifier. The target video template describes the rendering rules and hierarchical relationship of a set of video materials.
[0111] The initial video template includes a hierarchical relationship among various video materials. This hierarchy represents the layer stacking order of multiple material templates and can be used to organize and integrate target material templates, thereby forming a target video template with a clear hierarchy and corresponding rendering rules for the video materials. The target video template describes the hierarchical relationship of a group of video materials, which is consistent with the hierarchical relationship of that group of video materials in the initial video template.
[0112] In this embodiment, each of the N video instance identifiers corresponds to its own target video template. That is, each group of video materials has its own corresponding target video template, enriching the variety of target video templates used in batch video generation. Based on the rendering rules described by the N target video templates, video generation processing is performed on the grouped video materials corresponding to the N video instance identifiers, ensuring that the style of each batch-generated video matches the adopted video template, thereby increasing the diversity of the batch-generated video content.
[0113] Given the target video templates corresponding to each video instance identifier, video generation processing is performed based on the target video templates corresponding to each of the N video instance identifiers and a set of video materials to obtain N videos under the video category. Video generation processing refers to the process of filling and rendering the target video templates with video materials based on the target video templates corresponding to each video instance identifier and a set of video materials to obtain the video corresponding to that video instance identifier. For example, for N video instance identifiers, the set of video materials corresponding to each of the N video instance identifiers can be filled into N target video templates to obtain filled target video templates. Then, the filled target video templates can be rendered to obtain the videos under the video category.
[0114] In an optional embodiment, the corresponding video generation of each video instance identifier can be processed in batches, and M videos are processed in parallel in each batch, M is less than N, and M is an integer; after the generation of M videos is completed, the generation of M videos after the video generation processing is continued until the corresponding videos of the N video instance identifiers are all processed. Through batch video generation processing, the resource utilization rate of the server is improved, and the server overload caused by high concurrency is avoided.
[0115] In the batch video generation in the embodiments of the present application, the initial video template is subjected to variable processing to obtain a plurality of materialized templates of various material types, the materialized templates are reorganized to obtain the video template required for generating a video, and then a video instance identifier corresponding to each video generation is obtained, the video instance identifier is bound to a corresponding set of video materials, the target materialized template for reorganization is determined from the plurality of materialized templates according to the set of video materials corresponding to the video instance identifier, the target materialized template is organized and integrated in combination with the hierarchical relationship between the materialized templates provided by the initial template, and the video template corresponding to each video instance identifier is obtained for the video generation processing of the set of video materials corresponding to each video instance identifier; since the video template can be personalized and reorganized in combination with the video materials for generating a video, the flexibility of video generation is improved; and the video template obtained by reorganization has a corresponding relationship with the video instance identifier, which to some extent enriches the types of video templates, so that the style of each video generated in batch matches the video template used, and the diversification degree of the video content generated in batch is improved.
[0116] In the embodiments of the present application, the variable processing method has been described in detail in the foregoing embodiments, and will not be described again here.
[0117] In the embodiments of the present application, the initial video template is subjected to variable processing to obtain a plurality of materialized templates. In batch video generation, at least one target materialized template is determined from the plurality of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier.
[0118] In an optional embodiment, when at least one target materialized template is determined from the plurality of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier, it includes: identifying at least one material type contained in the set of video materials corresponding to the video instance identifier; selecting at least one initial materialized template from the plurality of materialized templates according to the at least one material type contained in the set of video materials, each material type corresponding to one initial materialized template; and performing parameter adjustment on at least part of the selected at least one initial materialized template to obtain at least one target materialized template.
[0119] In the embodiment, the materialized template included in the plurality of materialized templates obtained by the variable processing of the initial video template is referred to as an initial materialized template. In the case where at least one initial materialized template corresponding to a certain group is determined from the plurality of initial materialized templates, the rendering rule described by the initial materialized template can be referred to as an initial rendering rule. Further, parameter adjustment is performed on at least part of the at least one initial materialized template to obtain at least one target materialized template. Each initial materialized template is adjusted by parameters to obtain a corresponding target materialized template. The target materialized template obtained by the parameter adjustment includes a target rendering rule, which is different from the initial rendering rule described by the initial materialized template.
[0120] In the embodiment, since the initial materialized template is obtained by variable processing, the variable processing can convert the fixed rendering parameters in the initial video template into variable parameters that can be dynamically assigned, so as to realize flexible configuration of the template content. The values of the optional parameters are different, and the rendering rules are different. Optionally, each initial materialized template includes at least one variable parameter, and each variable parameter is associated with a plurality of candidate parameter values and a default parameter value.
[0121] In the embodiment, the parameter adjustment of the initial materialized template is performed to assign different candidate parameter values to the variable parameters of the initial materialized template, so as to obtain a plurality of target materialized templates with different candidate parameter values. The target materialized template of a certain material type can be used for video material rendering of different groups of video instance identifiers. The parameter values of the variable parameters are different, the rendering rules described by the target materialized templates of the same material type are different, the rendering results of the batch-generated videos are different, and the difference between different videos is formed, so as to improve the richness of the video content. How to perform parameter adjustment on the initial materialized template will be introduced below.
[0122] In an optional embodiment, when at least part of the at least one initial materialized template is adjusted by parameters to obtain at least one target materialized template, the method includes the following steps. The number of templates to be adjusted is determined, and the number of templates is less than or equal to the number of the at least one initial materialized template. The initial materialized template to be adjusted is selected from the at least one initial materialized template according to the number of templates to be adjusted. The variable parameter to be adjusted is determined from the initial materialized template to be adjusted. The target parameter value is randomly determined from the plurality of candidate parameter values associated with the variable parameter to be adjusted. The target parameter value is assigned to the variable parameter to be adjusted to obtain the target materialized template.
[0123] In the present embodiment, the manner of determining the number of templates to be adjusted is not limited. For example, the number of templates to be adjusted can be the number of at least one initial materialized template, i.e., the number of initial materialized templates is adjusted for each initial video template. For another example, a random integer within a preset range can be generated as the number of templates to be adjusted according to a random number generation algorithm, and the preset range refers to a number less than or equal to the number of at least one initial materialized template. The random number generation algorithm is not limited, including but not limited to: linear congruential generator (LCG) and Mersenne Twister.
[0124] Further, the initial materialized template to be adjusted is selected from the at least one initial materialized template according to the number of templates to be adjusted. In an optional embodiment, the selection of the initial materialized template to be adjusted is based on the priority of the material type and the number of templates to be adjusted. The priority of the material type refers to the importance of the material type to the presentation effect of the generated video. For example, the importance of the subtitle material type to the presentation effect is generally low, so the priority of the subtitle material type can be set to a low priority; relatively, the priority of the background picture can be higher than that of the subtitle, so it can be set to a medium priority; and the importance of the digital person to the presentation effect is high, so it can be set to a high priority. Further, the initial material template with a higher priority can be selected as the template to be adjusted according to the priority, and if the number of these initial material templates is less than the number of templates to be adjusted determined before, the initial material template with a lower priority can be selected, and the number of templates to be adjusted is less than or equal to the number of at least one initial materialized template.
[0125] Further, the variable parameter to be adjusted is determined from the initial materialized template to be adjusted. In an optional embodiment, all variable parameters in the initial materialized template to be adjusted can be selected as the variable parameter to be adjusted. In another optional embodiment, the target variable parameter in the initial materialized template to be adjusted is selected as the optional parameter to be adjusted, and the target variable parameter is a pre-selected optional parameter.
[0126] Further, the target parameter value is randomly determined from the plurality of candidate parameter values associated with the variable parameter to be adjusted. In some embodiments, the N video instance identifiers each correspond to at least one initial materialized template having the same variable parameter to be adjusted, and the plurality of candidate parameter values associated with the variable parameter to be adjusted can be randomly selected as the target parameter value of each of the N video instance identifiers, so that the optional parameters of the N video instance identifiers take values as different as possible.
[0127] Further, the target parameter value is assigned to the variable parameter to be adjusted to obtain the target materialization template.
[0128] In the case of obtaining the target materialization template, the target materialization template is combined to obtain the target video template. The embodiments are not limited in terms of the combination manner. Two combination manners are provided below, but are not limited thereto.
[0129] In an optional embodiment, according to the hierarchical relationship between the plurality of video materials included in the initial video template, a base video template is generated, which is used as a framework of the target video template and includes a plurality of blank structure positions corresponding to the plurality of video materials. The blank structure position refers to a placeholder of the target materialization template preset in the base video template, which is used to identify the position where the target materialization template can be inserted, and each blank structure position corresponds to the filling of the target materialization template of one material type. The positional relationship between the plurality of blank structure positions reflects the hierarchical relationship between the plurality of video materials. As described in the above embodiments, the hierarchical relationship represents the front-to-back superimposition order of the plurality of materials in the generated video. In some embodiments, the hierarchical relationship is extracted from the initial video template; or, it can also be preset based on at least one target materialization template, that is, the hierarchical relationship can be set on demand. For example, the structure position corresponding to the subtitle is located in the upper layer, the structure position corresponding to the digital person is located in the middle layer, and the structure position corresponding to the background picture is located in the bottom layer. Further, at least one target materialization template is inserted into the corresponding blank structure position of the base video template to obtain the target video template corresponding to the video instance identifier.
[0130] In another optional embodiment, according to at least one target materialization template, the structure position where the rendering rule of the video material of the same material type in the initial video template is located is covered to obtain the target video template corresponding to the video instance identifier. The difference between the structure position and the above blank structure position is that the rendering rule of each video material in the initial video template occupies a structure position, and the blank structure position is empty. The positional relationship between the structure positions reflects the hierarchical relationship between the plurality of video materials.
[0131] In the case of obtaining the target video template, the corresponding target video template and a set of video materials are identified for each video instance to perform video generation processing. In an optional embodiment, in the case of performing video generation processing on the corresponding target video template and a set of video materials of each of the N video instance identifications to obtain N videos under the video category, it includes: for each video instance identification, the corresponding set of video materials of the video instance identification is respectively filled into the target materialized template in the target video template corresponding to the video instance identification; for the filled target video template, according to the rendering rules and hierarchical relationships of the set of video materials described in the filled target video template, the set of video materials is rendered to obtain a video under the video category.
[0132] In the target materialized template, at least one placeholder corresponding to the material type is included, which is used to fill the video material of the material type.
[0133] In an optional embodiment, a set of video materials related to the video category is generated for each video instance identification, which includes: obtaining a set of video material description information related to the video category for each video instance identification, and each set of video material description information includes description information of multiple video materials; for each video instance identification, a plurality of material generation models based on artificial intelligence are called according to the corresponding set of video material description information of the video instance identification to generate multiple video materials corresponding to the video instance identification, and the multiple video materials corresponding to the video instance identification are uploaded to the content distribution network; accordingly, before performing video generation processing on the corresponding target video template and a set of video materials of each of the N video instance identifications to obtain N videos under the video category, it also includes: in response to a batch video generation trigger event, a plurality of sets of video materials corresponding to the N video instance identifications are obtained from the content distribution network. By hosting video materials on the content distribution network, the server-side storage pressure is reduced, so that the server can efficiently render a large number of videos and improve user experience.
[0134] In this embodiment, the description information of each set of video materials is used to generate the video materials of the video instance identification corresponding thereto. Each set of video materials contains video materials of multiple material types. Material type refers to the type of different video elements that make up the video content, including but not limited to: audio, video, background picture, subtitle, digital person, etc.
[0135] In this embodiment, the implementation of the description information of the set of video materials generated for each video instance identification related to the video category is not limited.
[0136] In an optional embodiment, for any video instance identifier, the description information of the video material can be randomly extracted from the plurality of material types stored in the base material library, and the extracted description information of the plurality of video materials is taken as the description information of the set of video materials of the video instance identifier. The base material library stores the description information of the plurality of video materials under the plurality of video categories.
[0137] In another optional embodiment, for any video instance identifier, the semantic similarity between the description information of the plurality of video materials in the base material library and the video category is calculated respectively to obtain a plurality of similarity information; the description information of the video material satisfying the similarity condition in the plurality of similarity information is taken as the description information of the set of video materials of the video instance identifier.
[0138] In yet another optional embodiment, for any video instance identifier, a material description information generation model is invoked according to the video category, and the model is used to generate the description information of the plurality of video materials related to the video category. The material description information generation model is obtained by training a large number of different sample video categories and sample video material description information of different material types, and by learning the semantic correlation between different sample video categories and sample video material description information of different material types, the model can be combined with different video categories to generate the description information of the plurality of video materials related thereto.
[0139] Further, for each video instance identifier, a plurality of material generation models based on artificial intelligence are invoked according to the video material description information corresponding to the video instance identifier, to generate a plurality of video materials corresponding to the video instance identifier, and the plurality of video materials corresponding to the video instance identifier are synchronously uploaded to the content distribution network.
[0140] Among them, one material generation model can generate at least part of the plurality of video materials. For example, one material generation model can generate one video material. The following is also described by way of example, but is not limited thereto.
[0141] In this embodiment, in the case of generating video materials of video instance identifiers, the target video template corresponding to each video instance identifier is determined according to the plurality of video materials corresponding to each video instance identifier, and the target video template corresponding to each video instance identifier is used to describe the rendering rule and hierarchical relationship of the plurality of video materials corresponding to the video instance identifier. The determination method of the target video template corresponding to each video instance identifier can refer to the above embodiments, which will not be described here.
[0142] In this embodiment, in response to the batch video generation trigger event, the N video instance identifiers correspond to multiple video materials are respectively obtained from the content distribution network; and the N video instance identifiers correspond to multiple video materials and the target video template are used for video generation to obtain N videos under the video category.
[0143] In this embodiment, the specific implementation of the batch video generation trigger event is not limited, and can be flexibly configured according to actual application requirements. For example, it can be in the case that the N video instance identifiers correspond to multiple materials are all generated; or, it can be in the case that each video instance identifier corresponds to multiple materials is generated, in which case, the video generation is performed for the multiple video materials corresponding to the video instance identifier; or, the batch video generation trigger event can be a preset trigger time, such as after a period of time after the input operation on the video generation page for the video quantity and the video category, for example, 2 hours, 1 day or 1 week, etc., the time span is not limited.
[0144] It should be noted that each video instance identifier corresponds to the generation of a video, and N video instance identifiers can correspond to the generation of N videos. When the N video instance identifiers correspond to multiple video materials are generated, the generation process of each video instance identifier corresponds to the video is asynchronous, and the generation of each video instance identifier corresponds to the video does not affect each other, so as to improve the generation efficiency of the N videos.
[0145] Further optionally, the description information of each group of video materials includes but is not limited to: subtitle description information, audio type description information, digital person description information and background picture description information. As shown in Figure 3 According to the video instance identifier corresponds to a group of video material description information, a plurality of material generation models based on artificial intelligence are called to generate a plurality of video materials corresponding to the video instance identifier, including: according to the subtitle description information, a generative language model is called to generate text information to obtain a target subtitle; according to the target subtitle and the audio type description information, a text-to-speech model is called to convert the target subtitle into a target audio that is adapted to the audio type description information; according to the target audio and the digital person description information, a multi-modal model is called to select a target digital person according to the digital person description information, and to generate a green screen video based on the target audio and the target digital person, to obtain a green screen video of the target digital person; according to the background picture description information, a text-to-image model is called to generate a background picture to obtain a target background picture.
[0146] In this embodiment, the APIs (Application Programming Interfaces) of multiple material generation models are associated with endpoints. In this embodiment, the video generation page is the presentation of the endpoint's front-end code. The endpoint is used for generating video materials, managing batch video generation tasks, and controlling the video generation process.
[0147] The endpoint includes front-end code and back-end code, adopting a client-server structure as described in the above embodiment. The front-end code refers to the video generation page built on the front-end framework. This page runs on the client side, interacts with the user, receives user input, and initiates batch video generation tasks to the server. The back-end code runs on the server side and is derived from the API of existing video editing software, such as Shortcut. In existing video editing software, the front-end UI code and video rendering function code are highly coupled. In this embodiment, the code is separated according to function to decouple the front-end UI and video rendering function code of the video editing software. The video rendering function code of the existing video editing software is then encapsulated into an independent API, allowing external code to call the video rendering function through a standard API, achieving automated video generation.
[0148] like Figure 3 In one example, the endpoint's front-end code could be a video generation page built on Astro, responsible for receiving callback notifications. For instance, when the material generation model is complete, it can send a notification to the endpoint to indicate that the subsequent video generation process can continue. The endpoint's back-end code could be obtained by API-izing the video rendering code of video editing software, exposing it externally for external calls via an API interface.
[0149] In one optional embodiment, the process of generating text information by calling a generative language model based on the subtitle description information to obtain target subtitles includes: matching corresponding keywords according to the video category, calling a pre-designed prompt word template, filling the prompt word template with keywords of the video category to obtain prompt words for the video category; inputting the prompt words of the video category into the generative language model to process the text information and obtain target subtitles related to the video category.
[0150] Keywords describe the theme of the video category; different video categories can correspond to different keywords, and different video types can also have different prompt word templates. Target subtitles include all the text content required for each video. By calling a generative language model, target subtitles can be automatically generated based on the video category without manual intervention, improving generation efficiency.
[0151] Further optionally, according to the target subtitle and the audio type description information, a text-to-speech model is called to convert the target subtitle into target audio that is adapted to the audio type description information. The target subtitle is used to provide text content, and the audio type description information is used to specify the audio type of the generated target audio, which includes but is not limited to audio format, audio language, audio tone, and audio quality, etc. Any audio type that can be used to specify the sound effect of the target audio is applicable to the present embodiment. Based on the target subtitle and the audio description information, a text-to-speech model is called to generate corresponding target audio and uploaded to a content distribution network, improving the automation level of audio generation.
[0152] Further, a multi-modal model is called to generate a green screen video of a digital human. The digital human refers to a virtual character generated based on AI technology, which can simulate the appearance, voice, and mouth shape of a real person, and can be used in scenarios such as intelligent customer service, short video production, and virtual anchor, but is not limited thereto. In the present embodiment, a plurality of video materials corresponding to each video instance identifier can be combined to generate a video with real person speaking effect.
[0153] In an optional embodiment, the digital human can be a pre-recorded real person video or picture. In subsequent embodiments, the real person video and picture are collectively referred to as video frames, and the number of video frames can be one or more. In this case, the description information of different digital humans can be implemented as identification information, which serves as the unique identity of the digital human and is used to obtain the video frames of the digital human corresponding to the identification information. In another optional embodiment, the video frames of the digital human can be dynamically generated based on a multi-modal model. In this case, the description information of the digital human can be a prompt word used to generate the digital human, and the description information of the digital human corresponding to each video instance identifier can be different.
[0154] Further optionally, in the step of calling the multi-modal model according to the target audio and the digital human description information, selecting the target digital human according to the digital human description information, and generating the green screen video based on the target audio and the target digital human to obtain the green screen video of the target digital human, the method further comprises: obtaining the video frame of the target digital human based on the digital human description information; calling the multi-modal model to perform multi-dimensional feature extraction on the target audio according to the video frame of the target digital human and the target audio, to obtain multi-dimensional speech features; wherein the multi-dimensional speech features comprise but are not limited to speech content features and speech emotion features; determining the lip control parameters of the target digital human according to the speech content features; determining the facial expression control parameters of the target digital human according to the speech emotion features; determining the body movement control parameters of the target digital human according to the speech content features and the speech emotion features; and generating the green screen video of the target digital human based on the lip control parameters, the facial expression control parameters and the body movement control parameters of the target digital human; wherein the lip movement, facial expression and body movement of the green screen video of the target digital human match the target audio.
[0155] In the present embodiment, the speech content features are used to reflect semantic information in the target audio, such as lexical content, grammatical structure, speech intent and speech rhythm, and are mainly used to drive the lip movement of the digital human to be synchronized with the semantics of the target audio. The speech emotion features are used to reflect the emotional state of the target audio, including but not limited to tone strength, speech speed, and tone change, and are mainly used to drive the facial expression and body movement of the digital human.
[0156] Further, the lip control parameters of the target digital human are determined according to the speech content features, the facial expression control parameters of the target digital human are determined according to the speech emotion features, and the body movement control parameters of the target digital human are determined according to the speech content features and the speech emotion features.
[0157] In the case where the lip control parameters, the facial expression control parameters and the body movement control parameters are obtained, the target digital human is driven to perform action rendering based on the lip control parameters, the facial expression control parameters and the body movement control parameters of the target digital human, to generate the green screen video of the corresponding target digital human, which is convenient for subsequent flexible replacement using a background image. Since the green screen video of the target digital human is generated based on the multi-dimensional speech features extracted from the content of the target audio, the dynamic performance of the lip movement, facial expression and body movement of the target digital human in the green screen video is ensured to be matched with the target audio in terms of semantics, timing and emotion, ensuring accurate alignment of the lip movement and improving the realism and viewing experience of the generated video.
[0158] Further, in the present embodiment, the target audio and the target subtitle are aligned, for example, by timestamp inference and alignment annotation of the target subtitle based on the timestamp of the speech content in the target audio, to ensure that the subsequent target subtitle display is completely synchronized with the target audio.
[0159] In this embodiment, one method for generating the background image is to call a text-based image model based on the background image description information to generate the target background image. Another method is to directly obtain a pre-generated or captured background image; this method is not limited.
[0160] In one optional embodiment, a set of video material generation states is maintained for each of the N video instance identifiers. The generation state of any video material of any type within each set includes: ready to generate, generating in progress, successfully generated, and failed generated. For example, for any video instance identifier, taking the multiple video materials that need to be generated for that video instance identifier, including: target subtitles, target audio, target digital human green screen video, and target background image, then the generation states for the target subtitles, target audio, target digital human green screen video, and target background image can be maintained separately.
[0161] Specifically, for each video instance identifier, the various video materials corresponding to that video instance identifier are marked as ready for generation. When calling multiple AI-based material generation models to generate the various video materials corresponding to that video instance identifier, if any material generation model returns a "generating in progress" response message, the video materials that the material generation model should generate are updated to the "generating in progress" state; if any material generation model returns a "generating successfully" message, the video materials that the material generation model should generate are updated to the "generating successfully" state; if any material generation model returns a "generating failed" message, the video materials that the material generation model should generate are updated to the "generating failed" state. For example... Figure 3 As shown, a subscription service is provided that can notify each material generation model of the generation status of the video material for that model, such as a successful generation status. The subscription service then returns a successful generation message to the server to inform it that the corresponding video material has been successfully generated. Figure 3 This example only uses the notification process for the target digital person as an example, but is not limited to this.
[0162] Optionally, if the generation status of the video footage is updated to a generation failure status, a failure notification message is output to the user who initiated the input operation. The failure notification message includes the footage type of the video footage that has been updated to a generation failure status and its corresponding video instance identifier. If a regeneration operation triggered by the user is received, the corresponding description information of the video footage is obtained based on the video instance identifier, and the corresponding AI-based footage generation model is invoked to regenerate the corresponding video footage.
[0163] In the embodiment, in the case that the plurality of video materials corresponding to each video instance identifier are generated, the plurality of video materials corresponding to each video instance identifier can be synchronously uploaded to the content distribution network, and the access links of the video materials corresponding to each video instance identifier in the content distribution network are obtained. In the above example, for any video instance identifier, in the case that the plurality of video materials corresponding to the video instance identifier include the target subtitle, the target audio, the green screen video of the target digital person and the target background picture, the target subtitle, the target audio, the green screen video of the target digital person and the target background picture are respectively uploaded to the content distribution network, and the access links of the target subtitle, the target audio, the green screen video of the target digital person and the target background picture are obtained.
[0164] Further, in response to the batch video generation trigger event, the plurality of video materials corresponding to the N video instance identifiers are respectively obtained according to the access links of the plurality of video materials corresponding to the N video instance identifiers in the content distribution network, and the video generation is performed according to the plurality of video materials corresponding to the N video instance identifiers and the target video template to obtain the N videos under the video category.
[0165] In the above example, in response to the batch video generation trigger event, the target subtitle, the target audio, the green screen video of the target digital person and the target background picture corresponding to the N video instance identifiers are respectively obtained according to the access links of the target subtitle, the target audio, the green screen video of the target digital person and the target background picture corresponding to the N video instance identifiers in the content distribution network, and the video generation is performed according to the target subtitle, the target audio, the green screen video of the target digital person, the target background picture corresponding to the N video instance identifiers and the target video template to obtain the N videos under the video category.
[0166] Further optionally, the videos corresponding to the N video instance identifiers and the target video template are uploaded to the content distribution network, and the access links of the videos corresponding to the N video instance identifiers and the target video template in the content distribution network are obtained, and the access links are added to the video generation result page; in response to a viewing operation of the video result page, the video generation result page is displayed, and the video generation result page includes the access links of the group of video materials, the videos and the target video template corresponding to at least one video instance identifier in the N video instance identifiers in the content distribution network.
[0167] In the embodiment, the video materials, the videos and the target video template corresponding to each video instance identifier are stored in the content distribution network, and the server does not need to store a large number of files, thereby reducing the disk I / O load and ensuring the service stability. Further, for the video materials, the videos and the target video template corresponding to each video instance identifier, the access links stored in the content distribution network can avoid data inflation, reduce query pressure, and support larger-scale data management.
[0168] In this optional embodiment, in response to the triggering operation of identifying a corresponding set of video materials, videos and / or target video templates for at least one video instance, the set of video materials, videos and / or target video templates corresponding to the at least one video instance identification information are accessed. By hosting the videos and video templates in the CDN, the video generation result page can directly load the videos from the content distribution network, which is less bandwidth-consuming and faster in loading than pulling the videos from the server, and can significantly improve the performance, which is suitable for large-scale video browsing scenarios.
[0169] The detailed implementation and beneficial effects of each step in the method of the embodiment have been described in the foregoing embodiments, and will not be described in detail here.
[0170] In addition, in some of the processes described in the foregoing embodiments and accompanying drawings, a plurality of operations appearing in a specific order are included, but it should be clearly understood that these operations can be executed in the order appearing in this document or in parallel, and the serial numbers of the operations, such as 11, 12, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the "first", "second", etc. described herein are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do "first" and "second" represent different types.
[0171] Figure 4 An electronic device structure schematic diagram is provided for the exemplary embodiments of the present application. As shown in the figure, the device includes: the device includes: a memory 44, a processor 45. Figure 4 As shown in the figure, the device includes: the device includes: a memory 44, a processor 45.
[0172] The memory 44 is used to store computer programs and can be configured to store various data to support operations on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, initial video templates, rendering rules of video materials, rendering parameter information, etc.
[0173] The memory 44 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0174] The processor 45 is coupled with the memory 44 and is configured to execute a computer program in the memory 44 to acquire an initial video template, the initial video template including rendering rules of a plurality of video materials required for generating a video and a hierarchical relationship between the plurality of video materials; parse the initial video template with a material type as a split variable to obtain information segments corresponding to the plurality of material types; extract rendering parameter information corresponding to the plurality of material types from the information segments corresponding to the plurality of material types respectively; and templateize the rendering parameter information corresponding to the plurality of material types respectively to obtain a plurality of materialized templates, each of the materialized templates being used to describe a rendering rule of a video material.
[0175] In an optional embodiment, the processor 45 parses the initial video template with a material type as a split variable to obtain information segments corresponding to the plurality of material types, including: loading an XML document corresponding to the initial video template, the XML document including a root element and a plurality of non-root elements connected with the root element, the plurality of non-root elements including a plurality of specific elements, each of the specific elements being used to describe a rendering rule of a material type; starting from the root element, traversing the non-root elements in the XML document to identify the plurality of specific elements; and extracting a plurality of information segments in which the plurality of specific elements are located as the information segments corresponding to the plurality of material types.
[0176] In an optional embodiment, the processor 45 starts from the root element to traverse the non-root elements in the XML document to identify the plurality of specific elements, including: S1, starting from the root element, traversing the non-root elements in the XML document; S2, for a currently traversed non-root element, acquiring an element tag included in the currently traversed non-root element; S3, if the element tag is a specific tag, determining whether the currently traversed non-root element includes a child element; S4, if the currently traversed non-root element includes the child element, taking the child element as the currently traversed non-root element and returning to step S2; S5, if the element tag is a non-specific tag, continuing with a next non-root element and returning to step S2; and S6, if the currently traversed non-root element does not include the child element, taking the currently traversed non-root element as a specific element.
[0177] S5, if the element tag is a non-specific tag, continuing with a next non-root element and returning to step S2; and S6, if the currently traversed non-root element does not include the child element, taking the currently traversed non-root element as a specific element.
[0178] In an optional embodiment, the processor 45 extracts the rendering parameter information corresponding to each of the plurality of material types from the information segments corresponding to the plurality of material types respectively, including: for each information segment, extracting path information of a material file and at least one attribute value of the material file from the information segment, the attribute value being used for rendering the material file; reading the material file according to the path information, and determining the material type described by the information segment according to the extension of the material file; taking at least one attribute to which the at least one attribute value belongs as at least one rendering parameter, and taking the at least one attribute value as a default parameter value of the at least one rendering parameter, to obtain the rendering parameter information corresponding to the material type described by the information segment.
[0179] In an optional embodiment, the processor 45 templates the rendering parameter information corresponding to each of the plurality of material types to obtain a plurality of materialization templates, including: for each material type, selecting at least one variable parameter from the rendering parameter information corresponding to the material type; associating a plurality of candidate parameter values with the at least one variable parameter; and generating a materialization template corresponding to the material type according to the plurality of candidate parameter values associated with the at least one variable parameter.
[0180] In an optional embodiment, the processor 45 selects the at least one variable parameter from the rendering parameter information corresponding to the material type, including: pre-configuring a weight value of each rendering parameter for the material type; and parsing each rendering parameter from the rendering parameter information corresponding to the material type, and selecting at least one rendering parameter with a weight value greater than a set weight threshold value from the each rendering parameter as the at least one variable parameter.
[0181] In an optional embodiment, the processor 45 generates the materialization template corresponding to the material type according to the plurality of candidate parameter values associated with the at least one variable parameter, including: adding each rendering parameter corresponding to the material type and a default parameter value of each rendering parameter to a preset template file; adding the plurality of candidate reference values associated with the at least one variable parameter in the preset template file; and adding a placeholder for carrying a material file corresponding to the material type in the preset template file, to obtain the materialization template corresponding to the material type.
[0182] In an optional embodiment, the processor 45 generates a batch video generation task in response to an input operation on the video generation page for the number of videos and the video category, the batch video generation task including the number of videos N and the video category, N being an integer greater than or equal to 2; generates N video instance identifiers according to the batch video generation task, and generates a set of video materials related to the video category for each video instance identifier; for each video instance identifier, determines at least one target materialization template from a plurality of materialization templates according to the material type in the set of video materials corresponding to the video instance identifier; combines the at least one target materialization template based on the hierarchical relationship between the plurality of video materials included in the initial video template to obtain a target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rule and the hierarchical relationship of the set of video materials; and performs video generation processing according to the target video template corresponding to each of the N video instance identifiers and the set of video materials to obtain N videos under the video category.
[0183] Further, as shown in Figure 4 , the electronic device further includes a communication component 46, a display 47, a power supply component 48, an audio component 49, and other components. Figure 4 Some components are only schematically shown in the computing platform, and it does not mean that the computing platform only includes Figure 4 the components shown. In addition, Figure 4 the components in the dashed box are optional components, not mandatory components, and the specific product form of the working node can be determined. The working node of the embodiment can be implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone, or an IOT device, or as a server device such as a conventional server, a cloud server, or a server array. If the working node of the embodiment is implemented as a terminal device such as a desktop computer, a notebook computer, or a smart phone, it can include Figure 4 the components in the dashed box; if the working node of the embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it can not include Figure 4 the components in the dashed box.
[0184] Correspondingly, the embodiment of the application also provides a computer readable storage medium storing a computer program, which can implement each step that can be executed by the electronic device in the above method embodiment when the computer program is executed.
[0185] The above-described memory can be implemented by any type of volatile or nonvolatile memory devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0186] The above-described communication component is configured to facilitate communication between the device in which the communication component is located and other devices in a wired or wireless manner. The device in which the communication component is located can access a wireless network based on a communication standard, such as a WiFi, 2G, 3G, 4G / LTE, 5G, or the like mobile communication network, or a combination thereof. In an example embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast managing system via a broadcast channel. In an example embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Blue Tooth (BT) technology, and other technologies.
[0187] The above-described display includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touch or a slide action, but also detect a duration and a pressure related to a touch or a slide operation.
[0188] The power component provides power to various components of the device in which the power component is located. The power component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which the power component is located.
[0189] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive an external audio signal when the device in which the audio component is located is in an operational mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in a memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0190] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable code.
[0191] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 means for performing the function specified by the flowchart illustrations and / or block diagrams block or blocks.
[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 means for performing the function specified by the flowchart illustrations and / or block diagrams block or blocks.
[0193] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 Figure 1
[0194] In one typical configuration, the computing device includes one or more processors (Central Processing Units, CPUs), input / output interfaces, network interfaces, and memory.
[0195] The memory can include non-persistent memory and / or persistent memory, both of which can be volatile and / or non-volatile. Non-persistent memory can include, for example, a random access memory (RAM), which can be a single RAM or dual RAM. Persistent memory can include, for example, a read-only memory (ROM), a flash memory, or a combination of both. The memory is an example of computer-readable media.
[0196] Computer-readable media includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology for storing information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carriers.
[0197] It should also be noted that the terms "comprising", "comprises", "including", "includes" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0198] The above embodiments of the present application are only used to illustrate the technical solutions of the present application, and not intended to limit the present application. Although the present application has been described in detail, it should be understood that those skilled in the art can make various modifications and changes without departing from the spirit and scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of the claims of the present application.< / textelement> < / textelement> < / textelement> < / fontsize> < / fontsize> < / fontsize> < / position> < / position> < / position> < / fontsize> < / position> < / textelement> < / textelement> < / textelement> < / textelement> < / videotemplate> < / path> < / settings> < / author> < / metadata> < / texts> < / images> < / video> < / text> < / images> < / audio> < / video> < / audio> < / video> < / media> < / media> < / metadata> < / template> < / template> < / position> < / path> < / images> < / metadata> < / template>
Claims
1. A method for video template variableization, characterized in that, include: Obtain an initial video template, which includes rendering rules for various video materials required to generate the video and the hierarchical relationship between the various video materials; Using the material type as a splitting variable, the initial video template is parsed to obtain information fragments corresponding to various material types; Extract the rendering parameter information corresponding to each of the various material types from the information fragments corresponding to each of the various material types; The rendering parameter information corresponding to each of the various material types is templated to obtain multiple material templates, including: for each material type, adding the rendering parameter information corresponding to the material type to a preset template file, wherein the rendering parameter information includes each rendering parameter and the default parameter value of each rendering parameter, and each rendering parameter includes at least one variable parameter; adding multiple candidate parameter values associated with at least one variable parameter to the preset template file; and adding placeholders to the preset template file to carry the material file corresponding to the material type, so as to obtain the material template corresponding to the material type; each material template is used to describe the rendering rules of a video material.
2. The method according to claim 1, characterized in that, Using the material type as a splitting variable, the initial video template is parsed to obtain information fragments corresponding to various material types, including: Load the XML document corresponding to the initial video template. The XML document includes a root element and multiple non-root elements connected to the root element. The multiple non-root elements include multiple specific elements, each of which is used to describe the rendering rules of a material type. Starting from the root element, the non-root elements in the XML document are traversed to identify the multiple specific elements; Extract multiple information fragments containing the multiple specific elements, and use them as information fragments corresponding to the multiple material types.
3. The method according to claim 2, characterized in that, Starting from the root element, the non-root elements in the XML document are traversed to identify the plurality of specific elements, including: S1. Starting from the root element, traverse the non-root elements in the XML document; S2. For the currently traversed non-root element, get the element tags contained in the currently traversed non-root element; S3. If the element tag is a specific tag, determine whether the currently traversed non-root element contains child elements; S4. If the currently traversed non-root element contains child elements, then the child element is taken as the currently traversed non-root element, and the process returns to step S2. S5. If the element label is a non-specific label, then continue to the next non-root element and return to step S2. S6. If the currently traversed non-root element does not contain child elements, then the currently traversed non-root element is treated as a specific element.
4. The method according to claim 1, characterized in that, Extract the rendering parameter information corresponding to each of the various material types from the information fragments corresponding to each of the various material types, including: For each information fragment, the path information of the material file and at least one attribute value of the material file are extracted from the information fragment, and the attribute value is used to render the material file; The material file is read based on the path information, and the material type described by the information fragment is determined based on the file extension. The at least one attribute to which the at least one attribute value belongs is used as at least one rendering parameter, and the at least one attribute value is used as the default parameter value of the at least one rendering parameter, so as to obtain the rendering parameter information corresponding to the material type described by the information fragment.
5. The method according to any one of claims 1-4, characterized in that, The rendering parameter information corresponding to each of the various material types is templated to obtain the various material templates, including: For each type of material, select at least one variable parameter from the rendering parameter information corresponding to that material type; Associate multiple candidate parameter values with the at least one variable parameter; Based on the multiple candidate parameter values associated with the at least one variable parameter, a material template corresponding to the material type is generated.
6. The method according to claim 5, characterized in that, Select at least one variable parameter from the rendering parameter information corresponding to the material type, including: Pre-configure the weight values of each rendering parameter for the material type; Each rendering parameter is parsed from the rendering parameter information corresponding to the material type, and at least one rendering parameter with a weight value greater than a set weight threshold is selected as the at least one variable parameter.
7. The method according to any one of claims 1-4 and 6, characterized in that, Also includes: In response to input operations on the video generation page regarding the number of videos and video categories, a batch video generation task is generated. The batch video generation task includes the number of videos N and the video categories, where N is an integer ≥ 2. Based on the batch video generation task, N video instance identifiers are generated, and a set of video materials related to the video category are generated for each video instance identifier; For each video instance identifier, at least one target material template is determined from the multiple material templates based on the material type in a set of video materials corresponding to the video instance identifier. Based on the hierarchical relationship between the various video materials included in the initial video template, the at least one target material template is combined to obtain the target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationship of the group of video materials. Based on the target video templates and a set of video materials corresponding to each of the N video instance identifiers, video generation processing is performed to obtain N videos under the video category.
8. An electronic device, characterized in that, include: A processor and a memory, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1-7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-7.
10. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Media information material processing method and device, electronic equipment and storage medium
CN116801008A
Multi-style video template creation method, system and device and storage medium
CN117412119A