Video batch generation method, device, storage medium and program product

By generating various materials for video instance identifiers and combining them with target templates to generate videos, the problems of monotonous video styles and homogenized content are solved, thereby achieving diversity in video content and improving generation efficiency.

CN120186429BActive Publication Date: 2026-03-27BEIJING 58 INFORMATION TTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies that generate videos in batches based on preset templates have monotonous styles and serious content homogenization.

Method used

For each video instance, the system retrieves descriptive information from various materials, calls an AI-based material generation model to generate video materials, and hosts the video generation through a content distribution network. It also combines target video templates to generate videos, thus enriching the video material resources and template types.

Benefits of technology

It improves the diversity and efficiency of video content generation, reduces server load, and enhances the flexibility and diversity of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186429B_ABST
    Figure CN120186429B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video batch generation method and device, a storage medium and a program product. In the embodiments of the present application, description information of multiple materials is obtained for each video instance identifier, and based on the description information of the multiple materials, a plurality of material generation models based on artificial intelligence are called to generate video materials, and each video instance identifier generates its corresponding multiple video materials, which enriches the video material resources and improves the diversity of video content. In addition, the video template of each video instance identifier is determined according to the corresponding multiple video materials, and has a corresponding relationship, which enriches the types of video templates and further improves the diversity of generated video content. Furthermore, the video materials are uploaded to a content distribution network to obtain the multiple video materials corresponding to the video instance identifiers from the content distribution network respectively, and video generation is performed in combination with the corresponding target video template, thereby reducing the load pressure of the server, improving the generation efficiency, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a video batch generation method, device, storage medium and program product. BACKGROUND

[0002] In the field of video generation, in order to reduce the learning threshold and improve the video generation efficiency, in some video editing software, a preset template is provided, the preset template contains some fixed content, the user's material is replaced with the fixed content such as text and picture of the preset template, and the required video can be generated. By generating a video through a preset template, the process of video generation is simplified, and the batch generation of videos is suitable. However, the video generated based on the preset template is monotonous in style and has serious content homogenization. SUMMARY

[0003] The embodiments of the present application provide a video batch generation method, device, storage medium and program product to improve the diversification degree of video generation.

[0004] The embodiments of the present application provide a video batch generation method, comprising: in response to an input operation on a video generation page for a video quantity and a video category, generating a batch video generation task, the batch video generation task comprising a video quantity N and a video category, N being an integer greater than or equal to 2; according to the batch video generation task, generating N video instance identifiers, and obtaining a set of video material description information related to the video category for each video instance identifier, each set of video material description information comprising description information of multiple video materials; for each video instance identifier, calling multiple material generation models based on artificial intelligence according to the set of video material description information corresponding to the video instance identifier, generating multiple video materials corresponding to the video instance identifier, and synchronously uploading the multiple video materials corresponding to the video instance identifier to a content distribution network; determining a target video template corresponding to each video instance identifier according to the multiple video materials corresponding to each video instance identifier, the target video template corresponding to each video instance identifier being used to describe the rendering rule and hierarchical relationship of the multiple video materials corresponding to the video instance identifier; in response to a batch video generation trigger event, obtaining the multiple video materials corresponding to the N video instance identifiers from the content distribution network respectively according to the N video instance identifiers; and performing video generation processing according to the multiple video materials corresponding to the N video instance identifiers and the target video templates to obtain N videos under the video category.

[0005] The embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory being configured to store a computer program, and the processor being coupled to the memory and configured to execute the computer program to implement the steps in the methods provided by the embodiments of the present application.

[0006] The embodiment of the present application further provides a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor can implement the steps in the above method.

[0007] The embodiment of the present application further provides a computer program product, the computer program product comprises computer programs / instructions, when the computer programs / instructions are executed by a processor, the processor can implement the steps in the above method embodiment.

[0008] In the embodiment of the present application, the description information of the corresponding multiple materials is obtained for the N video instance identifiers, based on the description information of the multiple materials, multiple material generation models based on artificial intelligence are respectively called to generate video materials, and multiple video materials corresponding to each video instance identifier are generated, so as to enrich the video material resources and improve the diversity of video content; and the video template of each video instance identifier is determined according to the corresponding multiple video materials, has a corresponding relationship, enriches the types of video templates, and further improves the diversity of generated video content. In addition, the video materials are uploaded to the content distribution network, so as to obtain the multiple video materials corresponding to the video instance identifier from the content distribution network, and generate a video in combination with the corresponding target video template, so as to reduce the load pressure of the server through the hosting of the content distribution network, and improve the generation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0009] The accompanying drawings for describing the present application are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0010] Figure 1 A structural schematic diagram of a video batch generation system according to an example embodiment of the present application is provided;

[0011] Figure 2 A structural schematic diagram of a video batch generation system according to another example embodiment of the present application is provided;

[0012] Figure 3 A flowchart of a video batch generation method according to an example embodiment of the present application is provided;

[0013] Figure 4 A structural schematic diagram of an electronic device according to still another example embodiment of the present application is provided. DETAILED DESCRIPTION

[0014] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in combination with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0015] It should be noted that, in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal. In addition, the various models (including but not limited to language models or large models) involved in the present application are in compliance with relevant laws and standard regulations.

[0016] In addition, it should be noted that, in the case where the embodiments of the present application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of the present application include but are not limited to: touch operations, gesture operations, voice operations, head movement operations, eye movement operations and various modes of interaction; wherein the touch operation includes but is not limited to: click operation, double-click operation, long-press operation, sliding operation, pinch operation or mouse hovering operation, etc. The sliding operation includes but is not limited to: straight line sliding, curve sliding, etc.

[0017] Further, it should be noted that, in the case where the embodiments of the present application involve the jump between the first interface and the second interface, the jump mode involved in the embodiments of the present application includes but is not limited to: directly jumping from the first interface to the second interface, jumping from the first interface to the task interface first and jumping to the second interface after completing the corresponding task operation on the task interface; completing the corresponding task operation on the task interface includes but is not limited to: in the case where the task interface is implemented as a game interface, completing the game operation on the game interface; in the case where the task interface is implemented as an identity authentication interface, completing the identity authentication on the identity authentication interface; in the case where the task interface is implemented as a recharge interface, completing the recharge operation on the recharge interface; etc.

[0018] In response to the technical problems of monotonous video style and serious content homogenization caused by batch generation based on preset templates, in the embodiments of the present application, description information of a plurality of materials corresponding to N video instance identifiers is obtained, and a plurality of material generation models based on artificial intelligence are called based on the description information of the plurality of materials to generate video materials, so as to generate a plurality of video materials corresponding to each video instance identifier, enrich video material resources, and improve the diversity of video content. In addition, the video template of each video instance identifier is determined according to the corresponding plurality of video materials, and has a corresponding relationship, which enriches the types of video templates and further improves the diversity of generated video content. In addition, the video materials are uploaded to the content distribution network to obtain the plurality of video materials corresponding to the video instance identifiers from the content distribution network, and the video generation is combined with the corresponding target video template, and the service load pressure is reduced through the hosting of the content distribution network, and the generation efficiency is improved.

[0019] The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0020] Figure 1 The structural schematic diagram of the video batch generation system provided by an exemplary embodiment of the present application is shown. The video batch generation system 100 includes a client 10, a server 11, a content distribution network 13, and a plurality of material generation models, such as material generation model a1, material generation model a2, and material generation model ax, x is greater than or equal to 2 and is an integer. Wherein, on the basis of adopting the client-server architecture, the client 10 provides a video generation page 12 for users, which is a page for providing batch video generation service, and the user can initiate a batch video generation task to the server 11 through the video generation page 12. Correspondingly, the server 11 receives the batch video generation task initiated by the client 10 through the video generation service page 12, on the one hand, calls a plurality of material generation models to generate a plurality of video materials required for each video, and uploads the generated video materials to the content distribution network 13; on the other hand, obtains a plurality of pre-generated materials from the content distribution network 13 to generate batch videos, so as to generate video materials and videos by means of the computing resources, network bandwidth and storage resources of the server 11 and the like.

[0021] In the present embodiment, in response to the input operation on the video generation page for the number of videos and the video category, a batch video generation task is generated, which includes the number of videos N and the video category, N is an integer greater than or equal to 2;

[0022] The batch video generation task is used to describe the generation of N videos under a certain video category. The specific implementation of the video category is not limited. For example, it can be implemented as a single-level video category, such as, but not limited to, business registration or legal consultation, etc. Alternatively, the video category can also be implemented as a multi-level video category. For example, the first-level video category can be business registration; the second-level video category of the first-level video category can be cleaning.

[0023] In this embodiment, according to the batch video generation task, N video instance identifiers are generated. The N video instance identifiers are all different, and each video instance identifier is used to uniquely represent a video to be generated, so as to track the required video materials and video templates of the video to be generated, that is, the video instance identifier can also be used as the unique identity of the related content (such as video materials) of the video to be generated. The way of generating N video instance identifiers is not limited. For example, it can be an increasing sequence starting from an arbitrary integer with a step of 1; or it can also be a string, etc.

[0024] In this embodiment, for each video instance identifier, a set of video material description information related to the video category is generated, and the set of video material description information is used to generate the video material corresponding to the video instance identifier. Each set of video material contains video materials of multiple material types. Material type refers to the type of different video elements that constitute video content, including but not limited to: audio, video, background picture, subtitle, digital person, etc.

[0025] In this embodiment, the implementation of generating a set of video material description information related to the video category for each video instance identifier is not limited.

[0026] In an optional implementation, for any video instance identifier, video material description information can be randomly extracted from the multiple material types stored in the basic material library, and the extracted multiple video material description information is used as a set of video material description information of any video instance identifier. The basic material library stores multiple video material description information of multiple video categories.

[0027] In another optional implementation, for any video instance identifier, the semantic similarity between the multiple video material description information in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the video material description information that meets the similarity condition in the multiple similarity information is used as a set of video material description information of the video instance identifier.

[0028] In yet another optional embodiment, for any video instance identifier, a material description information generation model is invoked according to the video category, the model being configured to generate description information of a plurality of video materials related to the video category. The material description information generation model is trained in combination with description information of a large number of sample video categories and sample video materials of different material types, and through learning semantic correlation between the sample video categories and the sample video materials of different material types, the model can be used to generate description information of a plurality of video materials related to different video categories.

[0029] Further, for each video instance identifier, a plurality of material generation models based on artificial intelligence are invoked to generate a plurality of video materials corresponding to the video instance identifier according to a set of video material description information corresponding to the video instance identifier, and the plurality of video materials corresponding to the video instance identifier are synchronously uploaded to a content distribution network.

[0030] The material generation model can generate at least part of the plurality of video materials. For example, one material generation model can generate one kind of video material, or one material generation model can generate two or more kinds of video materials, and the number of kinds of video materials generated by one material generation model is not limited.

[0031] In the embodiment, in the case of generating video materials of the video instance identifier, a target video template corresponding to each video instance identifier is determined according to the plurality of video materials corresponding to each video instance identifier, and the target video template corresponding to each video instance identifier is used to describe the rendering rule and the hierarchical relationship of the plurality of video materials corresponding to the video instance identifier.

[0032] In the embodiment, the determination method of the target video template is not limited. Two examples are provided below, but the present disclosure is not limited thereto.

[0033] In an optional embodiment, a video template library is constructed in advance, the video template library including a plurality of video templates, each video template including a pre-designed rendering rule and hierarchical relationship of video materials. The video templates in the video template library can be classified according to video categories. When determining the target video template corresponding to each video instance identifier, N video templates can be randomly determined from the plurality of video templates under the video category as the target video templates of the N video instance identifiers, and each video instance identifier corresponds to one target video template.

[0034] In yet another optional embodiment, the corresponding materialized templates corresponding to each video instance identifier can be generated respectively according to the corresponding plurality of materials of each video instance identifier, and then the plurality of materialized templates corresponding to the plurality of video materials are combined according to the combination logic to obtain the target video template corresponding to each video instance identifier. The combination logic includes a hierarchical relationship, which refers to the layer superposition order of the plurality of video materials in the video generation process, and is used to determine the front and back coverage relationship and occlusion logic of the plurality of video materials in the video generation. For example, in generating a video, the subtitle can be located in the relatively upper layer of the layer, the background picture can be located in the relatively bottom layer of the layer, the digital person can be superimposed on the upper layer of the background picture, and the layer below the subtitle.

[0035] In this embodiment, in response to the batch video generation trigger event, the plurality of video materials corresponding to the N video instance identifiers are respectively obtained from the content distribution network according to the N video instance identifiers; and the video generation processing is performed according to the plurality of video materials corresponding to the N video instance identifiers and the target video template to obtain the N videos under the video category. The video generation processing refers to the process of filling and rendering the target video template with video materials based on the target video template corresponding to each video instance identifier and a group of video materials to obtain a video corresponding to the video instance identifier. For example, for the N video instance identifiers, a group of video materials corresponding to each of the N video instance identifiers can be filled into the N target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain the N videos under the video category.

[0036] In this embodiment, the specific implementation of the batch video generation trigger event is not limited, and can be flexibly configured according to actual application requirements. For example, it can be in the case that the plurality of materials corresponding to the N video instance identifiers are all generated; or it can be in the case that each of the plurality of materials corresponding to the video instance identifier is generated, in which case the video generation is performed for the plurality of video materials of the video instance identifier; or the batch video generation trigger event can be a preset trigger time, such as a period of time after responding to the input operation of the video quantity and the video category on the video generation page, for example, 2 hours, 1 day or 1 week, etc., without limitation on the time span.

[0037] It should be noted that the generation of a video corresponding to each video instance identifier, N video instance identifiers can be N to-be-generated videos, and the video generation processing process corresponding to each video instance identifier is asynchronous when the video generation processing is performed on the plurality of video materials corresponding to the N video instance identifiers, and the video generation processing corresponding to each video instance identifier does not affect each other, so as to improve the generation efficiency of the N videos.

[0038] In the embodiments of the present application, the description information of the corresponding multiple materials is obtained for the N video instance identifiers, and based on the description information of the multiple materials, multiple material generation models based on artificial intelligence are respectively called to generate video materials, and multiple video materials corresponding to each video instance identifier are generated, which enriches the video material resources and improves the diversity of video content. In addition, the video template of each video instance identifier is determined according to the corresponding multiple video materials, and has a corresponding relationship, which enriches the types of video templates and further improves the diversity of generated video content. In addition, the video materials are uploaded to the content distribution network to obtain the multiple video materials corresponding to the video instance identifier from the content distribution network, and the video generation is combined with the corresponding target video template, and the load pressure of the server is reduced through the hosting of the content distribution network, and the generation efficiency is improved.

[0039] In the above embodiments, the generation of multiple materials and batch video generation combined with the content distribution network are introduced, and the specific implementation of the target video template in the embodiments of the present application is not limited as described in the above embodiments. The following gives an implementation of generating a target video template based on variable processing of materials, but is not limited thereto.

[0040] In the present embodiment, multiple materialized templates obtained by variable processing of an initial video template are obtained. Optionally, the initial video template can be implemented as an MLT (Media Lovin' Toolkit) template, which is an open source framework for multimedia processing and is based on an XML format file that records all parameters of video editing, such as video clips on the timeline, audio tracks, filter effects, transitions, etc., which can be used for video editing.

[0041] The initial video template includes rendering rules of multiple video materials required for generating a video, and a hierarchical relationship between the multiple video materials. The rendering rules corresponding to each video material are used to control the rendering logic of the video material in the rendering process to achieve the expected rendering result in the generated video. The hierarchical relationship between the multiple video materials in the initial video template refers to the layer stacking order of the multiple video materials in the video generation process in the initial video template, which is used to determine the front and back coverage relationship and occlusion logic of the multiple video materials in the video generation. For example, in generating a video, subtitles can be located in the upper layer of the layer, background pictures can be located in the bottom layer of the layer, digital people can be superimposed on the upper layer of the background picture, and can be located below the layer of the subtitles.

[0042] In the embodiment, the variable processing part is embodied in the material type as a split variable, the initial video template is structured and the template of the rendering parameter information is obtained. The material template corresponds to the material type. For example, the material type includes audio, video, background picture, subtitle, digital person and other material types. The corresponding material template types include but are not limited to audio material template, video material template, background picture material template, subtitle material template, digital person material template and the like. In other words, each material type corresponds to a material template, and each material template is used to describe the rendering rule of the video material.

[0043] In the embodiment, the timing of the variable processing is not limited. For example, the initial video template can be pre-processed. For another example, the initial video template can be dynamically processed. The details of how to process the variable can refer to the subsequent embodiments.

[0044] In the embodiment, the variable processing material template can be modularly reorganized. For each video instance identifier, at least one target material template is determined from a plurality of material templates according to the material type in the video material corresponding to the video instance identifier. The material type in the video material corresponds to the material template in a one-to-one manner.

[0045] For example, in the case of audio, video, background picture, subtitle material types included in the video material, the target material template can include the target material template corresponding to the audio, video, background picture, subtitle respectively; for another example, in the case of audio, video, background picture, subtitle, digital person material types included in the video material, the target material template can include the target material template corresponding to the audio, video, background picture, subtitle, digital person respectively.

[0046] Further, in the case of obtaining at least one target material template corresponding to a group of video materials, based on the hierarchical relationship between the plurality of video materials included in the initial video template, the at least one target material template is combined to obtain a target video template corresponding to the video instance identifier, and the target video template is used to describe the rendering rule and hierarchical relationship of a group of video materials.

[0047] The hierarchical relationship between the plurality of video materials included in the initial video template can represent the layer superposition order of the plurality of material templates, which can be used to organize and integrate the target material template, thereby forming a target video template with clear hierarchical relationship and rendering rule of the corresponding video material. The hierarchical relationship of the group of video materials described by the target video template is consistent with the hierarchical relationship of the group of video materials in the initial video template.

[0048] In the embodiment, the N video instance identifiers correspond to respective target video templates, that is, each set of video materials has a respective target video template, which enriches the types of target video templates used in batch video generation. Based on the rendering rules described by the N target video templates, the video generation processing is performed on the grouped video materials corresponding to the N video instance identifiers, so that the style of each batch-generated video matches the video template used, thereby improving the diversification of the batch-generated video content.

[0049] After obtaining the target video template corresponding to the video instance identifier, the video generation processing is performed on the set of video materials corresponding to the respective target video templates of the N video instance identifiers, to obtain N videos under the video category. For example, for the N video instance identifiers, the set of video materials corresponding to each of the N video instance identifiers can be filled into the N target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain the videos under the video category.

[0050] In an optional embodiment, the video generation corresponding to the N video instance identifiers can be processed in batches, and M videos are processed in parallel in each batch, where M is less than N and M is an integer; after the generation of M videos is completed, the generation of the next M videos is continued until the video corresponding to each of the N video instance identifiers is generated. Through batch processing, the resource utilization rate of the server is improved, and the server overload caused by high concurrency is avoided.

[0051] In the embodiment, in batch video generation, the materialized templates of multiple material types obtained by variable processing of the initial video template are used to reorganize the materialized templates to obtain the video templates required for video generation; then, for the generation of each video, a video instance identifier is bound to a set of video materials corresponding to the video instance identifier, and based on the set of video materials corresponding to the video instance identifier, a target materialized template for reorganization is determined from the multiple materialized templates, and the target materialized template is organized and integrated in combination with the hierarchical relationship between the materialized templates provided by the initial template to obtain the video template corresponding to the video instance identifier, which is used for video generation processing of the set of video materials corresponding to each video instance identifier; since the video template can be personalized and reorganized in combination with the video materials of the generated video, the flexibility of video generation is improved; and the reorganized video template has a corresponding relationship with the video instance identifier, which to some extent enriches the types of video templates, so that the style of each batch-generated video matches the video template used, thereby improving the diversification of the batch-generated video content.

[0052] The variable processing method is described below.

[0053] In an optional embodiment, the step of variable processing of the initial video template includes: obtaining the initial video template, the initial video template including rendering rules of various video materials required for generating a video and hierarchical relationships among the various video materials; parsing the initial video template with the material type as a split variable to obtain information segments corresponding to the various material types; respectively extracting rendering parameter information corresponding to the various material types from the information segments corresponding to the various material types respectively;

[0054] templating the rendering parameter information corresponding to the various material types to obtain various materialization templates, each materialization template being used to describe a rendering rule of a video material.

[0055] Further, after obtaining the initial video template, the initial video template is parsed with the material type as a split variable. The material type refers to different categories of video materials, which is used to distinguish the classification criteria of different video materials, and the split variable refers to the independent information segments obtained by dividing the initial video template according to the material type when parsing the initial video template. For example, the initial video template contains two video materials of image and audio, and the initial video template is divided into independent information segments of image and audio as the classification criteria, and subsequent processing is performed respectively.

[0056] The information segment refers to the information segment corresponding to the material type extracted from the initial video template, and the information segment can include rendering parameter information of the corresponding material type. For example, an information segment can include rendering parameter information such as path information, playing time, and special effect application of the material type.

[0057] In the embodiments of the present application, the rendering parameter information is used to describe at least one attribute of a material file of a material type, each attribute can be regarded as a rendering parameter, and the attribute value of each attribute in the initial video template can be regarded as the default parameter value of the corresponding rendering parameter. The rendering parameters with the default parameter values form a rendering parameter information corresponding to the material file of the material type. The "material file" refers to various resource files used in the materialization template, mainly including various video materials such as pictures, videos, and audios. The rendering parameter information can be used to control the rendering logic of the material file, so that the material file can obtain the rendering effect in the finally generated video. The rendering parameter information quantitatively describes the attributes of the material file, so that the attributes of the video material are converted into the attributes of the rendering parameter information, and the rendering effect of the video material in the finally generated video is controlled through the rendering parameter information. In other words, the attributes of the video material are described as corresponding rendering parameters by the rendering parameter information, and the rendering parameters with the default parameter values define the attribute values of certain attributes of the material file.

[0058] The rendering parameter information can be extracted from information segments corresponding to the plurality of material types respectively. The rendering parameter information corresponding to each material type includes at least one rendering parameter and a corresponding default parameter value. Each rendering parameter is used to describe the attribute of a material file of a material type. In the embodiments of the present application, the specific attribute corresponding to the rendering parameter corresponding to each of the plurality of material types is not limited.

[0059] In the embodiments of the present application, the templating can convert the rendering parameter information corresponding to each of the plurality of material types from the rendering parameter with a fixedly configured default parameter value to a variable parameter set with a dynamically replaceable parameter value. The templating of the rendering parameter information corresponding to each of the plurality of material types can obtain a plurality of materialization templates corresponding to the plurality of material types respectively. Each materialization template is used to describe the rendering rule of a material type, and the rendering rule includes at least one variable parameter. Each variable parameter is associated with a plurality of candidate parameter values, so as to control the rendering effect of the generated video.

[0060] In the embodiments of the present application, the materialization template is a result of the templating of the rendering parameter information corresponding to each of the plurality of material types. It includes the path information placeholder of the material file and the variable parameter and the plurality of candidate parameter values associated with the variable parameter. The materialization template describes the rendering rule of the material file of the corresponding material type in the generated video. The rendering rule is used to describe the rendering logic followed by the video material in the video rendering process, so as to control the expected rendering effect of the material file in the generated video. The materialization template formed by converting the rendering parameter information from the rendering parameter with a fixedly configured default parameter value to the variable parameter set with a dynamically replaceable parameter value can assign different candidate parameter values to the variable parameter, so as to flexibly adjust the rendering effect of the material file in the generated video in different scenarios. The templating of the rendering parameter information forms a plurality of materialization templates that can be modularly recombined. The materialization templates for different material types can be recombined to obtain different target video templates, and then a plurality of diversified videos can be batch generated without designing and adjusting the video template for each video. This not only saves time and human resources, but also improves the flexibility, content diversity, and efficiency of video template generation, thereby realizing efficient batch generation of diversified style videos.

[0061] For example, for the same materialization template, different candidate parameter values can be assigned to the variable parameter to flexibly set the attribute values corresponding to various attributes of the material type corresponding to the materialization template, thereby realizing fine control of the rendering rule of the material file. The materialization template improves the flexibility of video generation, can significantly improve the efficiency and quality of video generation, and meets the diversified creation demand.

[0062] In the embodiments of the present application, the specific content of the candidate parameter value of the variable parameter associated with the variable parameter information of the plurality of material types corresponding to the template is not limited.

[0063] By associating a plurality of candidate parameter values with the variable parameter, the variable parameter can be flexibly adjusted according to the requirements, the flexibility and diversity of the video template generation are improved, and efficient batch generation of diversified style videos is realized.

[0064] In an optional embodiment, the initial video template is parsed by taking the material type as a split variable to obtain information segments corresponding to a plurality of material types, including: loading an XML document corresponding to the initial video template, the XML document including a root element and a plurality of non-root elements connected to the root element, the plurality of non-root elements including a plurality of specific elements, each specific element being used to describe the rendering rule of a material type; starting from the root element, traversing the non-root elements in the XML document to identify the plurality of specific elements; extracting the plurality of information segments in which the plurality of specific elements are located as the information segments corresponding to the plurality of material types. Wherein, the XML document of the initial video template can be loaded into the memory to form a tree structure that can be parsed. In the embodiments of the present application, the XML document can be loaded into the memory by a DOM (Document Object Model) parser to form a parsing method of tree structure.

[0065] The tree structure includes a root element and a non-root element. The root element is the top node of the XML document, and the non-root element is directly or indirectly nested under the root element. An XML document has only one root element, and the root element is the first element of the XML document and can be used as the starting point of the XML document.

[0066] The non-root element is a child element directly or indirectly nested under the root element, and is used to divide different information segments of the initial video template according to the material type. The non-root element includes a plurality of specific elements, and the plurality of specific elements can be identified by traversing the non-root elements in the XML document starting from the root element. The specific element is an element that can directly describe the rendering rule of a certain material type, and each specific element corresponds to a material type.

[0067] By extracting these information segments, the system can quickly identify the attributes of the material file and render according to the attribute values. This design of extracting information segments significantly improves the flexibility and diversity of video generation, meets the diversified creation requirements, and provides a technical basis for efficient batch generation of videos.

[0068] In an alternative embodiment, starting from the root element, the non-root elements in the XML document are traversed to identify a plurality of specific elements, comprising: S1, starting from the root element, the non-root elements in the XML document are traversed; S2, for the currently traversed non-root element, the element tag contained by the currently traversed non-root element is acquired; S3, if the element tag is a specific tag, it is judged whether the currently traversed non-root element contains a child element; S4, if the currently traversed non-root element contains a child element, the child element is taken as the currently traversed non-root element, and the step S2 is returned to be executed; S5, if the element tag is a non-specific tag, the next non-root element is continued, and the step S2 is returned to be executed; S6, if the currently traversed non-root element does not contain a child element, the currently traversed non-root element is taken as a specific element.

[0069] In step S1, starting from the root element of the XML document, the non-root elements are accessed one by one to be traversed. In the embodiment of the application, the specific implementation strategy of the traversal is not limited. For example, the traversal can be implemented by using a depth-first search algorithm, or can be implemented by using a breadth-first search algorithm. In the embodiment of the application, the process of traversing the non-root elements in the XML document to identify a plurality of specific elements is described in detail by taking the depth-first search algorithm as an example.

[0070] Then, step S2 is executed, for the currently traversed non-root element, the element tag contained by the currently traversed non-root element is acquired. In the XML document, the element tag is an identifier in the angle brackets (< >), and is used to mark the type and semantic meaning of the element. Different material types are distinguished by the name of the element tag; the specific position of the material is located by the hierarchical relationship of the element tag.

[0071] After the element tag contained by the currently traversed non-root element is acquired, it is judged whether the element tag is a specific tag. The specific tag can be a set of key tags defined in advance, representing the element tag that needs to be specially processed in the XML document, and is used to identify the material type that needs to be extracted.

[0072] Then, step S3 or S5 is executed. If step S5 is executed, that is, the element tag is a non-specific tag, the next non-root element is continued, and the step S2 is returned to be executed.

[0073] If step S3 is executed, that is, the element tag is a specific tag, it is continued to be judged whether the currently traversed non-root element contains a child element. The child element can be other elements nested in the currently traversed non-root element.

[0074] Then, step S4 or S6 is executed. If step S4 is executed, i.e. the current traversed non-root element contains sub-elements, the sub-elements are taken as the current traversed non-root element, and the execution returns to step S2. If step S6 is executed, i.e. the current traversed non-root element does not contain sub-elements, the current traversed non-root element is taken as a specific element. Step S4 is a recursive process when the non-root element contains sub-elements, and the sub-elements of the current traversed non-root element are taken as new starting points for traversal, and steps S2-S6 are re-executed for each sub-element. After the recursion ends, other non-root elements are traversed. Step S6 is a terminal collection when the non-root element does not contain sub-elements, and the current traversed non-root element is taken as a specific element.

[0075] In an optional embodiment, the rendering parameter information corresponding to each of the plurality of material types is extracted from the information segment corresponding to each of the plurality of material types, including: for each information segment, extracting path information of a material file and at least one attribute value of the material file from the information segment, the attribute value being used for rendering the material file; reading the material file according to the path information, and determining the material type described by the information segment according to the extension name of the material file; taking at least one attribute to which the at least one attribute value belongs as at least one rendering parameter, and taking the at least one attribute value as a default parameter value of the at least one rendering parameter, to obtain the rendering parameter information corresponding to the material type described by the information segment.

[0076] The information segment is a part related to the material type extracted from the XML document, and includes path information of a material file and at least one attribute value of the material file. The material file can be used to construct various media files of a final generated video. In the embodiments of the present application, the material file can include but is not limited to: a file in a format such as MP4 or AVI, containing dynamic images and audio, and used to show a video file of a series of continuous pictures; a file in a format such as WAV or MP3, providing audio files of background music, narration or special effect sound, etc.

[0077] The at least one attribute value of the material file extracted from the information segment is a specific value corresponding to the attribute, to describe the specific state of the attribute. The attribute can describe the rendering effect of the material file. For example, for a material file being a video file, the attribute can be video length, video speed, filter effect, etc., and if the attribute is video length, the attribute value can be 2 minutes, 1 hour, 1 day, etc.

[0078] In an optional embodiment, the rendering parameter information corresponding to the material type described by the information segment includes at least one rendering parameter and a default parameter value of the at least one rendering parameter. The default parameter value can be an attribute value of an attribute corresponding to the rendering parameter before quantization. In this case, at least one attribute of the at least one attribute value of the extracted material file in the information segment can be taken as the at least one rendering parameter, and the at least one attribute value can be taken as the default parameter value of the at least one rendering parameter to obtain the rendering parameter information. That is, the rendering parameter and the default parameter value form a key-value pair set, which is used to control the rendering logic of the material file in the video to achieve the expected rendering effect.

[0079] In an optional embodiment, the rendering parameter information corresponding to each of the plurality of material types is templated to obtain a plurality of materialization templates, including: for each material type, selecting at least one variable parameter from the rendering parameter information corresponding to the material type; associating a plurality of candidate parameter values with the at least one variable parameter; and generating a materialization template corresponding to the material type according to the plurality of candidate parameter values associated with the at least one variable parameter. The variable parameter is obtained by variable processing on the rendering parameter with the default parameter value, each variable parameter is associated with a plurality of candidate parameter values, and different candidate parameter values correspond to different rendering logics to produce different rendering effects in the generated video, which can be used to generate diversified materialization templates. For each material type, at least one variable parameter is selected from the rendering parameter information corresponding to the material type, and these variable parameters can be changed within an adjustment range to obtain different materialization templates. For example, when the material file is a video material, the playback speed and the filter effect can be selected as the variable parameters. The variable parameter is associated with a plurality of candidate values to limit the adjustment range of the variable parameter, so as to ensure that the material file of the generated video meets the design requirements and avoid invalid settings.

[0080] In an optional embodiment, when selecting at least one variable parameter from the rendering parameter information corresponding to the material type, two selection methods are provided: one is to take all the rendering parameters as variable parameters, and the other is to select part of the rendering parameters as variable parameters according to the weight values. When all the rendering parameters are selected as variable parameters, there is no need to compare the weight values, and all the rendering parameters can be adjusted as variable parameters.

[0081] In an optional embodiment, the way of selecting the part of the rendering parameters as the variable parameters according to the weight values can be that at least one variable parameter is selected from the rendering parameter information corresponding to the material type, including: configuring the weight values of the rendering parameters in advance for the material type; and parsing the rendering parameters from the rendering parameter information corresponding to the material type, and selecting at least one rendering parameter with a weight value greater than a set weight threshold as at least one variable parameter. The weight value can be an importance score assigned to each rendering parameter, which is used to quantify the influence of the rendering parameter on the rendering effect of the material file in video generation, and at least one rendering parameter with a greater influence on user perception or application target can be selected as at least one variable parameter. The weight threshold is a pre-defined critical value, and the user can select the parameters with weight values higher than the weight threshold as variable parameters, so as to avoid selecting too many irrelevant rendering parameters as variable parameters, so as to prevent configuration conflicts or rendering logic confusion.

[0082] In the embodiments of the present application, the way of pre-configuring the weight values of the rendering parameters is not limited. For example, the weight values of the rendering parameters can be annotated according to experience, or automatically generated through user behavior data analysis.

[0083] Optionally, at least one variable parameter can also be randomly selected according to the number of set variable parameters from the rendering parameter information corresponding to the material type. In an optional embodiment, the materialized template corresponding to the material type is generated according to the at least one variable parameter associated with a plurality of candidate parameter values, including: adding the rendering parameters corresponding to the material type and the default parameter values of the rendering parameters to a preset template file, and adding the plurality of candidate parameter values associated with the at least one variable parameter to the preset template file; and adding a placeholder for carrying the material file corresponding to the material type in the preset template file to obtain the materialized template corresponding to the material type. The preset template file is a basic video template containing basic configurations, which contains the basic structure of the video template and some preset rendering parameters. Adding the rendering parameters corresponding to the material type and the default parameter values of the rendering parameters to the preset template file can ensure that the finally obtained materialized template has complete rendering parameter information.

[0084] In the embodiments, on the basis of adding the rendering parameters corresponding to the material type and the default parameter values of the rendering parameters to the preset template file, the plurality of candidate parameter values associated with the at least one variable parameter are added to the preset template file. The candidate parameter values provide multiple choices, allowing the user or the system to select different candidate parameter values according to the needs.

[0085] In an optional embodiment, on the basis of adding at least one candidate parameter value associated with a variable parameter in the preset template file, a placeholder for carrying a material file corresponding to a material type can be added in the preset template file to obtain a materialized template corresponding to the material type. The placeholder is a reserved position for filling the material file and supports dynamic replacement, for example, the path information of an actual material file corresponding to the material type can be filled. Through the placeholder, different material files can be flexibly replaced without modifying the template structure.

[0086] In the above embodiment, the materialized template integrates the rendering parameters of the material type and the default parameter values thereof into the preset base template, ensuring the completeness of the rendering parameter information of the generated materialized template. Meanwhile, by associating multiple candidate parameter values with the variable parameters, the dynamic selection of the parameter values of the variable parameters is realized to flexibly adjust the video style. The embedding of the placeholder further decouples the parameter configuration of the initial video template and the material file, supports dynamically replacing the material file without modifying the template structure, and thus allows assigning the corresponding candidate parameter values to the variable parameters according to the application requirements to obtain the freely combined materialized template. Finally, the flexibility and diversity of the video template generation are significantly improved, and the video content with different styles can be efficiently mass-produced, thereby solving the problems of complicated parameter configuration, complex adaptation process, and serious homogenization of generated video content in traditional video production.

[0087] On the basis of obtaining the multiple materialized templates, batch video generation or single video generation can be performed based on the multiple materialized templates. For the batch video generation scenario, the multiple materialized templates provided in the embodiments of the present application can be used to generate multiple videos with diverse styles and differentiated content. The following describes an embodiment of batch video generation based on the multiple materialized templates provided in the embodiments of the present application.

[0088] In the embodiments of the present application, the initial video template is subjected to variable processing to obtain multiple materialized templates. In batch video generation, for each video instance identifier, at least one target materialized template is determined from the multiple materialized templates according to the material types in the set of video materials corresponding to the video instance identifier.

[0089] In an optional embodiment, when determining the at least one target materialization template from the plurality of materialization templates according to the material type of the video instance identifier corresponding set of video materials, the method comprises: identifying the at least one material type contained in the video instance identifier corresponding set of video materials; selecting the at least one initial materialization template from the plurality of materialization templates according to the at least one material type contained in the set of video materials, each material type corresponding to an initial materialization template; and performing parameter adjustment on at least part of the selected at least one initial materialization template to obtain the at least one target materialization template.

[0090] In the embodiment, the plurality of materialization templates obtained by variable processing of the initial video template are referred to as initial materialization templates. In the case where at least one initial materialization template corresponding to a certain group is determined from the plurality of initial materialization templates, the rendering rule described by the initial materialization template can be referred to as initial rendering rule. Further, parameter adjustment is performed on at least part of the at least one initial materialization template to obtain at least one target materialization template. Each initial materialization template obtains a corresponding target materialization template after parameter adjustment. The target materialization template after parameter adjustment contains target rendering rule, which is different from the initial rendering rule described by the initial materialization template.

[0091] In the embodiment, since the initial materialization template is obtained by variable processing, the variable processing can convert the fixed rendering parameter information in the initial video template into variable parameters that can be dynamically assigned, so as to realize flexible configuration of template content. The values of optional parameters are different, and the rendering rules are different. Optionally, each initial materialization template includes at least one variable parameter, and each variable parameter is associated with a plurality of candidate parameter values and a default parameter value.

[0092] In the embodiment, the parameter adjustment on the initial materialization template is to assign different candidate parameter values to the variable parameters of the initial materialization template to obtain a plurality of target materialization templates with different candidate parameter values. The target materialization template for a certain material type can be used for video material rendering of the group of different video instance identifiers. The parameter values of the variable parameters are different, which means that the rendering rules described by the target materialization templates of the same material type are different, so that the rendering results of the batch-generated videos are different, different videos form differentiation, and the richness of video content is improved. How to perform parameter adjustment on the initial materialization template will be introduced below.

[0093] In an optional embodiment, when the at least one initial materialized template is parameter-adjusted to obtain the at least one target materialized template, the method comprises: determining the number of templates to be adjusted, the number of templates being less than or equal to the number of the at least one initial materialized template; selecting the initial materialized template to be adjusted from the at least one initial materialized template according to the number of templates to be adjusted; determining the variable parameter to be adjusted from the initial materialized template to be adjusted; randomly determining the target parameter value from a plurality of candidate parameter values associated with the variable parameter to be adjusted; and assigning the target parameter value to the variable parameter to be adjusted to obtain the target materialized template.

[0094] In the embodiment, the manner of determining the number of templates to be adjusted is not limited. For example, the parameter adjustment can be performed on each initial video template, and the number of templates to be adjusted is equal to the number of the at least one initial materialized template. For another example, a random integer in a preset range can be generated according to a random number generation algorithm as the number of templates to be adjusted, and the preset range refers to a number less than or equal to the number of the at least one initial materialized template. The random number generation algorithm is not limited, including but not limited to linear congruential generator (LCG) and Mersenne Twister.

[0095] Further, the initial materialized template to be adjusted is selected from the at least one initial materialized template according to the number of templates to be adjusted. In an optional embodiment, the initial materialized template to be adjusted is selected according to the priority of the material type and the number of templates to be adjusted. The priority of the material type refers to the importance of the material type to the presentation effect of the generated video. For example, the importance of the subtitle material type to the presentation effect is generally low, and thus the priority of the subtitle material type can be set to a low priority. In contrast, the priority of the background picture can be higher than that of the subtitle, and thus can be set to a medium priority. The importance of the digital person to the presentation effect is high, and thus the priority of the digital person can be set to a high priority. Further, the initial materialized template to be adjusted can be selected from the initial materialized template of the material type with a high priority according to the priority. If the number of the initial materialized template is less than the number of templates to be adjusted, the initial materialized template of the material type with a low priority can be selected, and the number of templates to be adjusted is less than or equal to the number of the at least one initial materialized template.

[0096] Further, from the initial materialized template to be adjusted, a variable parameter to be adjusted is determined. In an optional embodiment, all variable parameters in the initial materialized template to be adjusted can be taken as the variable parameter to be adjusted. In another optional embodiment, a target variable parameter in the initial materialized template to be adjusted is taken as the variable parameter to be adjusted, and the target variable parameter is a pre-selected optional parameter.

[0097] Further, a target parameter value is randomly determined from a plurality of candidate parameter values associated with the variable parameter to be adjusted. In some embodiments, the N video instance identifiers each correspond to at least one initial materialized template having the same variable parameter to be adjusted, and the plurality of candidate parameter values associated with the variable parameter to be adjusted can be randomly taken as the target parameter values of the N video instance identifiers respectively, so that the optional parameters of the N video instance identifiers take values as different as possible.

[0098] Further, the target parameter value is assigned to the variable parameter to be adjusted to obtain a target materialized template.

[0099] In the case of obtaining the target materialized template, the target materialized template is combined to obtain a target video template. As for the combination manner, the present embodiment is not limited. Two combination manners are provided below, but are not limited thereto.

[0100] In an optional embodiment, according to the hierarchical relationship between the plurality of video materials included in the initial video template, a basic video template is generated, which is taken as a framework of the target video template and includes a plurality of blank structure positions corresponding to the plurality of video materials. The blank structure position refers to a placeholder of the target materialized template preset in the basic video template, which is used to identify the position where the target materialized template can be inserted, and each blank structure position corresponds to the filling of the target materialized template of one material type. The positional relationship between the plurality of blank structure positions reflects the hierarchical relationship between the plurality of video materials, and as described in the above embodiments, the hierarchical relationship represents the front-to-back superimposition order of the plurality of materials in the generated video. In some embodiments, the hierarchical relationship is extracted from the initial video template; or, it can also be preset based on at least one target materialized template, that is, the hierarchical relationship can be set on demand. For example, the structure position corresponding to the subtitle is located in the upper layer, the structure position corresponding to the digital person is located in the middle layer, and the structure position corresponding to the background picture is located in the bottom layer. Further, at least one target materialized template is inserted into the corresponding blank structure position of the basic video template to obtain the target video template corresponding to the video instance identifier.

[0101] In another optional embodiment, a structure bit in which the rendering rule of the video material of the same material type in the initial video template is located is overwritten according to at least one target materialization template, to obtain a target video template corresponding to the video instance identifier. The difference between the structure bit and the blank structure bit is that the rendering rule of each video material in the initial video template occupies a structure bit, and the blank structure bit is empty. The positional relationship between the structure bits reflects the hierarchical relationship between the multiple video materials.

[0102] In the case of obtaining the target video template, video generation processing is performed according to the target video template corresponding to each video instance identifier and a set of video materials. In an optional embodiment, in the case of performing video generation processing according to the target video template corresponding to each of the N video instance identifiers and a set of video materials to obtain N videos under the video category, the method comprises: for each video instance identifier, filling the set of video materials corresponding to the video instance identifier into the target materialization template in the target video template corresponding to the video instance identifier; for the filled target video template, rendering the set of video materials according to the rendering rule and hierarchical relationship of the set of video materials described in the filled target video template, to obtain a video under the video category.

[0103] In the target materialization template, at least one placeholder corresponding to the material type is included, which is used to fill the video material of the material type.

[0104] Figure 2 The structural diagram of the video batch generation system provided by an exemplary embodiment of the present application is provided. The video batch generation system includes an endpoint, a content distribution network and multiple material generation models. In this embodiment, the multiple material generation models include generative language models, text-to-speech models, multi-modal models and text-to-image models, but are not limited thereto.

[0105] In this embodiment, the endpoint is used for video material generation, management of batch video generation tasks and control of the video generation process.

[0106] The endpoint includes front-end code and back-end code, that is, a client-server structure as described in the above embodiment. The front-end code refers to a video generation page built based on a front-end framework, which runs on the client and is used to interact with the user, receive the input operation of the user, and initiate a batch video generation task to the server. The back-end code of the endpoint runs on the server, and the back-end code is obtained by API (Application Programming Interface) based on the code of the existing video editing software. The existing video editing software can be, for example, shortcut. Among the existing video editing software, the front-end UI code is highly coupled with the video rendering function code. In this embodiment, the front-end UI of the video editing software is decoupled from the video rendering function code according to the function, and the video rendering function code of the existing video editing software is encapsulated as an independent API, so that the video rendering function can be called by an external standard API to realize automatic video generation.

[0107] As the endpoint in Figure 2 , in an example, the front-end code of the endpoint can be a video generation page built based on Astro, which is responsible for receiving callback notifications, such as when the material generation model is completed, a notification can be sent to the endpoint to inform that the subsequent process of video generation can continue. The back-end code of the endpoint can be obtained by API based on the video rendering code of the video editing software, so as to be exposed to the outside through the API interface for external calling.

[0108] In this embodiment, as shown in Figure 2 , in response to the input operation on the video generation page for the number of videos and the video category, a batch video generation task is generated, which includes the number of videos N and the video category. N is an integer greater than or equal to 2. For detailed content of the batch generation task, please refer to the above embodiment.

[0109] In this embodiment, according to the batch video generation task, N video instance identifiers are generated. For the implementation of generating N video instance identifiers, please refer to the above embodiment, which will not be repeated here. Among them, the N video instance identifiers are different from each other, and each video instance identifier can correspond to the generation of a video. The video instance identifier can also be used as the unique identity of the video in the generation process.

[0110] In this embodiment, a set of video material description information related to the video category is generated for each video instance identifier, and the set of video material description information is used to generate the video material corresponding to the video instance identifier. For the generation method of the description information, please refer to the related content of the above embodiment.

[0111] Each set of video materials includes description information of video materials of multiple material types. In this embodiment, the description information of each set of video materials includes subtitle description information, audio type description information, digital person description information, and background picture description information, but is not limited thereto.

[0112] Further, for each video instance identifier, a plurality of material generation models based on artificial intelligence are called according to the video instance identifier corresponding set of video material description information, a plurality of video materials corresponding to the video instance identifier are generated, and the plurality of video materials corresponding to the video instance identifier are synchronously uploaded to the content distribution network.

[0113] Among them, one material generation model can generate at least part of the plurality of video materials. For example, one material generation model can generate one kind of video material. The following is also described as an example, but is not limited thereto.

[0114] In an optional embodiment, when a plurality of material generation models based on artificial intelligence are called according to the video instance identifier corresponding set of video material description information, a plurality of video materials corresponding to each video instance identifier are generated, including: according to the subtitle description information, calling the generative language model to generate the text information to obtain the target subtitle; according to the target subtitle and the audio type description information, calling the text-to-speech model to convert the target subtitle into the target audio adapted to the audio type description information; according to the target audio and the digital person description information, calling the multi-modal model to select the target digital person according to the digital person description information, and generating the green screen video based on the target audio and the target digital person to obtain the green screen video of the target digital person; according to the background picture description information, calling the text-to-image model to generate the background picture to obtain the target background picture.

[0115] In an optional embodiment, when the generative language model is called to generate the text information according to the subtitle description information to obtain the target subtitle, it includes: according to the video category matching corresponding keywords, calling a pre-designed prompt word template, filling the keywords of the video category into the prompt word template to obtain the prompt words of the video category; inputting the prompt words of the video category into the generative language model to generate the text information to obtain the target subtitle related to the video category.

[0116] Among them, the keywords are used to describe the expression theme of the video category, different video categories can correspond to different keywords, and different video types can correspond to different prompt word templates. Among them, the target subtitle includes all the text content required by each video.

[0117] Further optionally, according to the target subtitle and the audio type description information, a text-to-speech model is called to convert the target subtitle into target audio that is adapted to the audio type description information. The target subtitle is used to provide text content, and the audio type description information is used to specify the audio type of the generated target audio, which includes but is not limited to audio format, audio language, audio tone, and audio quality, etc. Any audio type that can be used to specify the sound effect of the target audio is applicable to this embodiment.

[0118] Further, a multi-modal model is called to generate a green screen video of the digital human. The digital human refers to a virtual character generated based on AI technology, which can simulate the appearance, voice, mouth shape, and other behaviors of a real person in synchronization, and can be used in scenarios such as intelligent customer service, short video production, virtual anchor, etc., but is not limited thereto. In this embodiment, in combination with the corresponding multiple video materials identified by each video instance identifier, a video with real person speaking effect can be generated.

[0119] In an optional embodiment, the digital human can be a pre-recorded real person video or picture. In subsequent embodiments, the real person video and picture are collectively referred to as video frames, and the number of video frames can be one or more. In this case, the description information of different digital humans can be implemented as identification information, which serves as the unique identity of the digital human and is used to obtain the video frames of the digital human corresponding to the identification information. In another optional embodiment, the video frames of the digital human can be dynamically generated based on the multi-modal model. In this case, the description information of the digital human can be a prompt word used to generate the digital human, and the description information of the corresponding digital human for each video instance identifier can be different.

[0120] Further optionally, in the step of calling the multi-modal model according to the target audio and the digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain a green screen video of the target digital human, the method further includes: obtaining the video frames of the corresponding target digital human based on the digital human description information; calling the multi-modal model to perform multi-dimensional feature extraction on the target audio according to the video frames of the target digital human and the target audio, to obtain multi-dimensional voice features; wherein the multi-dimensional voice features include but are not limited to voice content features and voice emotion features; determining the mouth shape control parameters of the target digital human according to the voice content features; determining the facial expression control parameters of the target digital human according to the voice emotion features; determining the body movement control parameters of the target digital human according to the voice content features and the voice emotion features; generating the green screen video of the target digital human based on the mouth shape control parameters, the facial expression control parameters, and the body movement control parameters of the target digital human; and the mouth shape, facial expression, and body movement of the green screen video of the target digital human match the target audio.

[0121] In the embodiment, the speech content features are used to reflect semantic information in the target audio, such as lexical content, syntactic structure, speech intention and speech rhythm, mainly for driving the synchronization of the digital person's mouth shape and the semantics of the target audio. The speech emotion features are used to reflect the emotional state of the target audio, including but not limited to tone, speed and intonation, mainly for driving the facial expression and body movement of the digital person.

[0122] Further, according to the speech content features, the mouth shape control parameters of the target digital person are determined; according to the speech emotion features, the facial expression control parameters of the target digital person are determined; according to the speech content features and the speech emotion features, the body movement control parameters of the target digital person are determined.

[0123] In the case of obtaining the mouth shape control parameters, facial expression control parameters and body movement control parameters, based on the mouth shape control parameters, expression control parameters and body movement control parameters of the target digital person, the target digital person is driven to perform action rendering, and the corresponding green screen video of the target digital person is generated, which is convenient for subsequent flexible replacement using the background image. Since the green screen video of the target digital person is generated based on the multi-dimensional speech feature control of the target audio content, it is ensured that the dynamic performance of the mouth shape, facial expression and body movement of the target digital person in the green screen video is matched with the target audio in terms of semantics, timing and emotion.

[0124] Further, in the embodiment, the target audio and the target subtitle are aligned to ensure consistency on the time axis. For example, the target subtitle is timestamped and aligned by the timestamp of the speech content in the target audio to ensure that the target subtitle display is completely synchronized with the target audio.

[0125] In the embodiment, the generation of the background image is one of generating the background image according to the background image description information, calling the text image model to generate the background image, and obtaining the target background image. The other is to directly obtain the pre-generated or photographed background image, which is not limited.

[0126] In an optional embodiment, a corresponding set of video material generation states is maintained for each of the N video instance identifiers. The generation state of any video material of any material type in each set of video materials includes: a ready-to-generate state, a generating state, a generation success state and a generation failure state. For example, for any video instance identifier, the multiple video materials that need to be generated include: target subtitles, target audio, target digital person green screen video and target background image. Then the generation state of each can be maintained respectively for the target subtitles, target audio, target digital person green screen video and target background image.

[0127] Specifically, for each video instance identifier, the various video materials corresponding to that video instance identifier are marked as ready for generation. When calling multiple AI-based material generation models to generate the various video materials corresponding to that video instance identifier, if any material generation model returns a "generating in progress" response message, the video materials that the material generation model should generate are updated to the "generating in progress" state; if any material generation model returns a "generating successfully" message, the video materials that the material generation model should generate are updated to the "generating successfully" state; if any material generation model returns a "generating failed" message, the video materials that the material generation model should generate are updated to the "generating failed" state. For example... Figure 2 As shown, a subscription service is provided that can notify each material generation model of the generation status of the video material for that model, such as a successful generation status. Furthermore, the subscription service returns a successful generation message to the endpoint to inform it that the corresponding video material has been successfully generated. Figure 2 This example only uses the notification process for the target digital person as an example, but is not limited to this.

[0128] Optionally, if the generation status of the video footage is updated to a generation failure status, a failure notification message is output to the user who initiated the input operation. The failure notification message includes the footage type of the video footage that has been updated to a generation failure status and its corresponding video instance identifier. If a regeneration operation triggered by the user is received, the corresponding description information of the video footage is obtained based on the video instance identifier, and the corresponding AI-based footage generation model is invoked to regenerate the corresponding video footage.

[0129] In this embodiment, when multiple video materials corresponding to each video instance identifier are generated, these multiple video materials can be simultaneously uploaded to the content delivery network (CDN), and access links for the video materials corresponding to each video instance identifier can be obtained within the CDN. Continuing with the example above, for any video instance identifier, if the multiple video materials corresponding to that video instance identifier include target subtitles, target audio, target digital human green screen video, and target background image, the target subtitles, target audio, target digital human green screen video, and target background image are uploaded to the CDN respectively, and access links for the target subtitles, target audio, target digital human green screen video, and target background image are obtained.

[0130] Further, in response to the batch video generation trigger event, according to the N video instance identifiers respectively corresponding to the plurality of video materials, access links of the plurality of video materials in the content distribution network are obtained, and the N video instance identifiers respectively corresponding to the plurality of video materials are obtained; and video generation is performed according to the N video instance identifiers respectively corresponding to the plurality of video materials and the target video template, to obtain the N videos under the video category. When the video generation is performed according to the N video instance identifiers respectively corresponding to the plurality of video materials and the target video template, the N video instance identifiers respectively corresponding to the plurality of video materials can be filled into the respective target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain the N videos under the video category. Optionally, the rendering can be performed based on the API corresponding to the video rendering function exposed by the above-mentioned endpoint to the video generation page.

[0131] In the above example, in response to the batch video generation trigger event, according to the N video instance identifiers respectively corresponding to the target subtitles, the target audio, the green screen video of the target digital person, and the target background image, access links of the target subtitles, the target audio, the green screen video of the target digital person, and the target background image in the content distribution network are obtained, and the N video instance identifiers respectively corresponding to the target subtitles, the target audio, the green screen video of the target digital person, and the target background image are obtained; and video generation is performed according to the N video instance identifiers respectively corresponding to the target subtitles, the target audio, the green screen video of the target digital person, the target background image, and the target video template, to obtain the N videos under the video category.

[0132] Further optionally, the N video instance identifiers respectively corresponding to the videos and the target video template are uploaded to the content distribution network, and access links of the N video instance identifiers respectively corresponding to the videos and the target video template in the content distribution network are obtained, and the access links are added to the video generation result page; in response to a viewing operation of the video result page, the video generation result page is displayed, and the video generation result page includes access links of a group of video materials, videos, and a target video template corresponding to at least one video instance identifier in the N video instance identifiers in the content distribution network.

[0133] In this optional embodiment, in response to a trigger operation on the access links of the group of video materials, the videos, and / or the target video template corresponding to at least one video instance identifier in the content distribution network, the group of video materials, the videos, and / or the target video template corresponding to the at least one video instance identifier information are accessed.

[0134] Figure 3 A flowchart of a video batch generation method provided by an exemplary embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the method includes the following steps. Figure 3

[0135] ​S301: In response to an input operation on the video generation page for the number of videos and the video category, a batch video generation task is generated, the batch video generation task including the number of videos N and the video category, N being an integer greater than or equal to 2;

[0136] S302: According to the batch video generation task, N video instance identifiers are generated, and for each video instance identifier, a set of video material description information related to the video category is obtained, each set of video material description information including description information of multiple video materials;

[0137] S303: For each video instance identifier, according to the set of video material description information corresponding to the video instance identifier, a plurality of material generation models based on artificial intelligence are called to generate a plurality of video materials corresponding to the video instance identifier, and the plurality of video materials corresponding to the video instance identifier are synchronously uploaded to the content distribution network;

[0138] S304: According to the plurality of video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined, the target video template corresponding to each video instance identifier being used to describe the rendering rule and hierarchical relationship of the plurality of video materials corresponding to the video instance identifier;

[0139] S305: In response to a batch video generation trigger event, according to N video instance identifiers, a plurality of video materials corresponding to N video instance identifiers are obtained from the content distribution network; according to the plurality of video materials corresponding to N video instance identifiers and the target video template, video generation processing is performed to obtain N videos under the video category.

[0140] In an optional embodiment, when determining the target video template corresponding to each video instance identifier according to the plurality of video materials corresponding to each video instance identifier, a plurality of materialized templates obtained by variable processing of an initial video template are obtained, the initial video template including the rendering rule of the plurality of video materials required for generating videos and the hierarchical relationship between the plurality of video materials, each materialized template being used to describe the rendering rule of one video material; for each video instance identifier, at least one target materialized template is determined from the plurality of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; based on the hierarchical relationship between the plurality of video materials included in the initial video template, the at least one target materialized template is combined to obtain the target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rule and hierarchical relationship of the set of video materials.

[0141] In an optional embodiment, in the variable processing of the initial video template, the initial video template is parsed according to the material type as a split variable to obtain information segments corresponding to a plurality of material types; the rendering parameter information corresponding to the plurality of material types is respectively extracted from the information segments corresponding to the plurality of material types; and the rendering parameter information corresponding to the plurality of material types is templated to obtain a plurality of materialization templates.

[0142] In an optional embodiment, in the determination of at least one target materialization template from the plurality of materialization templates according to the material types in the video instance-identified set of video materials, at least one material type contained in the video instance-identified set of video materials is identified; at least one initial materialization template is selected from the plurality of materialization templates according to the at least one material type contained in the set of video materials, with each material type corresponding to one initial materialization template; and at least part of the initial materialization templates are parameter-adjusted to obtain the at least one target materialization template.

[0143] In an optional embodiment, in the combination of the at least one target materialization template based on the hierarchical relationship between the plurality of video materials included in the initial video template to obtain the target video template corresponding to the video instance-identified set, a base video template is generated according to the hierarchical relationship between the plurality of video materials included in the initial video template, the base video template including a plurality of blank structure positions corresponding to the plurality of video materials; the positional relationship between the plurality of blank structure positions reflects the hierarchical relationship between the plurality of video materials; the at least one target materialization template is respectively inserted into the corresponding blank structure position of the base video template to obtain the target video template corresponding to the video instance-identified set; or, according to the at least one target materialization template, the structure position in which the rendering rule of the video material of the same material type in the initial video template is located is covered to obtain the target video template corresponding to the video instance-identified set; wherein the rendering rule of each video material in the initial video template occupies one structure position, and the positional relationship between the structure positions reflects the hierarchical relationship between the plurality of video materials.

[0144] In an optional embodiment, each set of video material description information includes: subtitle description information, audio type description information, digital person description information, and background picture description information; and when a plurality of material generation models based on artificial intelligence are called to generate a plurality of video materials corresponding to the video instance identifier according to the corresponding set of video material description information of the video instance identifier, the method comprises: calling a generative language model to generate text information according to the subtitle description information to obtain target subtitles; calling a text-to-speech model to convert the target subtitles into target audio that matches the audio type description information according to the target subtitles and the audio type description information; calling a multi-modal model to select a target digital person according to the digital person description information and generate a green screen video of the target digital person based on the target audio and the target digital person to obtain a green screen video of the target digital person according to the target audio and the digital person description information; and calling a text-to-picture model to generate a background picture to obtain a target background picture according to the background picture description information.

[0145] In an optional embodiment, when a multi-modal model is called to select a target digital person according to the digital person description information and generate a green screen video of the target digital person based on the target audio and the target digital person to obtain a digital person green screen video according to the target audio and the digital person description information, the method comprises: obtaining a video frame of the target digital person based on the digital person description information; calling a multi-modal model to extract multi-dimensional features of the target audio based on the video frame of the target digital person and the target audio to obtain multi-dimensional speech features, wherein the multi-dimensional speech features include: speech content features, speech emotion features; determining mouth shape control parameters of the target digital person based on the speech content features; determining facial expression control parameters of the target digital person based on the speech emotion features; determining limb movement control parameters of the target digital person based on the speech content features and the speech emotion features; and generating a green screen video of the target digital person based on the mouth shape control parameters, the facial expression control parameters, and the limb movement control parameters of the target digital person; and the mouth shape, facial expression, and limb movement of the green screen video of the target digital person match the target audio.

[0146] In an optional embodiment, after calling a text-to-speech model to convert the target subtitles into target audio that matches the audio type description information according to the target subtitles and the audio type description information, the method further comprises: aligning the target audio and the target subtitles to synchronize the target audio and the target audio in time.

[0147] In an optional embodiment, the method further comprises: identifying, for each of the N video instance identifiers, a generation state of a corresponding set of video materials, the generation state of any video material of any type in each set of video materials comprising: a ready-to-generate state, a generating state, a generation success state, and a generation failure state; for each video instance identifier, marking the corresponding plurality of video materials of the video instance identifier as being in the ready-to-generate state; when invoking the plurality of material generation models based on artificial intelligence to generate the plurality of video materials corresponding to the video instance identifier, if any material generation model returns a generating response message, updating the video material to be generated by the material generation model to the generating state; if any material generation model returns a generation success message, updating the video material to be generated by the material generation model to the generation success state; if any material generation model returns a generation failure message, updating the video material to be generated by the material generation model to the generation failure state; if the generation state of the video material is updated to the generation failure state, outputting a failure reminder message to the user who initiated the input operation, the failure reminder message comprising the material type of the video material and the corresponding video instance identifier of the video material; if a user-triggered re-generation operation is received, obtaining the description information of the video material according to the corresponding video instance identifier of the video material, and invoking the corresponding material generation model based on artificial intelligence to regenerate the video material.

[0148] In an optional embodiment, the method further comprises: uploading the videos corresponding to the N video instance identifiers and the target video template to a content distribution network, and obtaining access links of the videos corresponding to the N video instance identifiers and the target video template in the content distribution network, and adding the access links to a video generation result page; in response to a viewing operation of the video result page, displaying the video generation result page, the video generation result page comprising access links of a set of video materials, a video, and a target video template corresponding to at least one video instance identifier in the N video instance identifiers in the content distribution network; in response to a triggering operation of the access links of the set of video materials, the video, and / or the target video template corresponding to at least one video instance identifier in the content distribution network, accessing the set of video materials, the video, and / or the target video template corresponding to the at least one video instance identifier.

[0149] The detailed implementation and beneficial effects of each step in the method of the present embodiment have been described in detail in the foregoing embodiments, and will not be described in detail here.

[0150] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 301 to 303 can be device A; or the execution subject of steps 301 and 302 can be device A, and the execution subject of step 303 can be device B; and so on.

[0151] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 301, 302, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0152] Figure 4 This is a schematic diagram of the structure of an electronic device provided as another exemplary embodiment of this application. For example... Figure 4 As shown, the device includes a memory 44 and a processor 45.

[0153] Memory 44 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0154] The processor 45 is coupled with the memory 44 and is configured to execute a computer program in the memory 44 to: in response to an input operation on the video generation page for a video quantity and a video category, generate a batch video generation task, the batch video generation task including the video quantity N and the video category, N being an integer greater than or equal to 2; according to the batch video generation task, generate N video instance identifiers, and obtain a set of video material description information related to the video category for each video instance identifier, each set of video material description information including description information of multiple video materials; for each video instance identifier, according to the set of video material description information corresponding to the video instance identifier, call multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously upload the multiple video materials corresponding to the video instance identifier to a content distribution network; according to the multiple video materials corresponding to each video instance identifier, determine a target video template corresponding to each video instance identifier, the target video template corresponding to each video instance identifier being used to describe a rendering rule and a hierarchical relationship of the multiple video materials corresponding to the video instance identifier; in response to a batch video generation trigger event, according to the N video instance identifiers, obtain the multiple video materials corresponding to the N video instance identifiers from the content distribution network respectively; and according to the multiple video materials corresponding to the N video instance identifiers and the target video templates, perform video generation processing to obtain N videos under the video category.

[0155] In an optional embodiment, when the processor 45 determines the target video template corresponding to each video instance identifier according to the multiple video materials corresponding to each video instance identifier, the processor 45 is specifically configured to: obtain multiple materialized templates obtained by performing variable processing on an initial video template, the initial video template including a rendering rule of multiple video materials required for generating a video and a hierarchical relationship between the multiple video materials, each materialized template being used to describe a rendering rule of a video material; for each video instance identifier, according to a material type in a set of video materials corresponding to the video instance identifier, determine at least one target materialized template from the multiple materialized templates; and based on the hierarchical relationship between the multiple video materials included in the initial video template, combine the at least one target materialized template to obtain the target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rule and the hierarchical relationship of the set of video materials.

[0156] In an optional embodiment, the processor 45, when performing the variable processing on the initial video template, is specifically configured to: parse the initial video template to obtain information segments corresponding to a plurality of material types, taking the material types as the split variables; extract rendering parameter information corresponding to the plurality of material types respectively from the information segments corresponding to the plurality of material types respectively; and template the rendering parameter information corresponding to the plurality of material types respectively to obtain the plurality of materialization templates.

[0157] In an optional embodiment, the processor 45, when determining at least one target materialization template from the plurality of materialization templates according to the material types in the set of video materials corresponding to the video instance identifier, is specifically configured to: identify at least one material type contained in the set of video materials corresponding to the video instance identifier; select at least one initial materialization template from the plurality of materialization templates according to the at least one material type contained in the set of video materials, each material type corresponding to one initial materialization template; and perform parameter adjustment on at least part of the at least one initial materialization template to obtain the at least one target materialization template.

[0158] In an optional embodiment, the processor 45, when combining the at least one target materialization template based on the hierarchical relationship between the plurality of video materials included in the initial video template to obtain the target video template corresponding to the video instance identifier, is specifically configured to: generate a basic video template according to the hierarchical relationship between the plurality of video materials included in the initial video template, the basic video template including a plurality of blank structure positions corresponding to the plurality of video materials; the positional relationship between the plurality of blank structure positions reflecting the hierarchical relationship between the plurality of video materials; insert the at least one target materialization template into the corresponding blank structure position of the basic video template respectively to obtain the target video template corresponding to the video instance identifier; or, according to the at least one target materialization template, cover the structure position where the rendering rule of the video material of the same material type in the initial video template is located to obtain the target video template corresponding to the video instance identifier; wherein the rendering rule of each video material in the initial video template occupies one structure position, and the positional relationship between the structure positions reflects the hierarchical relationship between the plurality of video materials.

[0159] In an optional embodiment, each set of video material description information includes: subtitle description information, audio type description information, digital person description information, and background picture description information; and the processor 45 is specifically configured to: according to the video instance identifier, identify a corresponding set of video material description information, call a plurality of material generation models based on artificial intelligence, and generate a plurality of video materials corresponding to the video instance identifier; and according to the subtitle description information, call a generative language model to generate text information to obtain target subtitles; according to the target subtitles and the audio type description information, call a text-to-speech model to convert the target subtitles into target audio that matches the audio type description information; according to the target audio and the digital person description information, call a multi-modal model to select a target digital person according to the digital person description information, and generate a green screen video based on the target audio and the target digital person to obtain a green screen video of the target digital person; and according to the background picture description information, call a text-to-picture model to generate a background picture to obtain a target background picture.

[0160] In an optional embodiment, the processor 45 is specifically configured to: according to the target audio and the digital person description information, call a multi-modal model to select a target digital person according to the digital person description information, and generate a green screen video based on the target audio and the target digital person to obtain a digital person green screen video; and based on the digital person description information, obtain a video frame of the target digital person; according to the video frame of the target digital person and the target audio, call a multi-modal model to perform multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features, the multi-dimensional speech features including: speech content features, speech emotion features; according to the speech content features, determine mouth shape control parameters of the target digital person; according to the speech emotion features, determine facial expression control parameters of the target digital person; according to the speech content features and the speech emotion features, determine limb movement control parameters of the target digital person; based on the mouth shape control parameters, the facial expression control parameters, and the limb movement control parameters of the target digital person, generate a green screen video of the target digital person; and the mouth shape, facial expression, and limb movement of the green screen video of the target digital person match the target audio.

[0161] In an optional embodiment, after the processor 45 calls a text-to-speech model to convert the target subtitles into target audio that matches the audio type description information according to the target subtitles and the audio type description information, the processor 45 is further configured to: perform alignment processing on the target audio and the target subtitles to synchronize the target audio and the target audio in time.

[0162] In an optional embodiment, the processor 45 is further configured to: identify, for each of the N video instance identifiers, a generation state of a corresponding set of video materials, the generation state of the video material of any material type in each set of video materials including a preparation generation state, a generation in progress state, a generation success state and a generation failure state; for each video instance identifier, mark the corresponding multiple video materials of the video instance identifier as being in the preparation generation state; when invoking the multiple material generation models based on artificial intelligence to generate the multiple video materials corresponding to the video instance identifier, if any material generation model returns a generation in progress response message, update the video material to be generated by the material generation model to be in the generation in progress state; if any material generation model returns a generation success message, update the video material to be generated by the material generation model to be in the generation success state; if any material generation model returns a generation failure message, update the video material to be generated by the material generation model to be in the generation failure state; if the generation state of the video material is updated to the generation failure state, output a failure reminder message to the user who initiates the input operation, the failure reminder message including the material type of the video material and the corresponding video instance identifier of the video material; if a user-triggered re-generation operation is received, obtain the description information of the video material according to the corresponding video instance identifier of the video material, and invoke the corresponding material generation model based on artificial intelligence to regenerate the video material.

[0163] In an optional embodiment, the processor 45 is further configured to: upload the videos corresponding to the N video instance identifiers and the target video template to a content distribution network, and obtain access links of the videos corresponding to the N video instance identifiers and the target video template in the content distribution network, and add the access links to a video generation result page; in response to a viewing operation of the video result page, display the video generation result page, the video generation result page including the access links of a set of video materials, a video and a target video template corresponding to at least one video instance identifier in the N video instance identifiers in the content distribution network; in response to a trigger operation of the access links of the set of video materials, the video and / or the target video template corresponding to at least one video instance identifier in the content distribution network, access the set of video materials, the video and / or the target video template corresponding to the at least one video instance identifier.

[0164] Further, as shown in Figure 4 , the electronic device further includes a communication component 46, a display 47, a power supply component 48, an audio component 49 and other components. Figure 4 Some components are only schematically shown in the electronic device, and it does not mean that the electronic device only includes the components shown in Figure 4 . In addition, Figure 4The components in the dotted box are optional components, not mandatory components, and can be determined according to the product form of the working node. The electronic device of the embodiment can be implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone, or an IOT device, or a server device such as a general server, a cloud server, or a server array. If the electronic device of the embodiment is implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone, etc., it can include Figure 4 components in the dotted box; if the electronic device of the embodiment is implemented as a server device such as a general server, a cloud server, or a server array, it can not include Figure 4 components in the dotted box.

[0165] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0166] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, or the like mobile communication network, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0167] The above-mentioned display includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a swipe, and a gesture on the touch panel. The touch sensor can not only sense the boundary of a touch or swipe action, but also detect the duration and pressure associated with the touch or swipe operation.

[0168] The power component provides power to various components of the device in which the power component is located. The power component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which the power component is located.

[0169] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive external audio signals when the device in which the audio component is located is in an operational mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in a memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0170] Accordingly, the embodiments of the present application also provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is enabled to implement each step in the above method embodiments. The computer readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of the computer readable storage medium include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store computer program or instructions

[0171] Accordingly, the embodiments of the present application also provide a computer program product, the computer program product includes computer programs or instructions, when the computer programs or instructions are executed by a processor, the processor is enabled to implement each step in the above method embodiments. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by the computer programs or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor, or other programmable data processing devices can be implemented as a device that implements the corresponding functions in the above method embodiments.

[0172] It should also be noted that the terms "comprising," "including," or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0173] The above embodiments of the present application have been described only for clarity's sake and should not be considered limiting. Various changes and modifications can be made to the application by those skilled in the art. Any such changes or modifications are intended to fall within the scope of the application as defined by the appended claims, which are to be interpreted in the broadest sense and under the doctrine of equivalents.

Claims

1. A method for batch video generation, characterized in that, The method comprises the following steps: In response to an input operation on a video generation page for the number of videos and the video category, a batch video generation task is generated, the batch video generation task comprising a video number N and a video category, N being an integer greater than or equal to 2; According to the batch video generation task, N video instance identifiers are generated, and a set of video material description information related to the video category is obtained for each video instance identifier, each set of video material description information comprising the description information of a plurality of video materials; For each video instance identifier, a plurality of material generation models based on artificial intelligence are called according to the corresponding set of video material description information of the video instance identifier to generate a plurality of video materials corresponding to the video instance identifier, and the plurality of video materials corresponding to the video instance identifier are synchronously uploaded to a content distribution network; According to the plurality of video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined, and each target video template corresponding to each video instance identifier is used to describe the rendering rule and hierarchical relationship of the plurality of video materials corresponding to the video instance identifier; The target video template comprises a target materialized template selected from a plurality of materialized templates and adapted to the material type in the set of video materials, the plurality of materialized templates being obtained by varying an initial video template with the material type as a split variable, different material types corresponding to different materialized templates, and the materialized template being obtained by templateizing the rendering parameter information corresponding to the material type, the templateization being a process of converting the rendering parameter information from a fixed configuration default parameter value to a variable parameter set that can dynamically replace the parameter value; A materialized template is used to describe the rendering rule of a video material; In response to a batch video generation trigger event, a plurality of video materials corresponding to N video instance identifiers are obtained from the content distribution network according to the N video instance identifiers; and video generation processing is performed on the plurality of video materials corresponding to the N video instance identifiers and the target video template to obtain N videos under the video category.

2. The method of claim 1, wherein, According to the plurality of video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined, comprising: A plurality of materialized templates obtained by variable processing of an initial video template are obtained, the initial video template comprising the rendering rule of a plurality of video materials required for generating a video and the hierarchical relationship between the plurality of video materials, and each materialized template being used to describe the rendering rule of a video material; For each video instance identifier, at least one target materialized template is determined from the plurality of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; and the at least one target materialized template is combined based on the hierarchical relationship between the plurality of video materials included in the initial video template to obtain a target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rule and hierarchical relationship of the set of video materials.

3. The method of claim 2, wherein, The step of variable processing of the initial video template comprises: The initial video template is parsed according to the material type as a split variable to obtain information segments corresponding to a plurality of material types; From the information segments corresponding to the plurality of material types respectively, the rendering parameter information corresponding to the plurality of material types respectively is extracted; The rendering parameter information corresponding to the plurality of material types respectively is templated to obtain the plurality of materialization templates.

4. The method of claim 2, wherein, According to the material type in the group of video materials corresponding to the video instance identifier, at least one target materialization template is determined from the plurality of materialization templates, comprising: Identifying at least one material type contained in the group of video materials corresponding to the video instance identifier; According to the at least one material type contained in the group of video materials, at least one initial materialization template is selected from the plurality of materialization templates, and each material type corresponds to an initial materialization template; Parameter adjustment is performed on at least part of the at least one initial materialization template to obtain the at least one target materialization template.

5. The method of claim 2, wherein, Based on the hierarchical relationship between the plurality of video materials included in the initial video template, the at least one target materialization template is combined to obtain the target video template corresponding to the video instance identifier, comprising: According to the hierarchical relationship between the plurality of video materials included in the initial video template, a basic video template is generated, the basic video template includes a plurality of blank structure positions corresponding to the plurality of video materials; the positional relationship between the plurality of blank structure positions reflects the hierarchical relationship between the plurality of video materials; The at least one target materialization template is inserted into the corresponding blank structure position of the basic video template to obtain the target video template corresponding to the video instance identifier; Or According to the at least one target materialization template, the structure position where the rendering rule of the video material of the same material type in the initial video template is located is covered to obtain the target video template corresponding to the video instance identifier; wherein the rendering rule of each video material in the initial video template occupies a structure position, and the positional relationship between the structure positions reflects the hierarchical relationship between the plurality of video materials.

6. The method of claim 1, wherein, Each group of video material description information includes: subtitle description information, audio type description information, digital person description information and background picture description information; then according to the group of video material description information corresponding to the video instance identifier, a plurality of material generation models based on artificial intelligence are called to generate a plurality of video materials corresponding to the video instance identifier, comprising: According to the subtitle description information, a generative language model is called to generate text information to obtain a target subtitle; According to the target subtitle and the audio type description information, a text-to-speech model is called to convert the target subtitle into a target audio that is adapted to the audio type description information; according to the target audio and the digital person description information, a multi-modal model is called to select a target digital person according to the digital person description information, and generate a green screen video based on the target audio and the target digital person to obtain a green screen video of the target digital person; According to the background picture description information, a background picture model is called to generate a background picture to obtain a target background picture.

7. The method of claim 6, wherein, According to the target audio and the digital person description information, a multi-modal model is called to select a target digital person according to the digital person description information, and a green screen video is generated based on the target audio and the target digital person to obtain a digital person green screen video, including: Based on the digital person description information, a corresponding target digital person video frame is obtained; According to the target digital person video frame and the target audio, a multi-modal model is called to extract multi-dimensional features of the target audio to obtain multi-dimensional speech features, including speech content features and speech emotion features; according to the speech content features, the mouth shape control parameters of the target digital person are determined; according to the speech emotion features, the facial expression control parameters of the target digital person are determined; according to the speech content features and the speech emotion features, the body movement control parameters of the target digital person are determined; Based on the mouth shape control parameters, expression control parameters and body movement control parameters of the target digital person, a green screen video of the target digital person is generated; the mouth shape, facial expression and body movement of the target digital person green screen video match the target audio.

8. The method of claim 1, wherein, Also includes: For N video instance identifiers, the generation state of a corresponding set of video materials is maintained respectively, and the generation state of any video material of any material type in each set of video materials includes: a ready-to-generate state, a generating state, a generation success state and a generation failure state; For each video instance identifier, the video instance identifier corresponding to the plurality of video materials is marked as a ready-to-generate state; When calling a plurality of material generation models based on artificial intelligence to generate a plurality of video materials corresponding to the video instance identifier, if any material generation model returns a generating response message, the video material to be generated by the material generation model is updated to a generating state; If any material generation model returns a generation success message, the video material to be generated by the material generation model is updated to a generation success state; If any material generation model returns a generation failure message, the video material to be generated by the material generation model is updated to a generation failure state; If the generation state of the video material is updated to a generation failure state, a failure reminder message is output to the user who initiates the input operation, and the failure reminder message includes the material type of the video material and the corresponding video instance identifier thereof; If a user-triggered re-generation operation is received, the description information of the video material is obtained according to the video instance identifier corresponding to the video material, and the corresponding material generation model based on artificial intelligence is called to regenerate the video material.

9. The method of claim 1, wherein, The method further includes: uploading the N video instance identifiers corresponding to the videos and the target video template to a content distribution network, and obtaining the access links of the N video instance identifiers corresponding to the videos and the target video template in the content distribution network, and adding the access links to a video generation result page; In response to a viewing operation of the video result page, a video generation result page is displayed, the video generation result page including a set of video materials corresponding to at least one video instance identifier among the N video instance identifiers, a video, and an access link of the target video template in the content distribution network; In response to a triggering operation of the access link of the set of video materials corresponding to at least one video instance identifier, the video, and / or the target video template in the content distribution network, the set of video materials corresponding to the at least one video instance identifier, the video, and / or the target video template is accessed.

10. An electronic device, comprising: A computer program product comprising a memory and a processor, the memory configured to store a computer program, the processor coupled to the memory and configured to execute the computer program to implement the steps of any of the methods of claims 1-9.

11. A computer readable storage medium storing computer programs / instructions, characterized in that, The computer program / instructions, when executed by the processor, cause the processor to implement the steps of any of the methods of claims 1-9.

12. A computer program product, characterised in that, The computer program / instructions, when executed by the processor, cause the processor to implement the steps of any of the methods of claims 1-9. The computer program / instructions, when executed by the processor, cause the processor to implement the steps of any of the methods of claims 1-9.

Citation Information

Patent Citations

  • Short video template generation method and device, server and storage medium

    CN113111222A

  • Video editing processing method and device, electronic equipment and storage medium

    CN117880581A