Video batch generation method and device, storage medium and program product
By combining the generation task and the artificial intelligence material generation model, the problems of monotony of video style and homogeneity of content in the existing technology are solved, and the diversity and efficiency of video generation are achieved.
Patent Information
- Application Number
- CN202510443557.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the existing video generation technology, batch generation based on preset templates leads to monotonous video style and serious homogeneity of content.
By generating batch video generation tasks, multiple material generation models based on artificial intelligence are called, multiple video materials are generated based on material description information related to video categories, and video generation processing is performed in combination with target video templates.
It has improved the diversification of video generation, enriched the types of video material resources and templates, reduced the load pressure on the server, and improved the generation efficiency.
Smart Images

Figure CN120186429A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, storage medium, and program product for batch video generation. Background Art
[0002] In the field of video generation, in order to reduce the learning threshold and improve the video generation efficiency, in some video editing software, preset templates are provided. The preset templates contain some fixed content. By replacing the fixed content such as text and pictures in the preset templates with the user's materials, the required videos can be generated. Generating videos through preset templates simplifies the video generation process and is suitable for batch video generation. However, the videos generated in batches based on preset templates have a monotonous style and serious content homogenization. Summary of the Invention
[0003] Embodiments of this application provide a method, device, storage medium, and program product for batch video generation to improve the diversity of video generation.
[0004] An embodiment of this application provides a method for batch video generation, including: responding to an input operation on a video generation page for the number of videos and video categories, generating a batch video generation task, where the batch video generation task includes the number of videos N and video categories, and N is an integer greater than or equal to 2; according to the batch video generation task, generating N video instance identifiers, and obtaining a set of video material description information related to the video category for each video instance identifier, where each set of video material description information includes description information of multiple video materials; for each video instance identifier, according to the set of video material description information corresponding to the video instance identifier, calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously uploading the multiple video materials corresponding to the video instance identifier to a content delivery network; according to the multiple video materials corresponding to each video instance identifier, determining a target video template corresponding to each video instance identifier, where the target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to the video instance identifier; in response to a batch video generation trigger event, according to the N video instance identifiers, respectively obtaining the multiple video materials corresponding to the N video instance identifiers from the content delivery network; performing video generation processing according to the multiple video materials and target video templates respectively corresponding to the N video instance identifiers to obtain N videos under the video category.
[0005] An embodiment of this application further provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor is coupled to the memory and used to execute the computer program to implement the steps in each method provided by the embodiments of this application.
[0006] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method.
[0007] An embodiment of the present application also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, enable the processor to implement the steps in the above-described method embodiments.
[0008] In the embodiments of the present application, description information of multiple types of materials corresponding to N video instance identifiers is obtained. Based on the description information of the multiple types of materials, multiple material generation models based on artificial intelligence are respectively called to generate video materials, and respective corresponding multiple video materials are generated for each video instance identifier, enriching the video material resources and improving the diversity of video content; moreover, the video template for each video instance identifier is determined according to its corresponding multiple video materials, having a corresponding relationship, enriching the types of video templates and further improving the diversity of the generated video content. In addition, the video materials are uploaded to a content delivery network to respectively obtain the multiple video materials corresponding to the video instance identifiers from the content delivery network, and video generation is performed in combination with the corresponding target video template. Through the hosting of the content delivery network, the load pressure on the server side is reduced and the generation efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0010] Figure 1 FIG. is a schematic structural diagram of a video batch generation system provided for an exemplary embodiment of the present application;
[0011] Figure 2 FIG. is a schematic structural diagram of a video batch generation system provided for another exemplary embodiment of the present application;
[0012] Figure 3 FIG. is a schematic flowchart of a video batch generation method provided for an exemplary embodiment of the present application;
[0013] Figure 4 FIG. is a schematic structural diagram of an electronic device provided for yet another exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0015] It should be noted that in the case where the embodiments of this application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject. Additionally, various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0016] Furthermore, it should be noted that in the case where the embodiments of this application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of this application include but are not limited to: interaction operations in various ways such as touch operations, gesture operations, voice operations, head movement operations, eye movement operations, etc.; among them, touch operations include but are not limited to: click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations, etc. Swipe operations include but are not limited to: straight-line swipes, curved swipes, etc.
[0017] Moreover, it should be noted that in the case where the embodiments of this application involve the jump between a first interface and a second interface, the jump methods involved in the embodiments of this application include but are not limited to: directly jumping from the first interface to the second interface, first jumping from the first interface to a task interface and then jumping to the second interface when corresponding task operations are completed on the task interface; completing the corresponding task operations on the task interface includes but is not limited to: when the task interface is implemented as a game interface, completing game operations on the game interface; when the task interface is implemented as an identity authentication interface, completing identity authentication on the identity authentication interface; when the task interface is implemented as a recharge interface, completing recharge operations on the recharge interface; and so on.
[0018] In view of the technical problems of monotonous video styles and serious content homogenization generated in batches based on preset templates, in the embodiments of the present application, description information of multiple types of materials corresponding to N video instance identifiers is obtained, and based on the description information of multiple types of materials, multiple material generation models based on artificial intelligence are respectively called to generate video materials, and multiple corresponding video materials are generated for each video instance identifier, enriching the video material resources and improving the diversity of video content; moreover, the video template of each video instance identifier is determined according to the corresponding multiple video materials, having a corresponding relationship, enriching the types of video templates and further improving the diversity of the generated video content. In addition, the video materials are uploaded to the content delivery network to respectively obtain multiple video materials corresponding to the video instance identifiers from the content delivery network, and video generation is performed in combination with the corresponding target video templates. Through the hosting of the content delivery network, the load pressure on the server side is reduced and the generation efficiency is improved.
[0019] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0020] Figure 1 FIG. 7 is a schematic structural diagram of a video batch generation system provided by an exemplary embodiment of the present application. The video batch generation system 100 includes: a client 10, a server 11, a content delivery network 13, and multiple material generation models, such as a material generation model a1, a material generation model a2, and a material generation model ax, where x is greater than or equal to 2 and is an integer. Among them, on the basis of adopting a client-server architecture, the client 10 provides a video generation page 12 for users. The video generation page 12 is a page for providing batch video generation services. Through this video generation page 12, users can initiate a batch video generation task to the server 11. Correspondingly, the server 11 receives the batch video generation task initiated by the client 10 through the video generation service page 12. On the one hand, it calls multiple material generation models to generate multiple video materials required for each video and uploads the generated video materials to the content delivery network 13; on the other hand, it obtains multiple pre-generated materials from the content delivery network 13 to perform batch video generation, so that video materials and videos can be generated by leveraging resources such as the computing resources, network bandwidth, and storage resources of the server 11.
[0021] In this embodiment, in response to an input operation on the video generation page for the number of videos and video categories, a batch video generation task is generated. The batch video generation task includes the number of videos N and video categories, where N is an integer greater than or equal to 2;
[0022] Among them, the batch video generation task is used to describe the generation of N videos under a certain video category. The specific implementation of the video category is not limited. For example, it can be implemented as a single-level video category, such as but not limited to industrial and commercial registration or legal consultation, etc. Optionally, the video category can also be implemented as a multi-level video category. For example, the first-level video category can be industrial and commercial registration; the second-level video category of this first-level video category can be cleaning.
[0023] In this embodiment, according to the batch video generation task, N video instance identifiers are generated. Among them, the N video instance identifiers are all different, and each video instance identifier is used to uniquely represent a video to be generated, so as to facilitate the tracking of the required video materials and video templates of the video to be generated. That is to say, this video instance identifier can also be used as the unique identity identifier of the relevant content (such as video materials) of the video to be generated. Among them, the method of generating N video instance identifiers is not limited. For example, it can be an increasing sequence starting from any integer with a step size of 1; or it can also be a string, etc.
[0024] In this embodiment, a set of description information of video materials related to the video category is generated for each video instance identifier, and this set of description information of video materials is used to generate the video materials corresponding to this video instance identifier. Among them, each set of video materials includes video materials of multiple material types. The material type refers to the type of different video elements that make up the video content, including but not limited to: audio, video, background image, subtitle, digital human and other material types.
[0025] In this embodiment, the implementation manner of generating a set of description information of video materials related to the video category for each video instance identifier is not limited.
[0026] In an optional implementation manner, for any video instance identifier, the description information of video materials can be randomly extracted from multiple material types stored in the basic material library, and the extracted description information of multiple video materials is used as a set of description information of video materials for any video instance identifier. Among them, the basic material library stores the description information of multiple video materials under multiple video categories.
[0027] In another optional implementation manner, for any video instance identifier, the semantic similarity between the description information of multiple video materials in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the description information of the video materials that meet the similarity condition among the multiple similarity information is used as a set of description information of video materials for this video instance identifier.
[0028] In yet another alternative embodiment, for any video instance identifier, according to the video category, a material description information generation model is called. This model is used to generate description information for various video materials related to the video category. Among them, the material description information generation model is trained by combining the description information of a large number of sample video materials of different sample video categories and different material types. By learning the semantic correlation between the description information of different sample video categories and the sample video materials of different material types, this model can specifically combine different video categories and generate description information for various video materials related to them.
[0029] Furthermore, for each video instance identifier, according to a set of video material description information corresponding to the video instance identifier, multiple material generation models based on artificial intelligence are called to generate various video materials corresponding to the video instance identifier, and the various video materials corresponding to the video instance identifier are synchronously uploaded to the content delivery network.
[0030] Among them, one material generation model can generate at least some of the various video materials. For example, one material generation model can generate one video material; or, one material generation model can also generate two or more video materials, and this is not limited.
[0031] In this embodiment, when generating the video materials of the video instance identifier, according to the various video materials corresponding to each video instance identifier, the target video template corresponding to each video instance identifier is determined. The target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the various video materials corresponding to the video instance identifier.
[0032] In this embodiment, the determination method of the target video template is not limited. The following provides two examples, but is not limited to this.
[0033] In an alternative embodiment, a video template library is pre-constructed. The video template library includes multiple video templates, and each video template includes the pre-designed rendering rules and hierarchical relationships of video materials. Among them, the video templates in the video template library can be classified differently according to the video category. When determining the target video template corresponding to each video instance identifier, N video templates can be randomly determined from the multiple video templates under the video category as the target video templates of N video instance identifiers, and each video instance identifier corresponds to one target video template.
[0034] In yet another alternative embodiment, corresponding materialized templates can be generated respectively according to various materials corresponding to each video instance identifier. Furthermore, the various materialized templates corresponding to various video materials are combined according to the combination logic to obtain the target video template corresponding to each video instance identifier. The combination logic includes a hierarchical relationship, which refers to the layer stacking order of various video materials in the video generation process and is used to determine the front-back covering relationship and occlusion logic of various video materials in video generation. For example, in generating a video, subtitles can be located in the relatively upper layer of the layer, the background image can be located in the relatively lower layer of the layer, and the digital human can be superimposed on the upper layer of the background image and below the layer where the subtitles are located.
[0035] In this embodiment, in response to a batch video generation trigger event, according to N video instance identifiers, various video materials corresponding to the N video instance identifiers are respectively obtained from the content delivery network; video generation processing is performed according to the various video materials and the target video templates respectively corresponding to the N video instance identifiers to obtain N videos under the video category. The video generation processing refers to the process of filling and rendering video materials for the target video template based on the target video template corresponding to each video instance identifier and a set of video materials to obtain the video corresponding to the video instance identifier. For example, for N video instance identifiers, a set of video materials respectively corresponding to the N video instance identifiers can be filled into the N target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain N videos under the video category.
[0036] In this embodiment, the specific implementation of the batch video generation trigger event is not limited and can be flexibly configured according to actual application requirements. For example, it can be in the case where all the various materials corresponding to the N video instance identifiers are generated; or, it can also be in the case where every time a set of materials corresponding to a video instance identifier is generated, in which case, video generation is performed for the various video materials corresponding to the video instance identifier; or, the batch video generation trigger event can also be a preset trigger time, such as it can be after a period of time in response to an input operation on the video quantity and video category on the video generation page. For example, it can be 2 hours, 1 day or 1 week, etc., and the time span is not limited.
[0037] It should be noted that each video instance identifier corresponds to the generation of a video, and the N video instance identifiers can be N videos to be generated. When performing video generation processing on the various video materials respectively corresponding to the N video instance identifiers, the video generation processing process corresponding to each video instance identifier is asynchronous, and the video generation processing corresponding to each video instance identifier does not affect each other, so as to improve the generation efficiency of the N videos.
[0038] In the embodiments of the present application, description information of multiple types of materials corresponding to N video instance identifiers is obtained. Based on the description information of the multiple types of materials, multiple material generation models based on artificial intelligence are respectively called to generate video materials, and respective corresponding multiple video materials are generated for each video instance identifier, enriching the video material resources and improving the diversity of video content; moreover, the video template for each video instance identifier is determined according to the corresponding multiple video materials, having a corresponding relationship, enriching the types of video templates and further improving the diversity of the generated video content. In addition, the video materials are uploaded to a content delivery network to respectively obtain the multiple video materials corresponding to the video instance identifiers from the content delivery network, and video generation is performed in combination with the corresponding target video template. Through the hosting of the content delivery network, the load pressure on the server side is reduced and the generation efficiency is improved.
[0039] In the above embodiments, the generation of multiple types of materials and batch video generation in combination with a content delivery network are introduced. As described in the above embodiments, the specific implementation of the target video template in the embodiments of the present application is not limited. The following gives an implementation manner of generating a target video template from a materialized template based on variable processing, but is not limited thereto.
[0040] In this embodiment, multiple materialized templates obtained by performing variable processing on an initial video template are obtained. Optionally, the initial video template can be implemented as an MLT (Media Lovin' Toolkit) template. The MLT template is an open-source framework for multimedia processing, which is a file based on the XML format and records all parameters of video editing, such as video clips on the timeline, audio tracks, filter effects, transitions, etc., and can be used for video editing.
[0042] Among them, the initial video template includes the rendering rules of multiple video materials required for generating a video, as well as the hierarchical relationship between the multiple video materials. The rendering rule corresponding to each video material is used to control the rendering logic of the video material during the rendering process to achieve the expected rendering result in the generated video. Among them, the formation of a hierarchical relationship between multiple video materials in the initial video template means that there is an overlay order of multiple video materials in the video generation process in the initial video template, which is used to determine the front-to-back covering relationship and occlusion logic of multiple video materials in video generation. For example, in generating a video, the subtitle can be located in the relatively upper layer of the layer, the background image can be located in the relatively lower layer of the layer, the digital human can be superimposed on the upper layer of the background image, and be located below the layer of the subtitle.
[0043] In this embodiment, part of the variational processing is reflected in taking the material type as the splitting variable, structurally splitting the initial video template and templatizing the rendering parameter information to obtain a materialized template. Among them, there is a corresponding relationship between the materialized template and the material type. For example, the material types include materials such as audio, video, background image, subtitle, digital human, etc. Then the types of the corresponding materialized templates include but are not limited to: audio materialized template, video materialized template, background image materialized template, subtitle materialized template, digital human materialized template, and so on. In other words, each material type corresponds to a materialized template, and each materialized template is used to describe the rendering rules of this type of video material.
[0044] In this embodiment, there is no limitation on the timing of the variational processing. For example, the initial video template can be variably processed in advance. Another example is that the initial video template can also be variably processed dynamically. For the detailed content on how to perform the variational processing, reference can be made to the subsequent embodiments.
[0045] In this embodiment, the materialized template obtained through variational processing can be modularly reorganized. For each video instance identifier, at least one target materialized template is determined from multiple materialized templates according to the material types in a group of video materials corresponding to the video instance identifier. There is a one-to-one correspondence between the material types in this group of video materials and the materialized templates.
[0046] For example, in the case where the material types of the audio, video, background image, and subtitle included in this group of video materials, the target materialized templates can include the target materialized templates corresponding to the audio, video, background image, and subtitle respectively; another example is that in the case where the material types of the audio, video, background image, subtitle, and digital human included in this group of video materials, the target materialized templates can include the target materialized templates corresponding to the material types of the audio, video, background image, subtitle, and digital human respectively.
[0047] Furthermore, in the case of obtaining at least one target materialized template corresponding to a group of video materials, based on the hierarchical relationship between multiple video materials included in the initial video template, at least one target materialized template is combined to obtain a target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationship of a group of video materials.
[0048] Among them, for the hierarchical relationship between multiple video materials included in the initial video template, this hierarchical relationship can represent the layer stacking order of multiple materialized templates, and can be used to organize and integrate the target materialized templates, so as to form a target video template with a clear hierarchy and corresponding video material rendering rules. Among them, the hierarchical relationship of a group of video materials described by the target video template is consistent with the hierarchical relationship of this group of video materials in the initial video template.
[0049] Among them, in this embodiment, the N video instance identifiers respectively correspond to their respective target video templates. That is to say, each group of video materials has its corresponding target video template, enriching the types of target video templates used for batch video generation. Among them, based on the rendering rules described by the N target video templates, video generation processing is performed on the grouped video materials corresponding to the N video instance identifiers, so that the style of each batch-generated video matches the adopted video template respectively, improving the diversification degree of batch-generated video content.
[0050] In the case of obtaining the target video template corresponding to the video instance identifier, video generation processing is performed according to the target video templates respectively corresponding to the N video instance identifiers and a group of video materials to obtain N videos under the video category. For example, for the N video instance identifiers, a group of video materials respectively corresponding to the N video instance identifiers can be filled into the N target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain the videos under the video category.
[0051] In an optional embodiment, the video generation corresponding to the N video instance identifiers can be processed in batches, with M videos processed in parallel in each batch, where M is less than N and M is an integer; after the generation of M videos is completed, the generation of the next M videos continues until the videos corresponding to the N video instance identifiers are all generated. Through batch processing, the resource utilization rate of the server is improved, and at the same time, the server pressure overload caused by high concurrency is avoided.
[0052] In the embodiment of the present application, in batch video generation, the materialized templates of various material types obtained by variable processing of the initial video template are used to reorganize the materialized templates to obtain the video templates required for video generation: furthermore, a video instance identifier corresponding to the generation of each video is generated, and the video instance identifier is bound to a respective group of video materials. According to the group of video materials corresponding to the video instance identifier, the target materialized template for reorganization is determined from various materialized templates, and combined with the hierarchical relationship between the materialized templates provided by the initial template, the target materialized template is organized and integrated to obtain the video templates respectively corresponding to the video instance identifiers for video generation processing of the respective groups of video materials corresponding to each video instance identifier; since the video template can be obtained by personalized reorganization in combination with the video materials for generating the video, the flexibility of video generation is improved; and, the reorganized video template has a corresponding relationship with the video instance identifier, enriching the types of video templates to a certain extent, so that the style of each batch-generated video matches the adopted video template, improving the diversification degree of batch-generated video content.
[0053] The following introduces the method of variable processing.
[0054] In an optional embodiment, the step of variable processing on the initial video template includes: obtaining the initial video template, where the initial video template includes rendering rules of various video materials required for generating a video and the hierarchical relationship between various video materials; parsing the initial video template with the material type as the splitting variable to obtain information segments corresponding to various material types; respectively extracting rendering parameter information corresponding to various material types from the information segments corresponding to various material types.
[0055] Template the rendering parameter information corresponding to various material types to obtain various materialized templates, and each materialized template is used to describe the rendering rules of a video material.
[0056] Furthermore, after obtaining the initial video template, parse the initial video template with the material type as the splitting variable. Among them, the material type refers to different types of video materials, which is a classification standard for distinguishing different video materials, and the splitting variable refers to the independent information segments obtained by dividing the initial video template according to the material type when parsing the initial video template. For example, if the initial video template contains two types of video materials, namely images and audio, the initial video template will be split into independent information segments of images and audio according to the material type as the classification standard, and subsequent processing will be carried out separately.
[0057] Among them, the information segment refers to the information segment extracted from the initial video template and corresponding to the material type, and the information segment may include rendering parameter information of the corresponding material type. For example, an information segment may include rendering parameter information such as path information, playing time, and special effect application of the material type.
[0058] In the embodiment of the present application, the rendering parameter information is used to describe at least one attribute of a material file of a material type. Each attribute can be used as a rendering parameter, and the attribute value of each attribute in the initial video template is used as the default parameter value of the corresponding rendering parameter. The rendering parameters with default parameter values form a rendering parameter information corresponding to the material file of this material type. Among them, the "material file" refers to various resource files used in the materialized template, mainly including various video materials, such as pictures, videos, and audio, etc. The rendering parameter information can be used to control the rendering logic of the material file so that the material file can obtain the rendering effect presented in the finally generated video. The rendering parameter information quantitatively describes the attributes of the material file, so that the attributes of the video material are converted into the attributes of the rendering parameter information, and the rendering effect of the video material in the finally generated video is controlled through the rendering parameter information. In other words, the attributes of the video material are described as corresponding rendering parameters by the rendering parameter information, and the rendering parameters with default parameter values define the attribute values of a certain attribute of the material file.
[0059] Among them, the rendering parameter information can be extracted separately from the information segments corresponding to various material types. The rendering parameter information corresponding to each material type includes at least one rendering parameter and its corresponding default parameter value, and each rendering parameter is used to describe the attributes of the material file of a material type. In the embodiments of the present application, the specific attributes corresponding to the rendering parameters of various material types are not limited.
[0060] In the embodiments of the present application, templatization can transform the rendering parameter information corresponding to various material types from the rendering parameters with fixed configuration default parameter values into a variable parameter set with dynamically replaceable parameter values. Among them, by templatizing the rendering parameter information corresponding to various material types, various materialization templates corresponding to various material types can be obtained. Each materialization template is used to describe the rendering rules of a material type, and the rendering rules include at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values, so as to control the rendering effect of video generation.
[0061] In the embodiments of the present application, the materialization template is the result of templatizing the rendering parameter information corresponding to various material types. It includes the path information placeholder of the material file, as well as variable parameters and multiple candidate parameter values associated with them. The materialization template describes the rendering rules of the material file of the corresponding material type in video generation. The rendering rules are used to describe the rendering logic followed by this type of video material during video rendering, so as to control the expected rendering effect of the presentation of this type of material file in video generation. Among them, the materialization template formed by transforming the rendering parameter information from the rendering parameters with fixed configuration and default parameter values into a dynamically replaceable variable parameter set can flexibly adjust the rendering effect of the material file in video generation in different scenarios by assigning different candidate parameter values to the variable parameters. Among them, the rendering parameter information is templatized to form a variety of materialization templates that can be modularly recombined. For the materialization templates of different material types, different target video templates can be recombined, and then diverse videos can be batch-generated without designing and adjusting a separate video template for each video, which not only saves time and human resources, but also improves the flexibility and content diversity of video template generation, and then realizes efficient batch generation of diverse style videos.
[0062] For example, for the same materialization template, by assigning different candidate parameter values to the variable parameters, the attribute values corresponding to various attributes of the material type corresponding to the materialization template can be flexibly set, so as to realize the refined control of the rendering rules of the material file. The materialization template improves the flexibility of video generation, can significantly improve the efficiency and quality of video generation, and at the same time meet diverse creative needs.
[0063] In the embodiments of the present application, the specific content of the candidate parameter values associated with the variable parameters obtained after templatizing the rendering parameter information corresponding to each of the multiple material types is not limited.
[0064] Among them, by associating multiple candidate parameter values with the variable parameters, the variable parameters can be flexibly adjusted according to requirements, improving the flexibility and diversity of video template generation, and thus realizing efficient batch generation of videos with diverse styles.
[0065] In an optional embodiment, using the material type as a splitting variable, the initial video template is parsed to obtain information segments corresponding to multiple material types, including: loading the XML document corresponding to the initial video template, where the XML document includes a root element and multiple non-root elements connected to the root element, and multiple specific elements are included in the multiple non-root elements, and each specific element is used to describe the rendering rules of a material type; starting from the root element, traversing the non-root elements in the XML document to identify multiple specific elements; extracting the multiple information segments where the multiple specific elements are located as the information segments corresponding to multiple material types. Among them, the XML document of the initial video template can be loaded into memory to form a parseable tree structure. In the embodiments of the present application, the XML document can be loaded into memory through a DOM (Document Object Model) parser to form a tree structure parsing method.
[0066] Among them, the tree structure includes a root element and non-root elements. The root element is the top-level node of the XML document, and the non-root elements are all directly or indirectly nested under the root element. An XML document has one and only one root element, and the root element is the first element of the XML document and can be used as the starting point of the XML document.
[0067] Among them, the non-root elements are the child elements directly or indirectly nested under the root element and are used to divide different information segments of the initial video template according to the material type. Multiple specific elements are included in the non-root elements. Starting from the root element, traversing the non-root elements in the XML document can identify multiple specific elements. A specific element is an element that can directly describe the rendering rules of a certain type of material, and each specific element corresponds to a material type.
[0068] By extracting these information segments, the system can quickly identify the attributes of the material files and perform rendering according to their attribute values. This design of extracting information segments significantly improves the flexibility and diversity of video generation, meets the diverse creation requirements, and provides a technical basis for efficient batch video generation.
[0069] In an optional embodiment, starting from the root element, non-root elements in the XML document are traversed to identify multiple specific elements, including: S1. Starting from the root element, traverse non-root elements in the XML document; S2. For the currently traversed non-root element, obtain the element tag included in the currently traversed non-root element; S3. If the element tag is a specific tag, determine whether the currently traversed non-root element contains sub-elements; S4. If the currently traversed non-root element contains sub-elements, use the sub-elements as the currently traversed non-root element and return to execute step S2; S5. If the element tag is a non-specific tag, continue to the next non-root element and return to execute step S2; S6. If the currently traversed non-root element does not contain sub-elements, use the currently traversed non-root element as a specific element.
[0070] Among them, in step S1, starting from the root element of the XML document, non-root elements are accessed one by one for traversal. In the embodiments of the present application, the specific implementation strategy of traversal is not limited. For example, traversal can be implemented using a depth-first search algorithm or a breadth-first search algorithm. In the embodiments of the present application, taking the depth-first search algorithm as an example, the process of traversing non-root elements in the XML document to identify multiple specific elements is described in detail.
[0071] Next, execute step S2. For the currently traversed non-root element, obtain the element tag included in the currently traversed non-root element. Among them, in the XML document, the element tag is the identifier within angle brackets (<>), which is used to mark the type and semantic meaning of the element. Different material types are distinguished by the name of the element tag; the specific position of the material is located through the hierarchical relationship of the element tags.
[0072] After obtaining the element tag included in the currently traversed non-root element, determine whether the element tag is a specific tag. Among them, the specific tag can be a predefined set of key tags, representing the element tags that need to be specially processed in the XML document, and are used to identify the material types to be extracted.
[0073] Next, execute step S3 or S5. If step S5 is executed, that is, the element tag is a non-specific tag, continue to the next non-root element and return to execute step S2.
[0074] If step S3 is executed, that is, the element tag is a specific tag, continue to determine whether the currently traversed non-root element contains sub-elements. Among them, the sub-elements can be other elements nested inside the currently traversed non-root element.
[0075] Next, perform step S4 or S6. If step S4 is executed, that is, the non-root element currently traversed contains child elements, then use the child elements as the non-root element currently traversed, and return to execute step S2. If step S6 is executed, that is, the non-root element currently traversed does not contain child elements, then use the non-root element currently traversed as a specific element. Among them, step S4 is the recursive processing when there are child elements. Set the child elements of the non-root element currently traversed as the new traversal starting point, and re-execute steps S2 - S6 for each child element. After the recursion ends, continue to traverse other non-root elements. Step S6 is the end collection when there are no child elements, and use the non-root element currently traversed as a specific element.
[0076] In an optional embodiment, render parameter information corresponding to each of multiple material types is respectively extracted from information segments corresponding to the multiple material types, including: for each information segment, extract the path information of the material file and at least one attribute value of the material file from the information segment, where the attribute value is used to render the material file; read the material file according to the path information, and determine the material type described by the information segment according to the extension of the material file; use at least one attribute to which at least one attribute value belongs as at least one render parameter, and use at least one attribute value as the default parameter value of at least one render parameter, so as to obtain the render parameter information corresponding to the material type described by the information segment.
[0077] Among them, the information segment is a part related to the material type extracted from the XML document, and contains the path information of the material file and at least one attribute value of the material file. Among them, the material file can be used to construct various media files for the finally generated video. In the embodiments of the present application, the material file may include but is not limited to: video files in formats such as MP4, AVI, etc., which contain dynamic images and audio and are used to display a series of consecutive pictures; audio files in formats such as WAV, MP3, etc., which provide background music, narration, or special effect sounds, etc.
[0078] Among them, at least one attribute value of the material file extracted from the information segment is the specific value corresponding to the attribute, so as to describe the specific state of the attribute. The attribute can describe the rendering effect of the material file. For example, for a material file that is a video file, the attribute can be video duration, video speed, filter effect, etc. If the attribute is video duration, the attribute value can be 2 minutes, 1 hour, 1 day, etc.
[0079] In an alternative embodiment, the rendering parameter information corresponding to the material type described by the information segment includes at least one rendering parameter and the default parameter value of at least one rendering parameter. The default parameter value may be the attribute value of the attribute corresponding to the rendering parameter before quantization. Among them, at least one attribute for extracting at least one attribute value of the material file in the information segment may be used as at least one rendering parameter, and at least one attribute value may be used as the default parameter value of at least one rendering parameter to obtain the rendering parameter information. That is to say, a key-value pair set is formed by the rendering parameter and the default parameter value to control the rendering logic of the material file in the video to achieve the expected rendering effect.
[0080] In an alternative embodiment, the rendering parameter information corresponding to each of multiple material types is templatized to obtain multiple material templates, including: for each material type, selecting at least one variable parameter from the rendering parameter information corresponding to the material type; associating multiple candidate parameter values with at least one variable parameter; and generating a material template corresponding to the material type according to the multiple candidate parameter values associated with at least one variable parameter. Among them, the variable parameter is obtained by variable processing of the rendering parameter with a default parameter value. Each variable parameter is associated with multiple candidate parameter values, and different candidate parameter values correspond to different rendering logics to produce different rendering effects in the generated video, which can be used to generate diverse material templates. For each material type, at least one variable parameter is selected from the rendering parameter information corresponding to it, and these variable parameters can vary within the adjustment range to obtain different material templates. For example, when the material file is a video material, the playback speed and filter effects can be selected as variable parameters. The variable parameter is associated with multiple candidate optional values to limit the adjustment range of the variable parameter, ensuring that the material file of the generated video meets the design requirements and avoiding invalid settings.
[0081] In an alternative embodiment, when selecting at least one variable parameter from the rendering parameter information corresponding to the material type, two selection methods are provided: one is to use all rendering parameters as variable parameters, and the other is to select some rendering parameters as variable parameters according to the weight value. Among them, when choosing to use all rendering parameters as variable parameters, there is no need to compare the weight values, and all rendering parameters can be adjusted as variable parameters.
[0082] In an optional embodiment, the manner of selecting some rendering parameters as variable parameters according to weight values may be to select at least one variable parameter from the rendering parameter information corresponding to the material type, including: pre-configuring the weight values of each rendering parameter for the material type; parsing each rendering parameter from the rendering parameter information corresponding to the material type, and selecting at least one rendering parameter with a weight value greater than a set weight threshold from each rendering parameter as at least one variable parameter. Among them, the weight value can be an importance score assigned to each rendering parameter, used to quantify the influence degree of the rendering parameter on the rendering effect of the material file in video generation, and at least one rendering parameter that has a greater impact on user perception or application goals can be selected as at least one variable parameter. The weight threshold is a predefined critical value. By screening the rendering parameters corresponding to the weight values higher than the weight threshold, the user can select the parameters with weight values higher than the weight threshold as variable parameters, avoiding taking too many irrelevant rendering parameters as variable parameters to prevent problems such as configuration conflicts or rendering logic confusion.
[0083] In the embodiments of the present application, the manner of pre-configuring the weight values of each rendering parameter is not limited. For example, the weight values of each rendering parameter can be pre-configured according to empirical annotation or automatically generated through user behavior data analysis.
[0084] Optionally, selecting at least one variable parameter from the rendering parameter information corresponding to the material type can also be randomly selected according to the set number of variable parameters. In an optional embodiment, generating a materialized template corresponding to the material type according to multiple candidate parameter values associated with at least one variable parameter includes: adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to a preset template file, and adding multiple candidate parameter values associated with at least one variable parameter to the preset template file; and adding a placeholder for carrying the material file corresponding to the material type in the preset template file to obtain a materialized template corresponding to the material type. Among them, the preset template file is a basic video template containing basic configurations, including the basic structure of the video template and some preset rendering parameters. Among them, adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to the preset template file can ensure that the finally obtained materialized template has complete rendering parameter information.
[0085] In this embodiment, on the basis of adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to the preset template file, multiple candidate parameter values associated with at least one variable parameter are added to the preset template file. Among them, the candidate parameter values provide multiple choices, allowing the user or the system to select different candidate parameter values according to needs.
[0086] In an alternative embodiment, based on adding multiple candidate parameter values associated with at least one variable parameter to a preset template file, a placeholder for carrying a material file corresponding to a material type can be added to the preset template file to obtain a materialized template corresponding to the material type. The placeholder is a reserved position for filling the material file and supports dynamic replacement. For example, the path information of the actual material file corresponding to the material type can be filled. Through the placeholder, different material files can be flexibly replaced without modifying the template structure.
[0087] In the above embodiment, by integrating the rendering parameters of the material type and their default parameter values into a preset basic template, the integrity of the rendering parameter information of the generated materialized template is ensured. At the same time, by associating multiple candidate parameter values with the variable parameter, dynamic selection of the parameter values of the variable parameter is realized to flexibly adjust the video style. The embedding of the placeholder further decouples the parameter configuration of the initial video template from the material file and supports dynamic replacement of the material file without modifying the template structure. Thus, it is allowed to assign corresponding candidate parameter values to the variable parameter according to the application requirements to obtain a freely combined materialized template, ultimately significantly improving the flexibility and diversity of video template generation, and then efficiently batch-producing video content with different styles, solving problems such as cumbersome parameter configuration of video templates, complex adaptation process, and serious homogenization of generated video content in traditional video production.
[0088] Based on obtaining the above multiple materialized templates, batch video generation or single video generation can be performed based on the multiple materialized templates. For the batch video generation scenario, using the multiple materialized templates provided by the embodiments of the present application, multiple videos with diverse styles and different contents can be generated. A detailed description of an implementation manner for batch video generation based on the multiple materialized templates provided by the embodiments of the present application is given below.
[0089] In the embodiments of the present application, the initial video template is variable-processed to obtain various materialized templates. During batch video generation, for each video instance identifier, at least one target materialized template is determined from various materialized templates according to the material type in a set of video materials corresponding to the video instance identifier.
[0090] In an optional embodiment, when determining at least one target materialization template from multiple materialization templates according to the material types in a group of video materials corresponding to the video instance identifier, it includes: identifying at least one material type included in the group of video materials corresponding to the video instance identifier; selecting at least one initial materialization template from multiple materialization templates according to at least one material type included in the group of video materials, with each material type corresponding to one initial materialization template; and adjusting the parameters of at least some of the selected at least one initial materialization templates to obtain at least one target materialization template.
[0091] In this embodiment, multiple materialization templates obtained by variable processing of the initial video template are called initial materialization templates. In the case of determining at least one initial materialization template corresponding to a certain grouping from multiple initial materialization templates, the rendering rules described by the initial materialization templates can be called initial rendering rules. Further, the parameters of at least some of the at least one initial materialization templates are adjusted to obtain at least one target materialization template. Each initial materialization template obtains a corresponding target materialization template after parameter adjustment. The target materialization template after parameter adjustment includes a target rendering rule, and the target rendering rule is different from the initial rendering rule described by the initial materialization template.
[0092] In this embodiment, since the initial materialization template is obtained by variable processing, the variable processing can convert the fixed rendering parameter information in the initial video template into variable parameters that can be dynamically assigned values to achieve flexible configuration of the template content. If the values of the optional parameters are different, the rendering rules will be different. Optionally, each initial materialization template includes at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values and default parameter values.
[0093] In this embodiment, the parameter adjustment of the initial materialization template is to assign different candidate parameter values to the variable parameters of the initial materialization template to obtain multiple target materialization templates with different candidate parameter values. Among them, the target materialization template for a certain material type can be used for rendering the video materials in different groupings of video instance identifiers. If the parameter values of the variable parameters are different, it means that the rendering rules described by the target materialization templates of the same material type are different, so that the rendering results of the batch-generated videos are different, and differences are formed between different videos, improving the richness of video content. Hereinafter, how to adjust the parameters of the initial materialization template will be introduced.
[0094] In an alternative embodiment, when adjusting parameters of at least a part of at least one initial materialized template to obtain at least one target materialized template, the following steps are included: determining the number of templates for which parameters are to be adjusted, where the number of templates is less than or equal to the number of at least one initial materialized template; selecting the initial materialized templates to be adjusted from at least one initial materialized template according to the number of templates to be adjusted; determining the variable parameters to be adjusted from the initial materialized templates to be adjusted; randomly determining target parameter values from multiple candidate parameter values associated with the variable parameters to be adjusted; and assigning the target parameter values to the variable parameters to be adjusted to obtain the target materialized template.
[0095] In this embodiment, there is no limitation on the method for determining the number of templates for which parameters are to be adjusted. For example, if parameter adjustment is performed on each initialized video template, then the number of templates to be adjusted is the number of at least one initial materialized template. Another example is that, according to the random number generation algorithm, a random integer within a preset range is generated as the number of templates to be adjusted, and the preset range refers to being less than or equal to the number of at least one initial materialized template. There is no limitation on the random number generation algorithm, including but not limited to: the Linear congruential generator (LCG) and the Mersenne Twister, etc.
[0096] Furthermore, select the initial materialized templates to be adjusted from at least one initial materialized template according to the number of templates to be adjusted. In an alternative embodiment, select the initial materialized templates to be adjusted according to the priority of the material types and the number of templates to be adjusted. Here, the priority of the material type refers to the importance degree of the material type to the presentation effect of the generated video. For example, the subtitle material type generally has a relatively low importance degree to the presentation effect, so the priority of the subtitle material type can be set to a lower priority; relatively speaking, the priority of the background image can be higher than that of the subtitle, so it can be set to a medium priority; and, the digital human has a relatively high importance degree to the presentation effect, so it can be set to a higher priority. Then, according to the high and low priorities, preferentially select the templates for which parameters are to be adjusted from the initial material templates with a higher priority of the material type. If the number of these initial material templates is less than the previously determined number of templates to be adjusted, then selection can be made from the initial materialized templates with a lower priority of the material type. The finally selected number of templates to be adjusted is less than or equal to the number of at least one initial materialized template.
[0097] Further, determine the variable parameters to be adjusted from the initial materialized template to be adjusted. In an optional embodiment, all variable parameters in the initial materialized template to be adjusted can be used as the variable parameters to be adjusted. In another optional embodiment, the target variable parameters in the initial materialized template to be adjusted are used as the optional parameters to be adjusted, and the target variable parameters are pre-selected optional parameters.
[0098] Furthermore, randomly determine the target parameter values from multiple candidate parameter values associated with the variable parameters to be adjusted. In some embodiments, among the at least one initial materialized template corresponding to each of the N video instance identifiers, there are the same variable parameters to be adjusted. The multiple candidate parameter values associated with the variable parameters to be adjusted can be randomly used as the target parameter values for each of the N video instance identifiers, so that the optional parameter values of the N video instance identifiers are as different as possible.
[0099] Further, assign the target parameter values to the variable parameters to be adjusted to obtain the target materialized template.
[0100] In the case of obtaining the target materialized template, combine the target materialized templates to obtain the target video template. The combination method is not limited in this embodiment. Two combination methods are provided below, but not limited thereto.
[0101] In an optional embodiment, according to the hierarchical relationship between multiple video materials included in the initial video template, generate a basic video template. This basic video template serves as the framework of the target video template and includes multiple blank structure positions corresponding to multiple video materials. Among them, the blank structure position is a placeholder for the target materialized template preset in the basic video template, used to identify the position where the target materialized template can be inserted. Each blank structure position corresponds to the filling of the target materialized template of one material type. The positional relationship between multiple blank structure positions reflects the hierarchical relationship between multiple video materials. As described in the above embodiment, the hierarchical relationship represents the front-to-back stacking order of multiple materials in the generated video. In some embodiments, this hierarchical relationship is extracted from the initial video template; alternatively, it can also be preset, preset based on at least one target materialized template, that is to say, the hierarchical relationship can be set as needed. For example, the structure position corresponding to the subtitle is located in the upper layer, the structure position corresponding to the digital human is located in the middle layer, and the structure position corresponding to the background image is located in the lower layer. Further, insert at least one target materialized template into the corresponding blank structure position in the basic video template to obtain the target video template corresponding to the video instance identifier.
[0102] In another alternative embodiment, according to at least one target materialization template, the structural position where the rendering rule of video materials of the same material type in the initial video template is overwritten to obtain the target video template corresponding to the video instance identifier. The difference between the structural position and the above-mentioned blank structural position is that the rendering rule of each video material in the initial video template occupies a structural position, and the blank structural position is empty. Among them, the positional relationship between the structural positions reflects the hierarchical relationship between multiple video materials.
[0103] In the case of obtaining the target video template, video generation processing is performed according to the target video template corresponding to each video instance identifier and a set of video materials. In an alternative embodiment, when performing video generation processing according to the target video templates corresponding to N video instance identifiers respectively and a set of video materials to obtain N videos under the video category, it includes: for each video instance identifier, filling the set of video materials corresponding to the video instance identifier into the target materialization template in the target video template corresponding to the video instance identifier; for the filled target video template, rendering the set of video materials according to the rendering rules and hierarchical relationship of the set of video materials described in the filled target video template to obtain a video under the video category.
[0104] Among them, the target materialization template includes at least one placeholder corresponding to the material type for filling video materials of the material type.
[0105] Figure 2 It is a schematic structural diagram of a video batch generation system provided by an exemplary embodiment of the present application. The video batch generation system includes an endpoint, a content delivery network, and multiple material generation models. In this embodiment, taking multiple material generation models including a generative language model, a text-to-speech model, a multimodal model, and a text-to-image model as an example, but not limited thereto.
[0106] In this embodiment, the endpoint is used for the generation of video materials, the management of batch video generation tasks, and the control of the video generation process.
[0107] Among them, the endpoints include front-end code and back-end code, that is, the client-server structure is adopted as described in the above embodiments. The front-end code refers to a video generation page built based on a front-end framework. This video generation page runs on the client and is used to interact with users, receive users' input operations, and initiate a batch video generation task to the server. The back-end code of the endpoint runs on the server. The back-end code is obtained by API (Application Programming Interface) -ifying the code of an existing video editing software. The existing video editing software can be, for example, shortcut, etc. Among them, in the existing video editing software, the front-end UI code is highly coupled with the video rendering function code. In this embodiment, it is split according to functions to decouple the front-end UI of the video editing software from the video rendering function code, and then encapsulate the video rendering function code of the existing video editing software into an independent API, so that the outside can call the video rendering function through the standard API method to achieve automated video generation.
[0108] For the endpoints as Figure 2 described, in one example, the front-end code of the endpoint can be a video generation page built based on Astro, which is responsible for receiving callback notifications. For example, when the material generation model processing is completed, it can send a notification to the endpoint to notify that the subsequent process of video generation can continue. The back-end code of the endpoint can be obtained by API -ifying the video rendering code of the video editing software and exposed to the outside through the API interface for external calls.
[0109] In this embodiment, as Figure 2 shown, in response to the input operations on the video generation page for the number of videos and video categories, a batch video generation task is generated. The batch video generation task includes the number of videos N and the video category. N is an integer greater than or equal to 2. For the detailed content of the batch generation task, reference can be made to the above embodiments.
[0110] In this embodiment, according to the batch video generation task, N video instance identifiers are generated. For the implementation method of generating N video instance identifiers, reference can be made to the above embodiments and will not be elaborated here. Among them, the N video instance identifiers are all different, and each video instance identifier can correspond to the generation of a video, and this video instance identifier can also be used as the unique identity identifier of the video during the generation process.
[0111] In this embodiment, a set of description information of video materials related to the video category is generated for each video instance identifier. This set of description information of video materials is used to generate the video materials corresponding to the video instance identifier. For the generation method of the description information, reference can be made to the relevant content of the above embodiments.
[0112] Among them, each group of video materials contains the description information of video materials of various material types. In this embodiment, the description information of each group of video materials includes, for example, subtitle description information, audio type description information, digital human description information, and background image description information, but is not limited thereto.
[0113] Further, for each video instance identifier, according to a group of video material description information corresponding to the video instance identifier, call multiple material generation models based on artificial intelligence to generate various video materials corresponding to the video instance identifier, and synchronously upload the various video materials corresponding to the video instance identifier to the content delivery network.
[0114] Among them, one material generation model can generate at least some of the various video materials. For example, one material generation model can generate one type of video material. The following also takes this as an example for illustration, but is not limited thereto.
[0115] In an optional embodiment, when calling multiple material generation models based on artificial intelligence according to a group of video material description information corresponding to each video instance identifier to generate various video materials corresponding to each video instance identifier, it includes: according to the subtitle description information, call a generative language model to generate text information to obtain a target subtitle; according to the target subtitle and the audio type description information, call a text-to-speech model to convert the target subtitle into a target audio adapted to the audio type description information; according to the target audio and the digital human description information, call a multimodal model, select a target digital human according to the digital human description information, and generate a green screen video based on the target audio and the target digital human to obtain a green screen video of the target digital human; according to the background image description information, call an image generation model from text to generate a background image to obtain a target background image.
[0116] In an optional embodiment, when calling a generative language model according to the subtitle description information to generate text information to obtain a target subtitle, it includes: matching corresponding keywords according to the video category, calling a pre-designed prompt template, filling the keywords of the video category into the prompt template to obtain a prompt for the video category; inputting the prompt for the video category into the generative language model to perform text information to obtain a target subtitle related to the video category.
[0117] Among them, the keywords are used to describe the expression theme of the video category. Different video categories can correspond to different keywords, and the prompt templates corresponding to different video types can also be different. Among them, the target subtitle includes all the text content required for each video.
[0118] Further optionally, according to the target subtitle and the audio type description information, call the text-to-speech model to convert the target subtitle into a target audio adapted to the audio type description information. Among them, the target subtitle is used to provide text content, and the audio type description information is used to specify the audio type of the generated target audio. The audio type includes, but is not limited to: audio format, audio language, audio tone, audio quality, etc. Any audio type that can be used to specify the sound effect of the target audio is applicable to this embodiment.
[0119] Further, call the multi-modal model to generate a green screen video of the virtual human. Among them, a virtual human refers to a virtual character generated based on AI technology, which can synchronously simulate the appearance, voice, lip shape and other behaviors of a real person, and can be used in scenarios such as intelligent customer service, short video production, virtual anchors, etc., but is not limited to this. In this embodiment, by combining various video materials corresponding to each video instance identifier, a video with the effect of real person speaking can be generated.
[0120] In an optional embodiment, the virtual human can be a pre-recorded real person video or picture. In subsequent embodiments, the real person video and picture are collectively referred to as video frames, and the number of video frames can be one or more. In this case, the description information for different virtual humans can be implemented as identification information, which serves as the unique identity identifier of the virtual human and is used to obtain the video frames of the virtual human corresponding to the identification information. In another optional embodiment, the video frames of the virtual human can be dynamically generated based on the multi-modal model. In this case, the description information of the virtual human can be the prompt words used to generate the virtual human, and the description information of the virtual human corresponding to each video instance identifier can be different.
[0121] Further optionally, when, according to the target audio and the virtual human description information, call the multi-modal model, select the target virtual human according to the virtual human description information, and generate a green screen video based on the target audio and the target virtual human to obtain the green screen video of the target virtual human, it includes: based on the virtual human description information, obtain the corresponding target virtual human video frames; according to the video frames of the target virtual human and the target audio, call the multi-modal model to perform multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features; among them, the multi-dimensional speech features include, but are not limited to: speech content features and speech emotion features; according to the speech content features, determine the lip movement control parameters of the target virtual human; according to the speech emotion features, determine the facial expression control parameters of the target virtual human; according to the speech content features and speech emotion features, determine the body movement control parameters of the target virtual human; based on the lip movement control parameters, expression control parameters and body movement control parameters of the target virtual human, generate the green screen video of the target virtual human; the lip shape, facial expression and body movement of the green screen video of the target virtual human match the target audio.
[0122] In this embodiment, the speech content features are used to reflect the semantic information in the target audio, such as lexical content, grammatical structure, speech intention, and speech rhythm, and are mainly used to drive the lip synchronization of the digital human with the semantics of the target audio. The speech emotion features are used to reflect the emotional state of the target audio, including but not limited to the strength of the tone, the speed of speech, and the intonation change, etc., and are mainly used to drive the facial expressions and body movements of the digital human.
[0123] Furthermore, according to the speech content features, determine the lip control parameters of the target digital human; according to the speech emotion features, determine the facial expression control parameters of the target digital human; according to the speech content features and speech emotion features, determine the body movement control parameters of the target digital human.
[0124] In the case of obtaining the lip control parameters, facial expression control parameters, and body movement control parameters, based on the lip control parameters, expression control parameters, and body movement control parameters of the target digital human, drive the target digital human to perform action rendering to generate the corresponding green screen video of the target digital human, which is convenient for subsequent flexible replacement with the background image. Since the green screen video of the target digital human is generated by controlling the multi-dimensional speech features extracted from the target audio content, it is ensured that the dynamic performances of the lip, facial expression, and body movement of the target digital human in the green screen video match the target audio in terms of semantics, timing, and emotion.
[0125] Furthermore, in this embodiment, the target audio and target subtitles are aligned to ensure the consistency of the target subtitles and the target audio on the time axis. For example, through the timestamps of the speech content in the target audio, timestamp inference and alignment annotation are performed on the target subtitles to ensure that the subsequent display of the target subtitles is completely synchronized with the target audio.
[0126] In this embodiment, there are two ways to generate the background image. One is to call the text-to-image model to generate the background image according to the background image description information to obtain the target background image. The other is to directly obtain the pre-generated or captured background image, and there is no limitation on this.
[0127] In an optional embodiment, for each of the N video instance identifiers, a corresponding set of video material generation states is maintained. The generation state of any video material of any material type in each set of video materials includes: ready-to-generate state, generating state, generation success state, and generation failure state. For example, for any video instance identifier, taking the multiple video materials that need to be generated for this video instance identifier including: target subtitles, target audio, green screen video of the target digital human, and target background image as an example. Then, the generation states of the target subtitles, target audio, green screen video of the target digital human, and target background image can be maintained respectively.
[0128] Among them, for each video instance identifier, mark multiple video materials corresponding to the video instance identifier as the ready-to-generate state; when calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, if a generation-in-progress response message is returned by any material generation model, update the video material that the material generation model should generate to the generation-in-progress state; if any material generation model returns a generation-success message, update the video material that the material generation model should generate to the generation-success state; if any material generation model returns a generation-failure message, update the video material that the material generation model should generate to the generation-failure state. As Figure 2 shown, a subscription service is provided. The subscription service can receive notifications from each material generation model to notify the subscription service of the generation state of the video materials of the material model, for example, it can be the generation-success state. Furthermore, the subscription service will return a generation-success message to the endpoint to inform the endpoint that the corresponding video material has been generated successfully. Figure 2 Only the notification process of the target digital human is taken as an example in this section, but it is not limited to this.
[0129] Further optionally, if the generation state of the video material is updated to the generation-failure state, a failure reminder message is output to the user who initiated the input operation. The failure reminder message includes the material type of the video material updated to the generation-failure state and its corresponding video instance identifier. If a regeneration operation triggered by the user is received, according to the video instance identifier corresponding to the video material, obtain the description information of the corresponding video material, and call the corresponding material generation model based on artificial intelligence to regenerate the corresponding video material.
[0130] In this embodiment, in the case of generating multiple video materials corresponding to each video instance identifier, the multiple video materials corresponding to each video instance identifier can be synchronously uploaded to the content delivery network, and the access link of the video material corresponding to each video instance identifier in the content delivery network can be obtained. Continuing with the above example, for any video instance identifier, in the case where the multiple video materials corresponding to the video instance identifier include the target subtitle, the target audio, the green screen video of the target digital human, and the target background image, the target subtitle, the target audio, the green screen video of the target digital human, and the target background image are respectively uploaded to the content delivery network, and the access links of the target subtitle, the target audio, the green screen video of the target digital human, and the target background image are obtained.
[0131] Further, in response to a batch video generation trigger event, access links of various video materials corresponding to each of the N video instance identifiers in a content delivery network are obtained, and various video materials corresponding to the N video instance identifiers are respectively acquired; and, video generation is performed based on the various video materials and a target video template corresponding to each of the N video instance identifiers to obtain N videos under a video category. Among them, when performing video generation based on the various video materials and the target video template corresponding to each of the N video instance identifiers, for the N video instance identifiers, the various video materials corresponding to each of the N video instance identifiers can be filled into their respective target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain N videos under the video category. Optionally, the rendering can be performed by calling an API corresponding to the video rendering function exposed to the video generation page based on the above endpoints.
[0132] Continuing with the above example, in response to a batch video generation trigger event, access links of the target subtitles, target audio, green screen videos of the target digital humans, and target background images corresponding to each of the N video instance identifiers in a content delivery network are obtained, and the target subtitles, target audio, green screen videos of the target digital humans, and target background images corresponding to the N video instance identifiers are respectively acquired; and, video generation is performed based on the target subtitles, target audio, green screen videos of the target digital humans, target background images, and a target video template corresponding to each of the N video instance identifiers to obtain N videos under a video category.
[0133] Further optionally, the videos corresponding to the N video instance identifiers and the target video template are uploaded to the content delivery network, and access links of the videos corresponding to the N video instance identifiers and the target video template in the content delivery network are obtained, and the access links are added to the video generation result page; in response to a viewing operation on the video result page, the video generation result page is displayed, and the video generation result page includes access links of a set of video materials, videos, and the target video template corresponding to at least one of the N video instance identifiers in the content delivery network.
[0134] In this optional embodiment, in response to a trigger operation on the access links of a set of video materials, videos, and / or the target video template corresponding to at least one video instance identifier in the content delivery network, access is performed on the set of video materials, videos, and / or the target video template corresponding to the at least one video instance identifier information.
[0135] Figure 3 The flowchart of a video batch generation method provided by an exemplary embodiment of this application. As Figure 3 shown, the method includes:
[0136] S301: In response to an input operation on the video quantity and video category on the video generation page, generate a batch video generation task. The batch video generation task includes a video quantity N and a video category, where N is an integer greater than or equal to 2;
[0137] S302: According to the batch video generation task, generate N video instance identifiers, and obtain a set of video material description information related to the video category for each video instance identifier. Each set of video material description information includes the description information of multiple video materials;
[0138] S303: For each video instance identifier, according to a set of video material description information corresponding to the video instance identifier, call multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously upload the multiple video materials corresponding to the video instance identifier to the content distribution network;
[0139] S304: According to the multiple video materials corresponding to each video instance identifier, determine the target video template corresponding to each video instance identifier. The target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to the video instance identifier;
[0140] S305: In response to a batch video generation trigger event, according to the N video instance identifiers, respectively obtain the multiple video materials corresponding to the N video instance identifiers from the content distribution network; perform video generation processing according to the multiple video materials and the target video templates respectively corresponding to the N video instance identifiers to obtain N videos under the video category.
[0141] In an optional embodiment, when determining the target video template corresponding to each video instance identifier according to the multiple video materials corresponding to each video instance identifier, it includes: obtaining multiple materialized templates obtained by variable processing of the initial video template. The initial video template includes the rendering rules of multiple video materials required for generating a video and the hierarchical relationships between the multiple video materials. Each materialized template is used to describe the rendering rule of a video material; for each video instance identifier, determine at least one target materialized template from the multiple materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; based on the hierarchical relationships between the multiple video materials included in the initial video template, combine the at least one target materialized template to obtain the target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationships of the set of video materials.
[0142] In an alternative embodiment, when performing variable processing on the initial video template, it includes: using the material type as a splitting variable to parse the initial video template to obtain information segments corresponding to multiple material types; respectively extracting the rendering parameter information corresponding to each of the multiple material types from the information segments corresponding to each of the multiple material types; and templatizing the rendering parameter information corresponding to each of the multiple material types to obtain the multiple materialized templates.
[0143] In an alternative embodiment, when determining at least one target materialized template from the multiple materialized templates according to the material types in a group of video materials corresponding to the video instance identifier, it includes: identifying at least one material type included in the group of video materials corresponding to the video instance identifier; selecting at least one initial materialized template from the multiple materialized templates according to the at least one material type included in the group of video materials, with each material type corresponding to one initial materialized template; and adjusting the parameters of at least some of the at least one initial materialized template to obtain the at least one target materialized template.
[0144] In an alternative embodiment, when combining the at least one target materialized template based on the hierarchical relationship between multiple video materials included in the initial video template to obtain the target video template corresponding to the video instance identifier, it includes: generating a basic video template according to the hierarchical relationship between multiple video materials included in the initial video template, where the basic video template includes multiple blank structure positions corresponding to the multiple video materials; the positional relationship between the multiple blank structure positions reflects the hierarchical relationship between the multiple video materials; inserting the at least one target materialized template into the corresponding blank structure positions in the basic video template respectively to obtain the target video template corresponding to the video instance identifier; or, covering the structure positions where the rendering rules of the video materials of the same material type in the initial video template are located according to the at least one target materialized template to obtain the target video template corresponding to the video instance identifier; where, in the initial video template, the rendering rule of each video material occupies one structure position, and the positional relationship between the structure positions reflects the hierarchical relationship between the multiple video materials.
[0145] In an alternative embodiment, each set of video material description information includes: subtitle description information, audio type description information, digital human description information, and background image description information. When generating multiple video materials corresponding to the video instance identifier by invoking multiple material generation models based on artificial intelligence according to the set of video material description information corresponding to the video instance identifier, it includes: invoking a generative language model according to the subtitle description information to generate text information to obtain a target subtitle; invoking a text-to-speech model according to the target subtitle and the audio type description information to convert the target subtitle into a target audio adapted to the audio type description information; invoking a multimodal model according to the target audio and the digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain the green screen video of the target digital human; invoking an image generation model according to the background image description information to generate a background image to obtain a target background image.
[0146] In an alternative embodiment, when generating a green screen video of a digital human by invoking a multimodal model according to the target audio and the digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video based on the target audio and the target digital human, it includes: obtaining corresponding target digital human video frames based on the digital human description information; invoking a multimodal model according to the video frames of the target digital human and the target audio to perform multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features, where the multi-dimensional speech features include: speech content features, speech emotion features; determining lip movement control parameters of the target digital human according to the speech content features; determining facial expression control parameters of the target digital human according to the speech emotion features; determining limb movement control parameters of the target digital human according to the speech content features and speech emotion features; generating a green screen video of the target digital human based on the lip movement control parameters, expression control parameters, and limb movement control parameters of the target digital human; the lip movement, facial expression, and limb movement of the green screen video of the target digital human match the target audio.
[0147] In an alternative embodiment, after invoking a text-to-speech model according to the target subtitle and the audio type description information to convert the target subtitle into a target audio adapted to the audio type description information, it further includes: performing alignment processing on the target audio and the target subtitle to make the target audio and the target audio time-synchronized.
[0148] In an optional embodiment, the above method further includes: maintaining the generation status of a corresponding set of video materials for each of the N video instance identifiers, where the generation status of the video materials of any material type in each set of video materials includes: a ready-to-generate status, a generating status, a generation success status, and a generation failure status; for each video instance identifier, marking multiple video materials corresponding to the video instance identifier as in the ready-to-generate status; when calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, if any material generation model returns a generating response message, updating the video materials that should be generated by the material generation model to the generating status; if any material generation model returns a generation success message, updating the video materials that should be generated by the material generation model to the generation success status; if any material generation model returns a generation failure message, updating the video materials that should be generated by the material generation model to the generation failure status; if the generation status of the video materials is updated to the generation failure status, outputting a failure reminder message to the user who initiated the input operation, where the failure reminder message includes the material type of the video materials and its corresponding video instance identifier; if a regeneration operation triggered by the user is received, obtaining the description information of the video materials according to the video instance identifier corresponding to the video materials, and calling the corresponding material generation model based on artificial intelligence to regenerate the video materials.
[0149] In an optional embodiment, the above method further includes: uploading the videos and the target video template corresponding to the N video instance identifiers to a content delivery network, and obtaining the access links of the videos and the target video template corresponding to the N video instance identifiers in the content delivery network, and adding the access links to the video generation result page; in response to a viewing operation of the video result page, displaying the video generation result page, where the video generation result page includes the access links of a set of video materials, videos, and the target video template corresponding to at least one of the N video instance identifiers in the content delivery network; in response to a trigger operation on the access links of a set of video materials, videos, and / or the target video template corresponding to at least one video instance identifier in the content delivery network, accessing the set of video materials, videos, and / or the target video template corresponding to the at least one video instance identifier information.
[0150] The detailed implementation manners and beneficial effects of the steps in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0151] It should be noted that the execution entity of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution entity. For example, the execution entity of steps 301 to 303 can be device A; for another example, the execution entity of steps 301 and 302 can be device A, and the execution entity of step 303 can be device B; and so on.
[0152] In addition, in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this document or in parallel. The operation numbers such as 301, 302, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0153] Figure 4 It is a schematic structural diagram of an electronic device provided by another exemplary embodiment of the present application. As Figure 4 shown, the device includes: a memory 44 and a processor 45.
[0154] The memory 44 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of these data include instructions for any application program or method for operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0155] A processor 45, coupled to a memory 44, is configured to execute a computer program in the memory 44 for: generating a batch video generation task in response to input operations on a video generation page for the number of videos and video categories, the batch video generation task including the number of videos N and video categories, where N is an integer greater than or equal to 2; generating N video instance identifiers according to the batch video generation task, and obtaining a set of video material description information related to the video category for each video instance identifier, each set of video material description information including description information of multiple video materials; for each video instance identifier, calling multiple material generation models based on artificial intelligence according to the set of video material description information corresponding to the video instance identifier, generating multiple video materials corresponding to the video instance identifier, and synchronously uploading the multiple video materials corresponding to the video instance identifier to a content delivery network; determining a target video template corresponding to each video instance identifier according to the multiple video materials corresponding to each video instance identifier, where the target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to the video instance identifier; in response to a batch video generation trigger event, obtaining the multiple video materials corresponding to the N video instance identifiers from the content delivery network respectively according to the N video instance identifiers; performing video generation processing according to the multiple video materials and target video templates respectively corresponding to the N video instance identifiers to obtain N videos under the video category.
[0156] In an optional embodiment, when the processor 45 determines the target video template corresponding to each video instance identifier according to the multiple video materials corresponding to each video instance identifier, it is specifically configured to: obtain multiple materialized templates obtained by variable processing of an initial video template, where the initial video template includes the rendering rules of multiple video materials required for generating a video and the hierarchical relationships between the multiple video materials, and each materialized template is used to describe the rendering rule of one video material; for each video instance identifier, determine at least one target materialized template from the multiple materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; based on the hierarchical relationships between the multiple video materials included in the initial video template, combine the at least one target materialized template to obtain the target video template corresponding to the video instance identifier, and the target video template is used to describe the rendering rules and hierarchical relationships of the set of video materials.
[0157] In an alternative embodiment, when the processor 45 performs variable processing on the initial video template, it is specifically configured to: use the material type as a splitting variable to parse the initial video template to obtain information segments corresponding to multiple material types; extract the rendering parameter information corresponding to each of the multiple material types from the information segments corresponding to each of the multiple material types; and template the rendering parameter information corresponding to each of the multiple material types to obtain the multiple materialized templates.
[0158] In an alternative embodiment, when the processor 45 determines at least one target materialized template from the multiple materialized templates according to the material types in a group of video materials corresponding to the video instance identifier, it is specifically configured to: identify at least one material type included in the group of video materials corresponding to the video instance identifier; select at least one initial materialized template from the multiple materialized templates according to the at least one material type included in the group of video materials, with each material type corresponding to one initial materialized template; and adjust the parameters of at least some of the at least one initial materialized template to obtain the at least one target materialized template.
[0159] In an alternative embodiment, when the processor 45 combines the at least one target materialized template based on the hierarchical relationship between multiple video materials included in the initial video template to obtain the target video template corresponding to the video instance identifier, it is specifically configured to: generate a basic video template according to the hierarchical relationship between multiple video materials included in the initial video template, where the basic video template includes multiple blank structure positions corresponding to the multiple video materials; the positional relationship between the multiple blank structure positions reflects the hierarchical relationship between the multiple video materials; insert the at least one target materialized template into the corresponding blank structure positions in the basic video template to obtain the target video template corresponding to the video instance identifier; or, cover the structure positions where the rendering rules of the video materials of the same material type in the initial video template are located according to the at least one target materialized template to obtain the target video template corresponding to the video instance identifier; where, in the initial video template, the rendering rule of each video material occupies one structure position, and the positional relationship between the structure positions reflects the hierarchical relationship between the multiple video materials.
[0160] In an alternative embodiment, each set of video material description information includes: subtitle description information, audio type description information, digital human description information, and background image description information; when the processor 45 generates multiple types of video materials corresponding to the video instance identifier by invoking multiple material generation models based on artificial intelligence according to the set of video material description information corresponding to the video instance identifier, it specifically is used for: according to the subtitle description information, invoking a generative language model to generate text information to obtain a target subtitle; according to the target subtitle and the audio type description information, invoking a text-to-speech model to convert the target subtitle into a target audio adapted to the audio type description information; according to the target audio and the digital human description information, invoking a multimodal model, selecting a target digital human according to the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain the green screen video of the target digital human; according to the background image description information, invoking an image generation model from text to generate a background image to obtain a target background image.
[0161] In an alternative embodiment, when the processor 45 invokes a multimodal model according to the target audio and the digital human description information, selects a target digital human according to the digital human description information, and generates a green screen video based on the target audio and the target digital human to obtain a digital human green screen video, it specifically is used for: based on the digital human description information, obtaining corresponding video frames of the target digital human; according to the video frames of the target digital human and the target audio, invoking a multimodal model to perform multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features, where the multi-dimensional speech features include: speech content features, speech emotion features; according to the speech content features, determining lip movement control parameters of the target digital human; according to the speech emotion features, determining facial expression control parameters of the target digital human; according to the speech content features and speech emotion features, determining limb movement control parameters of the target digital human; based on the lip movement control parameters, expression control parameters, and limb movement control parameters of the target digital human, generating the green screen video of the target digital human; the lip movement, facial expression, and limb movement of the green screen video of the target digital human match the target audio.
[0162] In an alternative embodiment, after the processor 45 invokes a text-to-speech model according to the target subtitle and the audio type description information to convert the target subtitle into a target audio adapted to the audio type description information, it is further used for: performing alignment processing on the target audio and the target subtitle to make the target audio and the target subtitle time-synchronized.
[0163] In an alternative embodiment, the processor 45 is further configured to: maintain the generation status of a corresponding set of video materials for each of the N video instance identifiers, where the generation status of the video materials of any material type in each set of video materials includes: a ready-to-generate status, a generating status, a generation success status, and a generation failure status; for each video instance identifier, mark multiple video materials corresponding to the video instance identifier as in the ready-to-generate status; when calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, if any material generation model returns a generating response message, update the video material to be generated by the material generation model to the generating status; if any material generation model returns a generation success message, update the video material to be generated by the material generation model to the generation success status; if any material generation model returns a generation failure message, update the video material to be generated by the material generation model to the generation failure status; if the generation status of the video material is updated to the generation failure status, output a failure reminder message to the user who initiated the input operation, where the failure reminder message includes the material type of the video material and its corresponding video instance identifier; if a re-generation operation triggered by the user is received, obtain the description information of the video material according to the video instance identifier corresponding to the video material, and call the corresponding material generation model based on artificial intelligence to re-generate the video material.
[0164] In an alternative embodiment, the processor 45 is further configured to: upload the videos and the target video template corresponding to the N video instance identifiers to a content delivery network, and obtain the access links of the videos and the target video template corresponding to the N video instance identifiers in the content delivery network, and add the access links to the video generation result page; in response to a viewing operation of the video result page, display the video generation result page, where the video generation result page includes the access links of a set of video materials, videos, and the target video template corresponding to at least one of the N video instance identifiers in the content delivery network; in response to a trigger operation on the access links of a set of video materials, videos, and / or the target video template corresponding to at least one video instance identifier in the content delivery network, access the set of video materials, videos, and / or the target video template corresponding to the at least one video instance identifier information.
[0165] Further, as Figure 4 shown, the electronic device further includes: other components such as a communication component 46, a display 47, a power supply component 48, and an audio component 49. Figure 4 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 4 the components shown. Additionally, Figure 4The components within the dashed-line box are optional components, rather than essential components, and can be determined according to the product form of the working node. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include Figure 4 the components within the dashed-line box; if the electronic device in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include Figure 4 the components within the dashed-line box.
[0166] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0167] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile communication network such as 2G, 3G, 4G / LTE, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.
[0168] The above-mentioned display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0169] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0170] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC). When the device where the audio component is located is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0171] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to be able to implement the steps in the above-mentioned method embodiments. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission medium
[0172] Accordingly, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processor, it causes the processor to be able to implement the steps in the above-mentioned method embodiments. It should be understood that each process or a combination of multiple processes in the above-mentioned method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices for implementing the corresponding functions in the above-mentioned method embodiments.
[0174] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0175] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for batch generating videos, characterized in that: include: In response to an input operation on the video generation page for the number of videos and the video category, a batch video generation task is generated, wherein the batch video generation task includes the number of videos N and the video category, where N is an integer ≥ 2; According to the batch video generation task, N video instance identifiers are generated, and a set of video material description information related to the video category is obtained for each video instance identifier, where each set of video material description information includes description information of multiple video materials; For each video instance identifier, based on a set of video material description information corresponding to the video instance identifier, multiple material generation models based on artificial intelligence are called to generate multiple video materials corresponding to the video instance identifier, and the multiple video materials corresponding to the video instance identifier are synchronously uploaded to a content distribution network; According to the multiple video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined, wherein the target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to the video instance identifier; In response to a batch video generation triggering event, multiple video materials corresponding to the N video instance identifiers are respectively obtained from the content distribution network according to the N video instance identifiers; video generation processing is performed according to the multiple video materials corresponding to each of the N video instance identifiers and a target video template to obtain N videos under the video category.
2. The method according to claim 1, characterized in that According to the multiple video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined, including: Acquire multiple material templates obtained by performing variable processing on the initial video template, wherein the initial video template includes rendering rules of multiple video materials required for generating a video and hierarchical relationships between the multiple video materials, and each material template is used to describe the rendering rules of a video material; For each video instance identifier, at least one target materialization template is determined from the multiple materialization templates according to the material type in a group of video materials corresponding to the video instance identifier; based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target materialization template is combined to obtain a target video template corresponding to the video instance identifier, and the target video template is used to describe the rendering rules and hierarchical relationship of the group of video materials.
3. The method according to claim 2, characterized in that The steps of performing variable processing on the initial video template include: Taking the material type as a splitting variable, parsing the initial video template to obtain information fragments corresponding to multiple material types; Extracting rendering parameter information corresponding to each of the multiple material types from the information fragments corresponding to each of the multiple material types; The rendering parameter information corresponding to each of the multiple material types is templated to obtain the multiple material templates.
4. The method according to claim 2, characterized in that: Determining at least one target materialization template from the multiple materialization templates according to the material type in a group of video materials corresponding to the video instance identifier includes: Identify at least one material type included in a group of video materials corresponding to the video instance identifier; According to at least one material type included in the group of video materials, selecting at least one initial material conversion template from the multiple material conversion templates, each material type corresponds to an initial material conversion template; Parameters of at least a portion of the at least one initial materialization template are adjusted to obtain the at least one target materialization template.
5. The method according to claim 2, characterized in that: Based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target material template is combined to obtain a target video template corresponding to the video instance identifier, including: Generate a basic video template according to the hierarchical relationship between the multiple video materials included in the initial video template, wherein the basic video template includes a plurality of blank structural positions corresponding to the multiple video materials; the positional relationship between the plurality of blank structural positions reflects the hierarchical relationship between the multiple video materials; Inserting the at least one target material template into corresponding blank structural positions in the basic video template respectively, so as to obtain a target video template corresponding to the video instance identifier; or According to the at least one target material template, the structural bits where the rendering rules of the video materials of the same material type in the initial video template are located are overwritten to obtain the target video template corresponding to the video instance identifier; wherein, in the initial video template, the rendering rules of each video material occupies a structural bit, and the positional relationship between the structural bits reflects the hierarchical relationship between the multiple video materials.
6. The method according to claim 1, characterized in that Each set of video material description information includes: subtitle description information, audio type description information, digital human description information and background image description information; then, according to a set of video material description information corresponding to the video instance identifier, multiple material generation models based on artificial intelligence are called to generate multiple video materials corresponding to the video instance identifier, including: According to the subtitle description information, calling a generative language model to generate text information to obtain a target subtitle; According to the target subtitles and the audio type description information, a text-to-speech model is called to convert the target subtitles into a target audio that matches the audio type description information; according to the target audio and the digital human description information, a multimodal model is called to select a target digital human according to the digital human description information, and a green screen video is generated based on the target audio and the target digital human to obtain a green screen video of the target digital human; According to the background image description information, the text image model is called to generate the background image to obtain the target background image.
7. The method according to claim 6, characterized in that According to the target audio and the digital human description information, a multimodal model is called, a target digital human is selected according to the digital human description information, and a green screen video is generated based on the target audio and the target digital human to obtain a digital human green screen video, including: Based on the digital human description information, obtaining a corresponding target digital human video frame; According to the video frame of the target digital human and the target audio, a multimodal model is called to extract multidimensional features of the target audio to obtain multidimensional speech features, wherein the multidimensional speech features include speech content features and speech emotion features; according to the speech content features, a lip shape control parameter of the target digital human is determined; according to the speech emotion features, a facial expression control parameter of the target digital human is determined; according to the speech content features and the speech emotion features, a body movement control parameter of the target digital human is determined; Based on the lip shape control parameters, facial expression control parameters and body movement control parameters of the target digital human, a green screen video of the target digital human is generated; the lip shape, facial expression and body movement of the green screen video of the target digital human are matched with the target audio.
8. The method according to claim 1, characterized in that Also includes: For each of the N video instance identifiers, a corresponding group of video material generation states are maintained. The generation state of any type of video material in each group of video materials includes: a state of preparing to generate, a state of generating, a state of successful generation, and a state of failed generation. For each video instance identifier, marking multiple video materials corresponding to the video instance identifier as being in a ready-to-generate state; When calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, if any material generation model returns a generating response message, the video material to be generated by the material generation model is updated to a generating state; If any material generation model returns a generation success message, the video material that should be generated by the material generation model is updated to a generation success state; If any material generation model returns a generation failure message, the video material that should be generated by the material generation model is updated to a generation failure state; If the generation state of the video material is updated to a generation failure state, a failure reminder message is output to the user who initiated the input operation, wherein the failure reminder message includes the material type of the video material and its corresponding video instance identifier; If a user-triggered regeneration operation is received, description information of the video material is obtained according to the video instance identifier corresponding to the video material, and the corresponding artificial intelligence-based material generation model is called to regenerate the video material.
9. The method according to claim 1, characterized in that: The method further comprises: Uploading the videos and target video templates corresponding to the N video instance identifiers to a content distribution network, obtaining access links of the videos and target video templates corresponding to the N video instance identifiers in the content distribution network, and adding the access links to a video generation result page; In response to a viewing operation on the video result page, displaying the video generation result page, the video generation result page including a group of video materials, videos and an access link of a target video template in the content distribution network corresponding to at least one video instance identifier among the N video instance identifiers; In response to a triggering operation of an access link in the content distribution network corresponding to a group of video materials, videos and / or target video templates corresponding to at least one video instance identification information, a group of video materials, videos and / or target video templates corresponding to the at least one video instance identification information is accessed.
10. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is coupled to the memory and is used to execute the computer program to implement the steps in the method according to any one of claims 1 to 9.
11. A computer-readable storage medium storing a computer program / instruction, characterized in that: When the computer program / instructions are executed by a processor, the processor is enabled to implement the steps in the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that include: A computer program / instruction, when executed by a processor, causes the processor to implement the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Short video template generation method and device, server and storage medium
CN113111222A
Video editing processing method and device, electronic equipment and storage medium
CN117880581A
Video generation method and device based on digital human, storage medium and program product
CN119277168A
Audio and video synchronization method and device, electronic device, and computer readable storage medium
US20240259126A1
Cited By
Advertisement picture batch generation method and device based on layer editing
CN120912701A