Video batch generation method and device, storage medium and program product
By variable processing and combining the initial video templates, diversified video templates are generated, which solves the problems of monotonous and homogeneous batch video generation style in the prior art, and achieves higher flexibility and diversity in video generation.
Patent Information
- Application Number
- CN202510443563.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-20
AI Technical Summary
When generating batch videos, existing video editing software has monotonous video style, serious homogeneity, and poor flexibility.
By variable processing of the initial video template, multiple materialized templates are generated, and a corresponding set of video materials is identified according to the video instance, the target materialized template is determined and the target video template is generated for video generation.
It improves the flexibility and diversity of video generation, so that each video style generated in batches matches the corresponding video templates, and increases the degree of diversification of video content.
Smart Images

Figure CN120186430A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, storage medium and program product for batch video generation. Background Art
[0002] In the field of video generation, video editing software can be used to manually edit materials to obtain videos that meet personal needs. However, video editing software has a certain learning threshold, and the labor cost is relatively high when using it, making it difficult to achieve batch video generation.
[0003] In order to reduce the learning threshold and improve the video generation efficiency, some video editing software provides preset templates, and the text or pictures in the preset templates are allowed to be replaced. The user's materials are used to replace the text or pictures in the preset templates to generate the videos required by the user. Among them, the method of directly replacing the text or pictures in the preset templates to generate new videos simplifies the video generation process and is suitable for batch video generation. However, the videos generated in batches based on preset templates have a monotonous style, serious homogenization, and poor flexibility in video generation. Summary of the Invention
[0004] Embodiments of this application provide a method, device, storage medium and program product for batch video generation, which are used to enrich the types of video templates, so that each batch-generated video has its own video template, thereby improving the flexibility and diversity of the batch-generated video content.
[0005] An embodiment of the present application provides a method for batch video generation, including: responding to an input operation on the video generation page for the number of videos and video categories, generating a batch video generation task, where the batch video generation task includes the number of videos N and video categories, and N is an integer greater than or equal to 2; generating N video instance identifiers according to the batch video generation task, and generating a set of video materials related to the video category for each video instance identifier; obtaining a variety of materialized templates obtained by variable processing of an initial video template, where the initial video template includes rendering rules for a variety of video materials required for video generation and the hierarchical relationship between a variety of video materials, and each materialized template is used to describe the rendering rules of a video material; for each video instance identifier, determining at least one target materialized template from the variety of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; based on the hierarchical relationship between a variety of video materials included in the initial video template, combining the at least one target materialized template to obtain a target video template corresponding to the video instance identifier, where the target video template is used to describe the rendering rules and hierarchical relationship of the set of video materials; performing video generation processing according to the target video templates and a set of video materials corresponding to the N video instance identifiers respectively to obtain N videos under the video category.
[0006] An embodiment of the present application also provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor is coupled to the memory and is used to execute the computer program to implement the steps in each method provided by the embodiment of the present application.
[0007] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it causes the processor to be able to implement the steps in the above-mentioned method.
[0008] An embodiment of the present application also provides a computer program product, which includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, it causes the processor to be able to implement the steps in the above-mentioned method embodiments.
[0009] In the embodiments of the present application, a materialized template of various material types obtained by variable processing of an initial video template is recombined to obtain a video template required for generating a video. Further, each video instance identifier corresponds to a video generation, and a respective set of video materials is bound thereto. According to a set of video materials corresponding to the video instance identifier, a target materialized template for recombination is determined, and in combination with the hierarchical relationship between the materialized templates, the target materialized template is recombined to obtain a video template corresponding to each video instance identifier for respective video generation processing. Since the video template is obtained by personalized recombination in combination with the video materials for generating the video, the flexibility of video generation is improved. In addition, the recombined video template has a corresponding relationship with the video instance identifier, which enriches the types of video templates to a certain extent, so that the style of each batch-generated video matches the corresponding video template, and the diversification degree of batch-generated video content is improved. Brief Description of the Drawings
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0011] Figure 1 It is a schematic flowchart of a method for batch video generation provided by an exemplary embodiment of the present application;
[0012] Figure 2 It is an interaction schematic diagram of a method for batch video generation in another exemplary embodiment of the present application;
[0013] Figure 3 It is a schematic structural diagram of an electronic device provided by another exemplary embodiment of the present application. Detailed Embodiments
[0014] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0015] It should be noted that in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse. Additionally, various models (including but not limited to language models or large models) involved in the present application comply with the relevant laws and standards.
[0016] In addition, it should be noted that in the case where the embodiments of the present application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of the present application include but are not limited to: interaction operations in various ways such as touch operations, gesture operations, voice operations, head movement operations, eye movement operations, etc.; among them, touch operations include but are not limited to: click operations, double-click operations, long-press operations, slide operations, pinch operations, or mouse hover operations, etc. Slide operations include but are not limited to: linear slides, curved slides, etc.
[0017] Furthermore, it should be noted that in the case where the embodiments of the present application involve the jump between the first interface and the second interface, the jump methods involved in the embodiments of the present application include but are not limited to: directly jumping from the first interface to the second interface, first jumping from the first interface to the task interface and then jumping to the second interface when corresponding task operations are completed on the task interface; completing the corresponding task operations on the task interface includes but is not limited to: when the task interface is implemented as a game interface, completing game operations on the game interface; when the task interface is implemented as an identity authentication interface, completing identity authentication on the identity authentication interface; when the task interface is implemented as a recharge interface, completing recharge operations on the recharge interface; and so on.
[0018] In view of the technical problems that the videos generated in batches according to a preset template have a monotonous style, serious homogenization, and poor flexibility in video generation, in the embodiments of the present application, a materialized template of various material types obtained by variable processing of the initial video template is used to reorganize the materialized template to obtain a video template required for video generation. Furthermore, each video instance identifier corresponds to a video generation, and a respective set of video materials is bound thereto. According to a set of video materials corresponding to the video instance identifier, a target materialized template for reorganization is determined, and in combination with the hierarchical relationship between the materialized templates, the target materialized template is reorganized to obtain a video template corresponding to each video instance identifier for respective video generation processing. Since the video template is obtained by personalized reorganization in combination with the video materials for video generation, the flexibility of video generation is improved. In addition, the reorganized video template has a corresponding relationship with the video instance identifier, which enriches the types of video templates to a certain extent, so that the style of each batch-generated video matches the corresponding video template, and the diversification degree of the batch-generated video content is improved.
[0019] The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0020] Figure 1 It is a schematic flowchart of a video batch generation method provided by an exemplary embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0021] S101: Respond to an input operation on the video generation page for the number of videos and video categories, and generate a batch video generation task. The batch video generation task includes the number of videos N and video categories, where N is an integer greater than or equal to 2;
[0022] S102: According to the batch video generation task, generate N video instance identifiers, and generate a set of video materials related to the video category for each video instance identifier;
[0023] S103: Obtain various materialized templates obtained by variable processing of the initial video template. The initial video template includes rendering rules for various video materials required for video generation and the hierarchical relationship between various video materials. Each materialized template is used to describe the rendering rules of a video material;
[0024] S104: For each video instance identifier, determine at least one target materialized template from various materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; based on the hierarchical relationship between various video materials included in the initial video template, combine at least one target materialized template to obtain a target video template corresponding to the video instance identifier. The target video template is used to describe the rendering rules and hierarchical relationship of a set of video materials;
[0025] S105: Perform video generation processing according to the target video templates and a set of video materials corresponding to N video instances respectively to obtain N videos under the video category.
[0026] In this embodiment, the execution subject of the above video batch generation method is not limited. For example, this method can be implemented as a service product, and this service product can adopt a client-server architecture. In the case of adopting a client-server architecture, on the one hand, a video generation page for batch video generation is provided to the user on the client to receive the user's input operations for the number of videos and the video category, and then a batch video generation task is initiated to the server. On the other hand, in response to the client initiating a batch video generation task through the video generation service page on the server, the server batch generates videos through resources such as the server's computing resources, network bandwidth, and storage resources, which is beneficial to improving the speed of video generation.
[0027] For another example, as the processing power of the corresponding hardware device of the client increases, the above method can also be executed by the client. The client can provide a video generation page to the user and respond to the user's input operations for the number of videos and the video category on the video generation page to generate a batch video generation task, and then batch generate videos through resources such as the client's own computing resources and storage resources. Among them, when the batch generation task is deployed to be executed on the client side, there is no need to transmit data to the server, which can save network latency.
[0028] In this embodiment, the batch video generation task includes the number of videos N and the video category. N is an integer greater than or equal to 2. The video category is used to describe the expression theme of the video content of the batch video generation, and the specific implementation of the video category is not limited. For example, it includes but is not limited to: product introduction, teaching explanation, knowledge popularization, beauty and skincare, etc. Optionally, the video category can be implemented as a single-level video category. For example, it can be implemented as business registration or legal consultation, etc. Optionally, the video category can also be implemented as a multi-level video category. For example, the first-level video category can be business registration; the second-level video category of this first-level video category can be cleaning or food operation, etc.
[0029] In this embodiment, according to the batch generation task, N video instance identifiers are generated. The N video instance identifiers are different from each other, and each video instance identifier can be used to uniquely represent a video to be generated, so as to facilitate tracking of the required video materials and video templates for the video to be generated. That is to say, this video instance identifier can also be used as the unique identity identifier of the relevant content (such as video materials) of the video to be generated.
[0030] Among them, the method for generating N video instance identifiers is not limited. For example, it includes but is not limited to numbers and strings, etc. For example, it can be an increasing sequence starting from any integer with a fixed step size to obtain N integers as the N video instance identifiers; or, it can also be N preset strings, etc.
[0031] In this embodiment, a set of video materials related to the video category is generated for each video instance identifier, and this set of video materials is used to generate the video corresponding to this video instance identifier. Among them, each set of video materials contains video materials of at least one material type. In some embodiments of this application, the video materials of one material type are simply referred to as one kind of video material.
[0032] Among them, the material type refers to the media resource type that constitutes the video, including but not limited to: material types such as audio, video, background image, subtitle, digital human, etc.
[0033] In this embodiment, the implementation manner of generating a set of video materials related to the video category for each video instance identifier is not limited.
[0034] In an optional implementation manner, for any video instance identifier, video materials can be randomly extracted from multiple material types stored in the basic material library, and at least one extracted video material is used as a set of video materials for this video instance identifier. Among them, the basic material library stores multiple video materials under multiple video categories.
[0035] In another optional implementation manner, for any video instance identifier, the semantic similarity between at least one video material in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the video materials that meet the similarity conditions among the multiple similarity information are used as a set of video materials for this video instance identifier.
[0036] Furthermore, in this embodiment, multiple materialized templates obtained by variable processing of the initial video template are acquired. Optionally, the initial video template can be implemented as an MLT (Media Lovin' Toolkit) template. The MLT template is an open-source framework for multimedia processing, which is a file based on the XML format and records all parameters of video editing, such as video clips on the timeline, audio tracks, filter effects, transitions, etc., and can be used for video editing.
[0037] In the embodiments of the present application, the initial video template may be an existing video production framework. Among them, the video production framework may be preset or exported after video editing by an editing software. In the embodiments of the present application, the specific method of obtaining the initial video template is not limited. For example, the initial video template may be a preset general video template provided by the system by default, or a custom video template created by the user according to requirements, or a video template imported from other external sources.
[0038] Among them, the initial video template includes the rendering rules of various video materials required for generating a video, as well as the hierarchical relationship between various video materials. In the embodiments of the present application, the rendering rule of each video material is used to describe the rendering logic followed by this video material during the video rendering process, so as to achieve the expected rendering effect of this video material in the generated video. Among them, the hierarchical relationship formed between various video materials in the initial video template refers to the layer stacking order of various video materials during the video generation process in the initial video template, which is used to determine the front-to-back covering relationship and occlusion logic of various video materials during video generation. For example, in generating a video, the subtitle can be located in the relatively upper layer of the layer, the background image can be located in the relatively lower layer of the layer, the digital human can be superimposed on the upper layer of the background image, and under the layer of the subtitle.
[0039] In this embodiment, part of the variable processing of the initial video template is reflected in splitting the initial video template structurally and templatizing the rendering parameter information by taking the material type as the splitting variable, so as to obtain a variety of materialized templates. Among them, there is a corresponding relationship between the materialized template and the material type. For example, the material types include audio, video, background image, subtitle, digital human and other material types. Then the types of the corresponding materialized templates include but are not limited to: audio materialized template, video materialized template, background image materialized template, subtitle materialized template, digital human materialized template, etc. In other words, each material type corresponds to a materialized template, and each materialized template is used to describe the rendering rule of this video material.
[0040] In this embodiment, the timing of the variable processing is not limited. For example, the initial video template can be variablized in advance. Another example is that the initial video template can also be variablized dynamically. For the detailed content on how to perform the variable processing, reference can be made to the subsequent embodiments.
[0041] In this embodiment, the materialized templates after variable processing can be modularly reorganized. For each video instance identifier, at least one target materialized template is determined from a variety of materialized templates according to the material types in a group of video materials corresponding to the video instance identifier. There is a corresponding relationship between the material types in this group of video materials and the materialized templates.
[0042] For example, in the case where the types of materials in this group of video materials include audio, video, background images, and subtitles, the target materialization template can include the target materialization templates corresponding to audio, video, background images, and subtitles respectively; for another example, in the case where the types of materials in this group of video materials include audio, video, background images, subtitles, and digital humans, the target materialization template can include the target materialization templates corresponding to the types of materials of audio, video, background images, subtitles, and digital humans respectively.
[0043] Furthermore, in the case of obtaining at least one target materialization template corresponding to a group of video materials, based on the hierarchical relationship between multiple video materials included in the initial video template, the at least one target materialization template is combined to obtain the target video template corresponding to the video instance identifier. Among them, the target video template is used to describe the rendering rules and hierarchical relationship of a group of video materials.
[0044] Among them, the hierarchical relationship between multiple video materials included in the initial video template can represent the layer stacking order of multiple materialization templates and can be used to organize and integrate the target materialization templates, thereby forming a target video template with a clear hierarchical relationship and corresponding rendering rules for video materials. Among them, the target video template describes the hierarchical relationship of a group of video materials and is consistent with the hierarchical relationship of this group of video materials in the initial video template.
[0045] Among them, in this embodiment, the N video instance identifiers respectively correspond to their own target video templates, that is to say, each group of video materials has its own corresponding target video template, enriching the types of target video templates used for batch video generation. Among them, based on the rendering rules described by the N target video templates, video generation processing is performed on the grouped video materials corresponding to the N video instance identifiers, so that the style of each video generated in batch matches the respective video template used, improving the diversification degree of the video content generated in batch.
[0046] In the case of obtaining the target video template corresponding to the video instance identifier, video generation processing is performed according to the target video templates respectively corresponding to the N video instance identifiers and a group of video materials to obtain N videos under the video category. Video generation processing refers to the process of filling and rendering video materials for the target video template based on the target video template corresponding to each video instance identifier and a group of video materials to obtain the video corresponding to the video instance identifier. For example, for the N video instance identifiers, a group of video materials respectively corresponding to the N video instance identifiers can be filled into the N target video templates to obtain the filled target video templates, and then the filled target video templates can be rendered to obtain the videos under the video category.
[0047] In an alternative embodiment, batch video generation processing can be performed on the video generation corresponding to each of the N video instance identifiers. M videos are generated in parallel for each batch, where M is less than N and M is an integer. After M videos are generated, the subsequent M videos are processed until all the videos corresponding to the N video instance identifiers are processed. Through batch video generation processing, the resource utilization rate of the server is improved, and at the same time, the server pressure overload caused by high concurrency is avoided.
[0048] In the embodiments of the present application, in batch video generation, materialized templates of various material types obtained by variable processing of an initial video template are used to reorganize the materialized templates to obtain a video template required for generating a video. Furthermore, a video instance identifier is generated for each video, and each video instance identifier is bound to a corresponding set of video materials. According to the set of video materials corresponding to the video instance identifier, a target materialized template for reorganization is determined from various materialized templates, and in combination with the hierarchical relationship between the materialized templates provided by the initial template, the target materialized template is organized and integrated to obtain a video template corresponding to each video instance identifier for video generation processing of a corresponding set of video materials for each video instance identifier. Since the video template can be obtained by personalized reorganization in combination with the video materials for generating the video, the flexibility of video generation is improved. In addition, the reorganized video template has a corresponding relationship with the video instance identifier, which enriches the types of video templates to a certain extent, so that the style of each batch-generated video matches the adopted video template respectively, and the diversification degree of the batch-generated video content is improved.
[0049] The following introduces the method of variable processing.
[0050] In an alternative embodiment, the steps of variable processing of the initial video template include: obtaining the initial video template, which includes the rendering rules of various video materials required for generating a video and the hierarchical relationship between various video materials; parsing the initial video template with the material type as the splitting variable to obtain information segments corresponding to various material types; respectively extracting the rendering parameter information corresponding to various material types from the information segments corresponding to various material types; and templatizing the rendering parameter information corresponding to various material types to obtain various materialized templates, and each materialized template is used to describe the rendering rule of a video material.
[0051] Further, after obtaining the initial video template, the initial video template is parsed using the material type as the splitting variable. Herein, the material type refers to different types of video materials, which is a classification criterion for distinguishing different video materials, and the splitting variable refers to the independent information segments obtained by dividing the initial video template according to the material type when parsing the initial video template. For example, if the initial video template contains two types of video materials, i.e., images and audio, the initial video template will be split into independent information segments of images and audio according to the material type as the classification criterion, and subsequent processing will be performed separately.
[0052] Herein, the information segment refers to the information segment extracted from the initial video template corresponding to the material type, and the information segment may include the rendering parameter information of the corresponding material type. For example, an information segment may include rendering parameter information such as path information, playing time, and special effect application for the material type.
[0053] In the embodiment of the present application, the rendering parameter information is used to describe at least one attribute of a material file of a certain material type. Each attribute can be used as a rendering parameter, and the attribute value of each attribute in the initial video template can be used as the default parameter value of the corresponding rendering parameter. The rendering parameters with default parameter values form a rendering parameter information corresponding to the material file of this material type. Herein, the "material file" refers to various resource files used in the materialized template, mainly including various video materials, such as pictures, videos, and audio, etc. The rendering parameter information can be used to control the rendering logic of the material file so that the rendering effect presented by the material file in the finally generated video can be obtained. The rendering parameter information quantifies the description of the attributes of the material file, so that the attributes of the video material are converted into the attributes of the rendering parameter information, and the rendering effect of the video material in the finally generated video is controlled through the rendering parameter information. In other words, the attributes of the video material are described as the corresponding rendering parameters by the rendering parameter information, and the rendering parameters with default parameter values define the attribute values of a certain attribute of the material file.
[0054] Herein, the rendering parameter information can be extracted separately from the information segments corresponding to various material types. The rendering parameter information corresponding to each material type includes at least one rendering parameter and its corresponding default parameter value, and each rendering parameter is used to describe the attribute of a material file of a certain material type. In the embodiment of the present application, the specific attributes corresponding to the rendering parameters of various material types are not limited.
[0055] In the embodiments of the present application, templatization can convert the rendering parameter information corresponding to various material types from the rendering parameters with fixed configuration default parameter values into a set of variable parameters with dynamically replaceable parameter values. Among them, by templatizing the rendering parameter information corresponding to various material types, multiple materialized templates corresponding to various material types can be obtained. Each materialized template is used to describe the rendering rules of a material type, and the rendering rules include at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values, so as to control the rendering effect of video generation.
[0056] In the embodiments of the present application, the materialized template is the result of templatizing the rendering parameter information corresponding to various material types. It includes the placeholder for the path information of the material file, as well as the variable parameter and the multiple candidate parameter values associated with it. The materialized template describes the rendering rules of the material file of the corresponding material type in the generated video. The rendering rules are used to describe the rendering logic followed by this type of video material during the video rendering process, so as to control the expected rendering effect of the presentation of this type of material file in the generated video. Among them, the materialized template formed by converting the rendering parameter information from the rendering parameters with fixed configuration and default parameter values into a set of dynamically replaceable variable parameters can flexibly adjust the rendering effect of the material file in the generated video in different scenarios by assigning different candidate parameter values to the variable parameters. Among them, the rendering parameter information is templatized to form a variety of materialized templates that can be modularly recombined. For the materialized templates of different material types, different target video templates can be recombined, and then diverse videos can be batch-generated without the need to design and adjust the individual video templates for each video, which not only saves time and human resources, but also improves the flexibility and content diversity of video template generation, and thus realizes the efficient batch generation of videos with diverse styles.
[0058] For example, for the same materialized template, by assigning different candidate parameter values to the variable parameters, the attribute values corresponding to various attributes of the material type corresponding to the materialized template can be flexibly set, so as to realize the refined control of the rendering rules of the material file. The materialized template improves the flexibility of video generation, can significantly improve the efficiency and quality of video generation, and at the same time meets diverse creative requirements.
[0059] In the embodiments of the present application, there is no limitation on the specific content of the candidate parameter values associated with the variable parameters obtained after templatizing the rendering parameter information corresponding to various material types.
[0060] Among them, by associating multiple candidate parameter values with the variable parameters, the variable parameters can be flexibly adjusted according to requirements, improving the flexibility and diversity of video template generation, and thus realizing the efficient batch generation of videos with diverse styles.
[0061] In an optional embodiment, taking the material type as a splitting variable, the initial video template is parsed to obtain information segments corresponding to multiple material types, including: loading the XML document corresponding to the initial video template, where the XML document includes a root element and multiple non-root elements connected to the root element, and multiple specific elements are included in the multiple non-root elements, and each specific element is used to describe the rendering rule of a material type; starting from the root element, traversing the non-root elements in the XML document to identify multiple specific elements; extracting multiple information segments where the multiple specific elements are located as information segments corresponding to multiple material types. Among them, the XML document of the initial video template can be loaded into memory to form a parseable tree structure. In the embodiment of the present application, the XML document can be loaded into memory through a DOM (Document Object Model) parser to form a tree structure parsing method.
[0062] Among them, the tree structure includes a root element and non-root elements. The root element is the top-level node of the XML document, and the non-root elements are all directly or indirectly nested under the root element. An XML document has one and only one root element, and the root element is the first element of the XML document and can be used as the starting point of the XML document.
[0063] Among them, the non-root elements are child elements directly or indirectly nested under the root element and are used to divide different information segments of the initial video template according to the material type. Multiple specific elements are included in the non-root elements. Starting from the root element, traversing the non-root elements in the XML document can identify multiple specific elements. A specific element is an element that can directly describe the rendering rule of a certain type of material, and each specific element corresponds to a material type.
[0064] By extracting these information segments, the system can quickly identify the attributes of the material file and perform rendering according to their attribute values. This design of extracting information segments significantly improves the flexibility and diversity of video generation, meets the diverse creation needs, and at the same time provides a technical basis for efficient batch video generation.
[0065] In an optional embodiment, starting from the root element, non-root elements in the XML document are traversed to identify multiple specific elements, including: S1, starting from the root element, traverse non-root elements in the XML document; S2, for the currently traversed non-root element, obtain the element tag included in the currently traversed non-root element; S3, if the element tag is a specific tag, determine whether the currently traversed non-root element contains sub-elements; S4, if the currently traversed non-root element contains sub-elements, use the sub-elements as the currently traversed non-root element and return to execute step S2; S5, if the element tag is a non-specific tag, continue to the next non-root element and return to execute step S2; S6, if the currently traversed non-root element does not contain sub-elements, use the currently traversed non-root element as a specific element.
[0066] Among them, in step S1, starting from the root element of the XML document, non-root elements are accessed one by one for traversal. In the embodiments of the present application, the specific implementation strategy of traversal is not limited. For example, the traversal can be implemented using a depth-first search algorithm or a breadth-first search algorithm. In the embodiments of the present application, taking the depth-first search algorithm as an example, the process of traversing non-root elements in the XML document to identify multiple specific elements is described in detail.
[0067] Next, execute step S2. For the currently traversed non-root element, obtain the element tag included in the currently traversed non-root element. Among them, in the XML document, the element tag is the identifier within angle brackets (<>), which is used to mark the type and semantic meaning of the element. Different material types are distinguished by the name of the element tag; the specific position of the material is located through the hierarchical relationship of the element tags.
[0068] After obtaining the element tag included in the currently traversed non-root element, determine whether the element tag is a specific tag. Among them, the specific tag can be a predefined set of key tags, which represent the element tags that need to be specially processed in the XML document and are used to identify the material types to be extracted.
[0069] Next, execute step S3 or S5. If step S5 is executed, that is, the element tag is a non-specific tag, continue to the next non-root element and return to execute step S2.
[0070] If step S3 is executed, that is, the element tag is a specific tag, continue to determine whether the currently traversed non-root element contains sub-elements. Among them, the sub-elements can be other elements nested inside the currently traversed non-root element.
[0071] Next, execute step S4 or S6. If step S4 is executed, that is, the currently traversed non-root element contains child elements, then use the child elements as the currently traversed non-root element, and return to execute step S2. If step S6 is executed, that is, the currently traversed non-root element does not contain child elements, then use the currently traversed non-root element as a specific element. Among them, step S4 is the recursive processing when there are child elements. Set the child elements of the currently traversed non-root element as the new traversal starting point, and re-execute steps S2 - S6 for each child element. After the recursion ends, continue to traverse other non-root elements. Step S6 is the end collection when there are no child elements, and use the currently traversed non-root element as a specific element.
[0072] In an optional embodiment, rendering parameter information corresponding to various material types is respectively extracted from information segments corresponding to various material types, including: for each information segment, extracting path information of a material file and at least one attribute value of the material file from the information segment, where the attribute value is used to render the material file; reading the material file according to the path information, and determining the material type described by the information segment according to the extension of the material file; taking at least one attribute to which at least one attribute value belongs as at least one rendering parameter, and taking at least one attribute value as the default parameter value of at least one rendering parameter, so as to obtain the rendering parameter information corresponding to the material type described by the information segment.
[0073] Among them, the information segment is a part related to the material type extracted from the XML document, and contains the path information of the material file and at least one attribute value of the material file. Among them, the material file can be used to construct various media files for the finally generated video. In the embodiments of the present application, the material file may include but is not limited to: files in formats such as MP4 and AVI, which are video files containing dynamic images and audio and are used to display a series of continuous pictures; files in formats such as WAV and MP3, which are audio files providing background music, narration, or special effect sounds, etc.
[0074] Among them, at least one attribute value of the material file extracted from the information segment is the specific value corresponding to the attribute, so as to describe the specific state of the attribute. The attribute can describe the rendering effect of the material file. For example, for a material file that is a video file, the attribute can be video duration, video speed, filter effect, etc. If the attribute is video duration, the attribute value can be 2 minutes, 1 hour, 1 day, etc.
[0075] In an alternative embodiment, the rendering parameter information corresponding to the material type described by the information segment includes at least one rendering parameter and the default parameter values of at least one rendering parameter. The default parameter value can be the attribute value of the attribute corresponding to the rendering parameter before quantization. Among them, at least one attribute for extracting at least one attribute value of the material file in the information segment can be used as at least one rendering parameter, and at least one attribute value can be used as the default parameter value of at least one rendering parameter to obtain the rendering parameter information. That is to say, a key-value pair set is formed by the rendering parameter and the default parameter value to control the rendering logic of the material file in the video to achieve the expected rendering effect.
[0076] In an alternative embodiment, the rendering parameter information corresponding to each of multiple material types is templated to obtain multiple material templates, including: for each material type, at least one variable parameter is selected from the rendering parameter information corresponding to the material type; multiple candidate parameter values are associated with at least one variable parameter; and a material template corresponding to the material type is generated according to the multiple candidate parameter values associated with at least one variable parameter. Among them, the variable parameter is obtained by variable processing of the rendering parameter with a default parameter value. Each variable parameter is associated with multiple candidate parameter values, and different candidate parameter values correspond to different rendering logics to produce different rendering effects in the generated video, which can be used to generate diverse material templates. For each material type, at least one variable parameter is selected from its corresponding rendering parameter information, and these variable parameters can vary within the adjustment range to obtain different material templates. For example, when the material file is a video material, the playback speed and filter effects can be selected as variable parameters. The variable parameter is associated with multiple candidate optional values to limit the adjustment range of the variable parameter, ensuring that the material file of the generated video meets the design requirements and avoiding invalid settings.
[0077] In an alternative embodiment, when selecting at least one variable parameter from the rendering parameter information corresponding to the material type, two selection methods are provided: one is to use all rendering parameters as variable parameters, and the other is to select some rendering parameters as variable parameters according to the weight value. Among them, when choosing to use all rendering parameters as variable parameters, there is no need to compare the weight values, and all rendering parameters can be adjusted as variable parameters.
[0078] In an optional embodiment, the manner of selecting some rendering parameters as variable parameters according to weight values may be to select at least one variable parameter from the rendering parameter information corresponding to the material type, including: configuring the weight values of each rendering parameter for the material type in advance; parsing each rendering parameter from the rendering parameter information corresponding to the material type, and selecting at least one rendering parameter with a weight value greater than a set weight threshold from each rendering parameter as at least one variable parameter. Among them, the weight value can be an importance score assigned to each rendering parameter, used to quantify the influence degree of the rendering parameter on the rendering effect of the material file in video generation, and at least one rendering parameter that has a greater impact on user perception or application goals can be selected as at least one variable parameter. The weight threshold is a predefined critical value. By screening the rendering parameters corresponding to the weight values higher than the weight threshold, the user can select the parameters with weight values higher than the weight threshold as variable parameters, avoiding taking too many irrelevant rendering parameters as variable parameters to prevent problems such as configuration conflicts or rendering logic confusion.
[0079] In the embodiments of the present application, the manner of pre-configuring the weight values of each rendering parameter is not limited. For example, the weight values of each rendering parameter can be pre-configured according to empirical annotations or automatically generated through user behavior data analysis.
[0080] Optionally, selecting at least one variable parameter from the rendering parameter information corresponding to the material type can also be randomly selected according to the set number of variable parameters. In an optional embodiment, according to multiple candidate parameter values associated with at least one variable parameter, a materialized template corresponding to the material type is generated, including: adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to a preset template file, and adding multiple candidate parameter values associated with at least one variable parameter to the preset template file; and adding a placeholder for carrying the material file corresponding to the material type in the preset template file to obtain a materialized template corresponding to the material type. Among them, the preset template file is a basic video template containing basic configurations, including the basic structure of the video template and some preset rendering parameters. Among them, adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to the preset template file can ensure that the finally obtained materialized template has complete rendering parameter information.
[0081] In this embodiment, on the basis of adding each rendering parameter corresponding to the material type and the default parameter values of each rendering parameter to the preset template file, multiple candidate parameter values associated with at least one variable parameter are added to the preset template file. Among them, the candidate parameter values provide multiple choices, allowing the user or the system to select different candidate parameter values according to needs.
[0082] In an optional embodiment, based on adding multiple candidate parameter values associated with at least one variable parameter to a preset template file, a placeholder for carrying a material file corresponding to a material type can be added to the preset template file to obtain a materialized template corresponding to the material type. The placeholder is a reserved position for filling the material file and supports dynamic replacement. For example, the path information of the actual material file corresponding to the material type can be filled. Through the placeholder, different material files can be flexibly replaced without modifying the template structure.
[0083] In the above embodiment, by integrating the rendering parameters of the material type and their default parameter values into the preset basic template, the integrity of the rendering parameter information of the generated materialized template is ensured. At the same time, by associating multiple candidate parameter values with the variable parameter, dynamic selection of the parameter values of the variable parameter is realized to flexibly adjust the video style. The embedding of the placeholder further decouples the parameter configuration of the initial video template from the material file and supports dynamically replacing the material file without modifying the template structure. Thus, it is allowed to assign corresponding candidate parameter values to the variable parameter according to the application requirements to obtain a freely combined materialized template, ultimately significantly improving the flexibility and diversity of video template generation, and then efficiently batch-producing video content with various styles, solving the problems such as cumbersome parameter configuration of video templates, complex adaptation process, and serious homogenization of generated video content in traditional video production.
[0084] Based on obtaining the above multiple materialized templates, batch video generation or single video generation can be performed. For the batch video generation scenario, using the multiple materialized templates provided by the embodiments of the present application, multiple videos with diverse styles and different contents can be generated. A specific implementation manner of batch video generation based on the multiple materialized templates provided by the embodiments of the present application will be described in detail below.
[0085] In the embodiments of the present application, the initial video template is variable-processed to obtain various materialized templates. During batch video generation, for each video instance identifier, at least one target materialized template is determined from the various materialized templates according to the material type in a group of video materials corresponding to the video instance identifier.
[0086] In an optional embodiment, when determining at least one target materialization template from multiple materialization templates according to the material types in a group of video materials corresponding to the video instance identifier, it includes: identifying at least one material type included in the group of video materials corresponding to the video instance identifier; selecting at least one initial materialization template from multiple materialization templates according to at least one material type included in the group of video materials, with each material type corresponding to one initial materialization template; and adjusting the parameters of at least some of the selected at least one initial materialization templates to obtain at least one target materialization template.
[0087] In this embodiment, the materialization templates included in the multiple materialization templates obtained by variable processing of the initial video template are referred to as initial materialization templates. In the case of determining at least one initial materialization template corresponding to a certain group from multiple initial materialization templates, the rendering rules described by the initial materialization templates can be referred to as initial rendering rules. Furthermore, the parameters of at least some of the at least one initial materialization templates are adjusted to obtain at least one target materialization template. Each initial materialization template obtains a corresponding target materialization template after parameter adjustment. The target materialization template after parameter adjustment includes a target rendering rule, and the target rendering rule is different from the initial rendering rule described by the initial materialization template.
[0088] In this embodiment, since the initial materialization template is obtained by variable processing, the variable processing can convert the fixed rendering parameters in the initial video template into variable parameters that can be dynamically assigned values to achieve flexible configuration of the template content. Different values of the optional parameters result in different rendering rules. Optionally, each initial materialization template includes at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values and a default parameter value.
[0089] In this embodiment, adjusting the parameters of the initial materialization template is to obtain multiple target materialization templates with different candidate parameter values by assigning different candidate parameter values to the variable parameters of the initial materialization template. Among them, the target materialization template for a certain material type can be used for rendering video materials in groups with different video instance identifiers. Different parameter values of the variable parameters represent different rendering rules described by the target materialization templates of the same material type, so that the rendering results of the batch-generated videos are different, forming differences between different videos and improving the richness of video content. The following will introduce how to adjust the parameters of the initial materialization template.
[0090] In an alternative embodiment, when adjusting parameters of at least part of at least one initial materialized template to obtain at least one target materialized template, it includes: determining the number of templates with parameters to be adjusted, where the number of templates is less than or equal to the number of at least one initial materialized template; selecting the initial materialized templates to be adjusted from at least one initial materialized template according to the number of templates to be adjusted; determining the variable parameters to be adjusted from the initial materialized templates to be adjusted; randomly determining target parameter values from multiple candidate parameter values associated with the variable parameters to be adjusted; and assigning the target parameter values to the variable parameters to be adjusted to obtain the target materialized template.
[0091] In this embodiment, there is no limitation on the method for determining the number of templates with parameters to be adjusted. For example, if parameter adjustment is performed on each initialized video template, the number of templates to be adjusted is the number of at least one initial materialized template. Another example is that, according to the random number generation algorithm, a random integer within a preset range is generated as the number of templates to be adjusted, and the preset range means less than or equal to the number of at least one initial materialized template. There is no limitation on the random number generation algorithm, including but not limited to: the Linear congruential generator (LCG) and the Mersenne Twister, etc.
[0092] Furthermore, select the initial materialized templates to be adjusted from at least one initial materialized template according to the number of templates to be adjusted. In an alternative embodiment, select the initial materialized templates to be adjusted according to the priority of the material types and the number of templates to be adjusted. Here, the priority of the material type refers to the importance degree of the material type to the presentation effect of the generated video. For example, the subtitle material type generally has a lower importance degree to the presentation effect, so the priority of the subtitle material type can be set to a lower priority; relatively speaking, the priority of the background image can be higher than that of the subtitle, so it can be set to a medium priority; and, the digital human has a higher importance degree to the presentation effect, so it can be set to a higher priority. Then, according to the high or low priority, preferentially select the templates with parameters to be adjusted from the initial material templates with a higher priority of the material type. If the number of these initial material templates is less than the previously determined number of templates to be adjusted, then it can be further selected from the initial materialized templates with a lower priority of the material type. The finally selected number of templates to be adjusted is less than or equal to the number of at least one initial materialized template.
[0093] Further, determine the variable parameters to be adjusted from the initial materialized template to be adjusted. In an alternative embodiment, all variable parameters in the initial materialized template to be adjusted can be used as the variable parameters to be adjusted. In another alternative embodiment, the target variable parameters in the initial materialized template to be adjusted are used as the optional parameters to be adjusted, and the target variable parameters are pre-selected optional parameters.
[0094] Furthermore, randomly determine the target parameter values from multiple candidate parameter values associated with the variable parameters to be adjusted. In some embodiments, among the at least one initial materialized template corresponding to each of the N video instance identifiers, there are the same variable parameters to be adjusted, and the multiple candidate parameter values associated with the variable parameters to be adjusted can be randomly used as the target parameter values for each of the N video instance identifiers respectively, so that the optional parameter values of the N video instance identifiers are as different as possible.
[0095] Further, assign the target parameter values to the variable parameters to be adjusted to obtain the target materialized template.
[0096] In the case of obtaining the target materialized template, combine the target materialized templates to obtain the target video template. The combination method is not limited in this embodiment. The following provides two combination methods, but is not limited thereto.
[0097] In an alternative embodiment, according to the hierarchical relationship between multiple video materials included in the initial video template, generate a basic video template, which serves as the framework of the target video template and includes multiple blank structure positions corresponding to multiple video materials. Among them, the blank structure position is a placeholder for the target materialized template preset in the basic video template, used to identify the position where the target materialized template can be inserted, and each blank structure position corresponds to the filling of the target materialized template of one material type. The positional relationship between multiple blank structure positions reflects the hierarchical relationship between multiple video materials. As described in the above embodiment, the hierarchical relationship represents the front-to-back stacking order of multiple materials in the generated video. In some embodiments, this hierarchical relationship is extracted from the initial video template; or, it can also be preset, preset based on at least one target materialized template, that is to say, the hierarchical relationship can be set as needed. For example, the structure position corresponding to the subtitle is located in the upper layer, the structure position corresponding to the digital human is located in the middle layer, and the structure position corresponding to the background image is located in the lower layer. Further, insert at least one target materialized template into the corresponding blank structure position in the basic video template to obtain the target video template corresponding to the video instance identifier.
[0098] In another alternative embodiment, according to at least one target materialization template, the structural positions where the rendering rules of video materials of the same material type in the initial video template are overwritten to obtain the target video template corresponding to the video instance identifier. The difference between the structural position and the above-mentioned blank structural position is that the rendering rule of each video material in the initial video template occupies a structural position, and the blank structural position is empty. Among them, the positional relationship between the structural positions reflects the hierarchical relationship between multiple video materials.
[0099] In the case of obtaining the target video template, video generation processing is performed according to the target video template corresponding to each video instance identifier and a set of video materials. In an alternative embodiment, when performing video generation processing according to the target video templates corresponding to N video instance identifiers and a set of video materials to obtain N videos in the video category, it includes: for each video instance identifier, filling the set of video materials corresponding to the video instance identifier into the target materialization template in the target video template corresponding to the video instance identifier; for the filled target video template, rendering the set of video materials according to the rendering rules and hierarchical relationship of the set of video materials described in the filled target video template to obtain a video in the video category.
[0100] Among them, the target materialization template includes at least one placeholder corresponding to the material type for filling video materials of the material type.
[0101] In an alternative embodiment, generating a set of video materials related to the video category for each video instance identifier includes: obtaining a set of video material description information related to the video category for each video instance identifier, and each set of video material description information includes description information of multiple video materials; for each video instance identifier, according to the set of video material description information corresponding to the video instance identifier, calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously uploading the multiple video materials corresponding to the video instance identifier to the content delivery network; correspondingly, before performing video generation processing according to the target video templates corresponding to N video instance identifiers and a set of video materials to obtain N videos in the video category, it also includes: in response to a batch video generation trigger event, respectively obtaining multiple sets of video materials corresponding to N video instance identifiers from the content delivery network. Hosting video materials through the content delivery network reduces the storage pressure on the server side, enabling the server side to efficiently render a large number of videos and improving the user experience.
[0102] In this embodiment, the description information of each group of video materials is used to generate video materials corresponding to the video instance identifier thereof. Among them, each group of video materials includes video materials of multiple material types. The material type refers to the type of different video elements that make up the video content, including but not limited to: audio, video, background image, subtitle, digital human and other material types.
[0103] In this embodiment, the implementation manner of generating the description information of a group of video materials related to the video category for each video instance identifier is not limited.
[0104] In an alternative implementation manner, for any video instance identifier, the description information of video materials can be randomly extracted from multiple material types stored in the basic material library, and the description information of multiple video materials extracted is used as the description information of a group of video materials for any video instance identifier. Among them, the basic material library stores the description information of multiple video materials under multiple video categories.
[0105] In another alternative implementation manner, for any video instance identifier, the semantic similarity between the description information of multiple video materials in the basic material library and the video category is calculated respectively to obtain multiple similarity information; the description information of the video materials that meet the similarity condition among the multiple similarity information is used as the description information of a group of video materials for this video instance identifier.
[0106] In yet another alternative implementation manner, for any video instance identifier, according to the video category, a material description information generation model is called, and this model is used to generate the description information of multiple video materials related to this video category. Among them, this material description information generation model is trained by combining the description information of sample video materials of a large number of different sample video categories and different material types. By learning the semantic correlation between the description information of sample video materials of different sample video categories and different material types, this model can specifically combine different video categories to generate the description information of multiple video materials related to them.
[0107] Further, for each video instance identifier, according to the description information of a group of video materials corresponding to this video instance identifier, multiple material generation models based on artificial intelligence are called to generate multiple video materials corresponding to this video instance identifier, and the multiple video materials corresponding to this video instance identifier are synchronously uploaded to the content distribution network.
[0108] Among them, one material generation model can generate at least some of the multiple video materials. For example, one material generation model can generate one video material. The following also takes this as an example for illustration, but is not limited thereto.
[0109] In this embodiment, when generating video materials with video instance identifiers, according to multiple video materials corresponding to each video instance identifier, a target video template corresponding to each video instance identifier is determined. The target video template corresponding to each video instance identifier is used to describe the rendering rules and hierarchical relationships of the multiple video materials corresponding to this video instance identifier. The determination method of the target video template corresponding to each video instance identifier can refer to the above embodiments and will not be elaborated here.
[0110] In this embodiment, in response to a batch video generation trigger event, according to N video instance identifiers, multiple video materials corresponding to the N video instance identifiers are respectively obtained from the content delivery network; video generation is performed according to the multiple video materials and target video templates respectively corresponding to the N video instance identifiers to obtain N videos under the video category.
[0111] In this embodiment, the specific implementation of the batch video generation trigger event is not limited and can be flexibly configured according to actual application requirements. For example, it can be when all the multiple materials corresponding to the N video instance identifiers are generated; or, it can also be when multiple materials corresponding to each video instance identifier are generated. In this case, video generation will be performed for the multiple video materials corresponding to this video instance identifier; or the batch video generation trigger event can also be a preset trigger time. For example, it can be after a period of time in response to an input operation on the video quantity and video category on the video generation page. For example, it can be 2 hours, 1 day, or 1 week, etc. The time span is not limited.
[0112] It should be noted that each video instance identifier corresponds to the generation of one video. The N video instance identifiers can correspond to the generation of N videos. When performing video generation for the multiple video materials respectively corresponding to the N video instance identifiers, the generation process of the video corresponding to each video instance identifier is asynchronous, and the video generations corresponding to the respective video instance identifiers do not affect each other, so as to improve the generation efficiency of the N videos.
[0113] Further optionally, the description information of each group of video materials includes but is not limited to: subtitle description information, audio type description information, digital human description information, and background image description information. Such as Figure 2As shown in the figure, when calling multiple material generation models based on artificial intelligence according to a set of video material description information corresponding to the video instance identifier to generate various video materials corresponding to the video instance identifier, it includes: calling a generative language model according to the subtitle description information to generate text information to obtain the target subtitle; calling a text-to-speech model according to the target subtitle and the audio type description information to convert the target subtitle into a target audio adapted to the audio type description information; calling a multi-modal model according to the target audio and the digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video of the target digital human based on the target audio description to obtain the green screen video of the target digital human; calling an image generation model according to the background image description information to generate a background image to obtain the target background image.
[0114] In this embodiment, the APIs (Application Programming Interfaces) of multiple material generation models are associated with endpoints. In this embodiment, the video generation page is the presentation of the front-end code of the endpoint. Among them, the endpoint is used for the generation of video materials, the management of batch video generation tasks, and the control of the video generation process.
[0115] Among them, the endpoint includes front-end code and back-end code, that is, the client-server structure is adopted as described in the above embodiment. The front-end code refers to the video generation page built based on the front-end framework. This video generation page runs on the client and is used to interact with the user, receive the user's input operations, and initiate a batch video generation task to the server. The back-end code of the endpoint runs on the server. The back-end code is obtained by API-ifying the code of the existing video editing software. Existing video editing software includes, for example, shortcut, etc. Among them, in the existing video editing software, the front-end UI code and the video rendering function code are highly coupled. In this embodiment, it is split according to functions to decouple the front-end UI of the video editing software from the video rendering function code, and then encapsulate the video rendering function code of the existing video editing software into an independent API, so that the external can call the video rendering function through the standard API method to achieve automated video generation.
[0116] As Figure 2 For the endpoint in the figure, in one example, the front-end code of the endpoint can be a video generation page built based on Astro, which is responsible for receiving callback notifications. For example, when the material generation model finishes processing, it can send a notification to the endpoint to notify that the subsequent process of video generation can continue. The back-end code of the endpoint can be obtained by API-ifying the video rendering code of the video editing software, and is exposed to the outside through the API interface for external calls.
[0117] In an optional embodiment, when calling a generative language model to generate text information according to subtitle description information to obtain target subtitles, it includes: matching corresponding keywords according to the video category, calling a pre-designed prompt template, and filling the keywords of the video category into the prompt template to obtain the prompt for this video category; inputting the prompt of this video category into the generative language model to generate text information and obtain target subtitles related to the video category.
[0118] Among them, the keywords are used to describe the theme content expressed by the video category. Different video categories can correspond to different keywords, and the prompt templates corresponding to different video types can also be different. Among them, the target subtitles include all the text content required for each video. By calling the generative language model, target subtitles can be automatically generated based on the video category without manual intervention, improving the generation efficiency.
[0119] Further optionally, according to the target subtitles and audio type description information, call a text-to-speech model to convert the target subtitles into target audio adapted to the audio type description information. Among them, the target subtitles are used to provide text content, and the audio type description information is used to specify the audio type of the generated target audio. The audio type includes but is not limited to: audio format, audio language, audio tone, audio quality, etc. Any audio type that can be used to specify the sound effect of the target audio is applicable to this embodiment. Based on the target subtitles and audio description information, call a text-to-speech model to generate the corresponding target audio and upload it to the content delivery network to improve the automation degree of audio generation.
[0120] Further, call a multimodal model to generate a green screen video of a virtual human. Among them, a virtual human refers to a virtual character generated based on AI technology, which can synchronously simulate the appearance, voice, lip shape and other behaviors of a real person, and can be used in scenarios such as intelligent customer service, short video production, virtual anchors, etc., but is not limited to this. In this embodiment, by combining various video materials corresponding to each video instance identifier, a video with the effect of a real person speaking can be generated.
[0121] In an optional embodiment, the virtual human can be a pre-recorded real person video or picture. In subsequent embodiments, real person videos and pictures are collectively referred to as video frames, and the number of video frames can be one or more. In this case, the description information for different virtual humans can be implemented as identification information, and the identification information serves as the unique identity identifier of the virtual human and is used to obtain the video frames of the virtual human corresponding to this identification information. In another optional embodiment, the video frames of the virtual human can be dynamically generated based on a multimodal model. In this case, the description information of the virtual human can be the prompt for generating this virtual human, and the description information of the virtual human corresponding to each video instance identifier can be different.
[0122] Further optionally, when calling a multimodal model according to the target audio and digital human description information, selecting a target digital human according to the digital human description information, and generating a green screen video based on the target audio and the target digital human to obtain the green screen video of the target digital human, it includes: obtaining corresponding target digital human video frames based on the digital human description information; calling a multimodal model according to the video frames of the target digital human and the target audio, and performing multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features; wherein, the multi-dimensional speech features include but are not limited to: speech content features and speech emotion features; determining the lip movement control parameters of the target digital human according to the speech content features; determining the facial expression control parameters of the target digital human according to the speech emotion features; determining the body movement control parameters of the target digital human according to the speech content features and the speech emotion features; generating a green screen video of the target digital human based on the lip movement control parameters, expression control parameters and body movement control parameters of the target digital human; the lip movement, facial expression and body movement of the green screen video of the target digital human match the target audio.
[0123] In this embodiment, the speech content features are used to reflect the semantic information in the target audio, such as lexical content, grammatical structure, speech intention and speech rhythm, and are mainly used to drive the lip movement of the digital human to be synchronized with the semantics of the target audio. The speech emotion features are used to reflect the emotional state of the target audio, including but not limited to the strength of tone, speech speed and intonation changes, etc., and are mainly used to drive the facial expressions and body movements of the digital human.
[0124] Further, determine the lip movement control parameters of the target digital human according to the speech content features; determine the facial expression control parameters of the target digital human according to the speech emotion features; determine the body movement control parameters of the target digital human according to the speech content features and the speech emotion features.
[0125] In the case of obtaining the lip movement control parameters, facial expression control parameters and body movement control parameters, based on the lip movement control parameters, expression control parameters and body movement control parameters of the target digital human, drive the target digital human to perform action rendering, and generate the corresponding green screen video of the target digital human, which is convenient for subsequent flexible replacement with the background image. Since the green screen video of the target digital human is generated by controlling the multi-dimensional speech features extracted from the target audio content, it is ensured that the dynamic performance of the lip movement, facial expression and body movement of the target digital human in the green screen video matches the target audio in terms of semantics, timing and emotion, ensuring the precise alignment of the lip movement and improving the realism and visual experience of the generated video.
[0126] Further, in this embodiment, the target audio and the target subtitle are aligned, for example, the timestamp of the target subtitle is inferred and aligned and marked through the timestamp of the speech content in the target audio, so as to ensure that the subsequent display of the target subtitle is completely synchronized with the target audio.
[0127] In this embodiment, one way to generate the background image is to call the text image model to generate the background image according to the background image description information to obtain the target background image. Another way is to directly obtain the pre-generated or photographed background image, which is not limited.
[0128] In an optional embodiment, a corresponding set of video material generation states is maintained for each of the N video instance identifiers, and the generation states of any type of video material in each set of video materials include: a state of preparing to generate, a state of being generated, a state of successful generation, and a state of failed generation. For example, for any video instance identifier, the multiple video materials that need to be generated for the video instance identifier include: target subtitles, target audio, green screen video of a target digital person, and target background image. Then, the generation states of the target subtitles, target audio, green screen video of a target digital person, and target background image can be maintained respectively.
[0129] Among them, for each video instance identifier, mark the multiple video materials corresponding to the video instance identifier as being ready to be generated; when calling multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, if any material generation model returns a generating response message, update the video material that should be generated by the material generation model to the generating state; if any material generation model returns a generating success message, update the video material that should be generated by the material generation model to the generating success state; if any material generation model returns a generating failure message, update the video material that should be generated by the material generation model to the generating failure state. Figure 2 As shown, a subscription service is provided, which can generate a notification of each material model to notify the subscription service of the generation status of the video material of the material model, for example, the generation success status. Then, the subscription service returns a generation success message to the server to inform the server that the corresponding video material has been successfully generated. Figure 2 The notification process of the target digital person is taken as an example, but is not limited to this.
[0130] Further optionally, if the generation status of the video material is updated to a generation failure status, a failure reminder message is output to the user who initiated the input operation, and the failure reminder message includes the material type of the video material updated to the generation failure status and its corresponding video instance identifier. If a regeneration operation triggered by the user is received, the description information of the corresponding video material is obtained according to the video instance identifier corresponding to the video material, and the corresponding artificial intelligence-based material generation model is called to regenerate the corresponding video material.
[0131] In this embodiment, when generating multiple types of video materials corresponding to each video instance identifier, the multiple types of video materials corresponding to each video instance identifier can be synchronously uploaded to the content delivery network, and the access links of the video materials corresponding to each video instance identifier in the content delivery network can be obtained. Continuing with the above example, for any video instance identifier, when the multiple types of video materials corresponding to the video instance identifier include a target subtitle, a target audio, a green screen video of a target digital human, and a target background image, the target subtitle, the target audio, the green screen video of the target digital human, and the target background image are respectively uploaded to the content delivery network, and the access links of the target subtitle, the target audio, the green screen video of the target digital human, and the target background image are obtained.
[0132] Further, in response to a batch video generation trigger event, according to the access links of the multiple types of video materials corresponding to each of the N video instance identifiers in the content delivery network, the multiple types of video materials corresponding to the N video instance identifiers are respectively obtained; and, video generation is performed according to the multiple types of video materials corresponding to each of the N video instance identifiers and a target video template to obtain N videos under the video category.
[0133] Continuing with the above example, in response to a batch video generation trigger event, according to the access links of the target subtitle, the target audio, the green screen video of the target digital human, and the target background image corresponding to each of the N video instance identifiers in the content delivery network, the target subtitle, the target audio, the green screen video of the target digital human, and the target background image corresponding to the N video instance identifiers are respectively obtained; and, video generation is performed according to the target subtitle, the target audio, the green screen video of the target digital human, the target background image, and the target video template corresponding to each of the N video instance identifiers to obtain N videos under the video category.
[0134] Further optionally, the videos corresponding to the N video instance identifiers and the target video template are uploaded to the content delivery network, and the access links of the videos corresponding to the N video instance identifiers and the target video template in the content delivery network are obtained, and the access links are added to the video generation result page; in response to a viewing operation of the video result page, the video generation result page is displayed, and the video generation result page includes the access links of a set of video materials, videos, and the target video template corresponding to at least one video instance identifier among the N video instance identifiers in the content delivery network.
[0135] In this embodiment, the video materials, videos, and target video templates corresponding to each video instance identifier are stored in the content delivery network. The server does not need to store a large number of files, reducing the disk I / O load and ensuring service stability. Further, for the video materials, videos, and target video templates corresponding to each video instance identifier, the access links stored in the content delivery network can avoid data expansion, reduce query pressure, and at the same time support larger-scale data management.
[0136] In this alternative embodiment, in response to a triggering operation for access links of a set of video materials, videos, and / or target video templates corresponding to at least one video instance identifier in a content delivery network, a set of video materials, videos, and / or target video templates corresponding to at least one video instance identifier information is accessed. By hosting videos and video templates through the CDN, the video generation result page can directly load videos from the content delivery network. Compared with pulling videos from the server, it occupies less bandwidth and has a faster loading speed, which can significantly improve performance and is suitable for large-scale video browsing scenarios.
[0137] The detailed implementation manners and beneficial effects of the steps in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0138] It should be noted that the execution subject of each step of the method provided in the foregoing embodiment can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 101 to 103 can be device A; for another example, the execution subject of steps 101 and 402 can be device A, and the execution subject of step 103 can be device B; and so on.
[0139] In addition, in some processes described in the foregoing embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 101 and 102 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are different types.
[0140] Figure 3 This is a schematic structural diagram of an electronic device provided for another exemplary embodiment of the present application. As Figure 3 shown, the device includes: a memory 34 and a processor 35.
[0141] The memory 34 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0142] A processor 35, coupled to a memory 34, is configured to execute a computer program in the memory 34 for: generating a batch video generation task in response to input operations on a video generation page for the number of videos and video categories, the batch video generation task including the number of videos N and video categories, where N is an integer greater than or equal to 2; generating N video instance identifiers according to the batch video generation task, and generating a set of video materials related to the video category for each video instance identifier; obtaining a variety of materialized templates obtained by variable processing of an initial video template, the initial video template including rendering rules for a variety of video materials required for video generation and the hierarchical relationship between the variety of video materials, and each materialized template is used to describe the rendering rules of a video material; for each video instance identifier, determining at least one target materialized template from the variety of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier; combining the at least one target materialized template based on the hierarchical relationship between the variety of video materials included in the initial video template to obtain a target video template corresponding to the video instance identifier, the target video template being used to describe the rendering rules and hierarchical relationship of the set of video materials; performing video generation processing according to the target video templates and the set of video materials corresponding to the N video instance identifiers respectively to obtain N videos under the video category.
[0143] In an alternative embodiment, when the processor 35 performs variable processing on the initial video template, it is specifically configured to: parse the initial video template with the material type as the splitting variable to obtain information segments corresponding to a variety of material types; extract the rendering parameter information corresponding to the variety of material types respectively from the information segments corresponding to the variety of material types; and template the rendering parameter information corresponding to the variety of material types to obtain the variety of materialized templates.
[0144] In an alternative embodiment, when the processor 35 determines at least one target materialized template from the variety of materialized templates according to the material type in the set of video materials corresponding to the video instance identifier, it is specifically configured to: identify at least one material type included in the set of video materials corresponding to the video instance identifier; select at least one initial materialized template from the variety of materialized templates according to the at least one material type included in the set of video materials, where each material type corresponds to one initial materialized template; and adjust the parameters of at least some of the at least one initial materialized template to obtain the at least one target materialized template.
[0145] In an alternative embodiment, the initial materialized template includes at least one variable parameter, and each variable parameter is associated with multiple candidate parameter values and a default parameter value. When the processor 35 adjusts the parameters of at least some of the at least one initial materialized template to obtain the at least one target materialized template, it is specifically configured to: determine the number of templates with parameters to be adjusted, where the number of templates is less than or equal to the number of the at least one initial materialized template; select the initial materialized templates to be adjusted from the at least one initial materialized template according to the number of templates; determine the variable parameters to be adjusted from the initial materialized templates to be adjusted; randomly determine the target parameter value from the multiple candidate parameter values associated with the variable parameters to be adjusted; and assign the target parameter value to the variable parameters to be adjusted to obtain the target materialized template.
[0146] In an alternative embodiment, when the processor 35 combines the at least one target materialized template based on the hierarchical relationship between multiple video materials included in the initial video template to obtain the target video template corresponding to the video instance identifier, it is specifically configured to: generate a basic video template according to the hierarchical relationship between multiple video materials included in the initial video template, where the basic video template includes multiple blank structure positions corresponding to the multiple video materials; the positional relationship between the multiple blank structure positions reflects the hierarchical relationship between the multiple video materials; insert the at least one target materialized template into the corresponding blank structure positions in the basic video template respectively to obtain the target video template corresponding to the video instance identifier; or, overwrite the structure positions where the rendering rules of the video materials of the same material type in the initial video template are located according to the at least one target materialized template to obtain the target video template corresponding to the video instance identifier; where, in the initial video template, the rendering rule of each video material occupies one structure position, and the positional relationship between the structure positions reflects the hierarchical relationship between the multiple video materials.
[0147] In an alternative embodiment, when the processor 35 performs video generation processing on the basis of the target video templates corresponding to N video instance identifiers and a set of video materials to obtain N videos in the video category, it is specifically configured to: for each video instance identifier, fill the set of video materials corresponding to the video instance identifier into the target materialized template in the target video template corresponding to the video instance identifier; for the filled target video template, render the set of video materials according to the rendering rules and hierarchical relationship of the set of video materials described in the filled target video template to obtain one video in the video category.
[0148] In an alternative embodiment, when the processor 35 generates a set of video materials related to the video category for each video instance identifier, it is specifically configured to: obtain, for each video instance identifier, a set of video material description information related to the video category, where each set of video material description information includes description information of multiple video materials; for each video instance identifier, according to the set of video material description information corresponding to the video instance identifier, call multiple material generation models based on artificial intelligence to generate multiple video materials corresponding to the video instance identifier, and synchronously upload the multiple video materials corresponding to the video instance identifier to the content delivery network; correspondingly, before the processor 35 performs video generation processing based on the target video templates and a set of video materials respectively corresponding to N video instance identifiers to obtain N videos under the video category, it is further configured to: in response to a batch video generation trigger event, according to the N video instance identifiers, respectively obtain, from the content delivery network, multiple sets of video materials corresponding to the N video instance identifiers.
[0149] In an alternative embodiment, the description information of each set of video materials includes: subtitle description information, audio type description information, digital human description information, and background image description information; then when the processor 35 calls multiple material generation models based on artificial intelligence according to the set of video material description information corresponding to the video instance identifier to generate multiple video materials corresponding to the video instance identifier, it is specifically configured to: according to the subtitle description information, call a generative language model to generate text information to obtain a target subtitle; according to the target subtitle and the audio type description information, call a text-to-speech model to convert the target subtitle into a target audio adapted to the audio type description information; according to the target audio and the digital human description information, call a multimodal model, select a target digital human according to the digital human description information, and generate a green screen video of the target digital human based on the target audio and the target digital human to obtain the green screen video of the target digital human; according to the background image description information, call an image generation model from text to generate a background image to obtain a target background image.
[0150] In an optional embodiment, when the processor 35 invokes a multimodal model according to the target audio and the digital human description information, selects a target digital human according to the digital human description information, and generates a green screen video based on the target audio and the target digital human to obtain a digital human green screen video, it is specifically configured to: obtain corresponding target digital human video frames based on the digital human description information; invoke a multimodal model according to the video frames of the target digital human and the target audio, and perform multi-dimensional feature extraction on the target audio to obtain multi-dimensional speech features, where the multi-dimensional speech features include: speech content features, speech emotion features; determine lip movement control parameters of the target digital human according to the speech content features; determine facial expression control parameters of the target digital human according to the speech emotion features; determine limb movement control parameters of the target digital human according to the speech content features and speech emotion features; generate a green screen video of the target digital human based on the lip movement control parameters, expression control parameters, and limb movement control parameters of the target digital human; the lip movement, facial expression, and limb movement of the green screen video of the target digital human match the target audio.
[0151] Further, as Figure 3 shown, the electronic device further includes: other components such as a communication component 36, a display 37, a power supply component 38, an audio component 39, etc. Figure 3 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 3 the components shown. Additionally, Figure 3 the components within the dashed box in Figure 3 are optional components, not mandatory components, and specifically depend on the product form of the electronic device. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smartphone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smartphone, it may include Figure 3 the components within the dashed box in
[0152] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0153] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a communication standard-based wireless network, such as 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.
[0154] The above-mentioned display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0155] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device where the power supply component is located.
[0156] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0157] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above method embodiment. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission medium
[0158] Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instruction, and when the computer program or instruction is executed by a processor, enables the processor to implement the steps in the above method embodiment. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by the computer program or instruction. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiment.
[0159] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0160] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for batch generating videos, characterized in that: include: In response to an input operation on the video generation page for the number of videos and the video category, a batch video generation task is generated, wherein the batch video generation task includes the number of videos N and the video category, where N is an integer ≥ 2; Generate N video instance identifiers according to the batch video generation task, and generate a group of video materials related to the video category for each video instance identifier; Acquire multiple material templates obtained by performing variable processing on the initial video template, wherein the initial video template includes rendering rules of multiple video materials required for generating a video and hierarchical relationships between the multiple video materials, and each material template is used to describe the rendering rules of a video material; For each video instance identifier, determining at least one target materialization template from the multiple materialization templates according to a material type in a group of video materials corresponding to the video instance identifier; Based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target material template is combined to obtain a target video template corresponding to the video instance identifier, wherein the target video template is used to describe the rendering rules and hierarchical relationship of the group of video materials; Video generation processing is performed according to target video templates and a group of video materials corresponding to the N video instance identifiers, so as to obtain N videos under the video category.
2. The method according to claim 1, characterized in that The steps of performing variable processing on the initial video template include: Taking the material type as a splitting variable, parsing the initial video template to obtain information fragments corresponding to multiple material types; Extracting rendering parameter information corresponding to each of the multiple material types from the information fragments corresponding to each of the multiple material types; The rendering parameter information corresponding to each of the multiple material types is templated to obtain the multiple material templates.
3. The method according to claim 1, characterized in that Determining at least one target materialization template from the multiple materialization templates according to the material type in a group of video materials corresponding to the video instance identifier includes: Identify at least one material type included in a group of video materials corresponding to the video instance identifier; According to at least one material type included in the group of video materials, selecting at least one initial material conversion template from the multiple material conversion templates, each material type corresponds to an initial material conversion template; Parameters of at least a portion of the at least one initial materialization template are adjusted to obtain the at least one target materialization template.
4. The method according to claim 3, characterized in that: The initial materialization template includes at least one variable parameter, and each variable parameter is associated with a plurality of candidate parameter values and a default parameter value; and adjusting parameters of at least part of the at least one initial materialization template to obtain the at least one target materialization template includes: Determining the number of templates whose parameters are to be adjusted, wherein the number of templates is less than or equal to the number of the at least one initial materialized template; According to the number of templates, selecting an initial materialized template to be adjusted from the at least one initial materialized template; Determining the variable parameters to be adjusted from the initial materialized template to be adjusted; Randomly determine a target parameter value from a plurality of candidate parameter values associated with the variable parameter to be adjusted; The target parameter value is assigned to the variable parameter to be adjusted to obtain a target material template.
5. The method according to claim 1, characterized in that Based on the hierarchical relationship between the multiple video materials included in the initial video template, the at least one target material template is combined to obtain a target video template corresponding to the video instance identifier, including: Generate a basic video template according to the hierarchical relationship between the multiple video materials included in the initial video template, wherein the basic video template includes multiple blank structural positions corresponding to the multiple video materials; the positional relationship between the multiple blank structural positions reflects the hierarchical relationship between the multiple video materials; insert the at least one target material template into the corresponding blank structural positions in the basic video template respectively, so as to obtain a target video template corresponding to the video instance identifier; or According to the at least one target material template, the structural bits where the rendering rules of the video materials of the same material type in the initial video template are located are overwritten to obtain the target video template corresponding to the video instance identifier; wherein, in the initial video template, the rendering rules of each video material occupies a structural bit, and the positional relationship between the structural bits reflects the hierarchical relationship between the multiple video materials.
6. The method according to claim 1, characterized in that Video generation processing is performed according to the target video templates and a group of video materials corresponding to the N video instance identifiers, so as to obtain N videos under the video category, including: For each video instance identifier, a group of video materials corresponding to the video instance identifier is respectively filled into a target material template in a target video template corresponding to the video instance identifier; For the filled target video template, the group of video materials are rendered according to the rendering rules and hierarchical relationships of the group of video materials described by the filled target video template to obtain a video under the video category.
7. The method according to claim 1, characterized in that A set of video materials related to the video category is generated for each video instance identifier, including: Obtaining, for each video instance identifier, a set of video material description information related to the video category, each set of video material description information including description information of multiple video materials; For each video instance identifier, based on a set of video material description information corresponding to the video instance identifier, multiple material generation models based on artificial intelligence are called to generate multiple video materials corresponding to the video instance identifier, and the multiple video materials corresponding to the video instance identifier are synchronously uploaded to a content distribution network; Accordingly, before performing video generation processing according to the target video templates and a group of video materials corresponding to the N video instance identifiers to obtain N videos under the video category, the method further includes: In response to a batch video generation triggering event, multiple groups of video materials corresponding to the N video instance identifiers are respectively obtained from the content distribution network according to the N video instance identifiers.
8. The method according to claim 6, characterized in that The description information of each set of video materials includes: subtitle description information, audio type description information, digital human description information and background image description information; then, according to the set of video material description information corresponding to the video instance identifier, multiple material generation models based on artificial intelligence are called to generate multiple video materials corresponding to the video instance identifier, including: According to the subtitle description information, calling a generative language model to generate text information to obtain a target subtitle; According to the target subtitles and the audio type description information, calling a text-to-speech model to convert the target subtitles into target audio that matches the audio type description information; According to the target audio and the digital human description information, a multimodal model is called, a target digital human is selected according to the digital human description information, and a green screen video is generated based on the target audio and the target digital human to obtain a green screen video of the target digital human; According to the background image description information, the text image model is called to generate the background image to obtain the target background image.
9. The method according to claim 8, characterized in that According to the target audio and the digital human description information, a multimodal model is called, a target digital human is selected according to the digital human description information, and a green screen video is generated based on the target audio and the target digital human to obtain a digital human green screen video, including: Based on the digital human description information, obtaining a corresponding target digital human video frame; According to the video frame of the target digital human and the target audio, a multimodal model is called to extract multidimensional features of the target audio to obtain multidimensional speech features, wherein the multidimensional speech features include speech content features and speech emotion features; according to the speech content features, a lip shape control parameter of the target digital human is determined; according to the speech emotion features, a facial expression control parameter of the target digital human is determined; according to the speech content features and the speech emotion features, a body movement control parameter of the target digital human is determined; Based on the lip shape control parameters, facial expression control parameters and body movement control parameters of the target digital human, a green screen video of the target digital human is generated; the lip shape, facial expression and body movement of the green screen video of the target digital human are matched with the target audio.
10. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is coupled to the memory and is used to execute the computer program to implement the steps in the method according to any one of claims 1 to 9.
11. A computer-readable storage medium storing a computer program / instruction, characterized in that: When the computer program / instructions are executed by a processor, the processor is enabled to implement the steps in the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that include: A computer program / instruction, when executed by a processor, causes the processor to implement the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and computer readable storage medium
CN113556484A
Media information material processing method and device, electronic equipment and storage medium
CN116801008A
Video generation method and device, equipment and storage medium
CN117354603A
Multi-style video template creation method, system and device and storage medium
CN117412119A
Generating videos
WO2021262137A1
Cited By
Advertisement picture batch generation method and device based on layer editing
CN120912701A