Video generation method, system and device

The video subtitle file and initial video are generated through video scripts, and the animation material templates and target attribute information in the animation material library are used to automatically generate target videos, solving the problem of low efficiency of manual increase in special effects after video synthesis and achieving improvement in video production efficiency.

CN120321432BActive Publication Date: 2025-09-02ALIBABA HEALTH TECH (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510740407.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the existing video production process, the manual increase in video effects after video synthesis is low, resulting in low video production efficiency.

Method used

Generate video subtitle files and initial videos through video scripts, use animation material templates and target attribute information in the animation material library to automatically generate target videos, and integrate subtitle text and animation materials.

Benefits of technology

It greatly improves the efficiency of video animation efficiency generation, reduces the process of manual animation efficiency editing, and improves the overall efficiency of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321432B_ABST
    Figure CN120321432B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a video generation method, system and device, wherein the video generation method includes: generating a video subtitle file and an initial video based on an acquired video script; determining at least one animation material template from an animation material library based on the video subtitle file, determining target attribute information of each animation material template, and generating at least one target animation material through each target attribute information; generating a target video corresponding to the video script based on each target animation material, the video subtitle file and the initial video; by advancing the manual motion effect editing process after video production to the automated node of video generation and production, the motion effect generation efficiency of the entire video production is greatly improved, and the efficiency of video motion effect production is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of video generation technology, and in particular to a video generation method. One or more embodiments of this specification also relate to a video generation system, a video generation apparatus, a computing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] During the video editing process, combining the current video content with appropriate motion effects can greatly enhance the video. With the development of digital human projects, the demand for motion effects in high-quality video production is increasing.

[0003] In the original video special effects production process, when synthesizing videos through cloud editing, the products of video synthesis are relied upon for secondary processing. That is, the conventional video production process is to generate digital human videos and video subtitle files based on the video script, synthesize the digital human videos and video subtitle files through cloud editing, and then manually add video special effects to the synthesized video to complete the video production. Therefore, when the synthesized video requires manual addition of video special effects, the efficiency of video production is low. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a video generation method. One or more embodiments of this specification also relate to a video generation system, a video generation apparatus, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a video generation method is provided, including:

[0006] Generate a video subtitle file and an initial video according to the acquired video script;

[0007] Determining at least one animation material template from an animation material library according to the video subtitle file, determining target attribute information of each animation material template, and generating at least one target animation material according to each target attribute information;

[0008] A target video corresponding to the video script is generated according to each target animation material, the video subtitle file and the initial video.

[0009] According to a second aspect of the embodiments of this specification, a video generation system is provided, including a client and a cloud, wherein:

[0010] The client is configured to send a video generation request to the cloud in response to an interactive operation on the user interaction interface, wherein the video generation request carries a video script;

[0011] The cloud is configured to generate a video subtitle file and an initial video based on the acquired video script; determine at least one animation material template from an animation material library based on the video subtitle file, determine target attribute information of each animation material template, and generate at least one target animation material based on each target attribute information; and generate a target video corresponding to the video script based on each target animation material, the video subtitle file, and the initial video;

[0012] The client is further configured to receive the target video returned by the cloud and display the target video on the user interaction interface.

[0013] According to a third aspect of the embodiments of this specification, a video generating apparatus is provided, including:

[0014] A file generation module is configured to generate a video subtitle file and an initial video according to the acquired video script;

[0015] a material generation module configured to determine at least one animation material template from an animation material library according to the video subtitle file, determine target attribute information of each animation material template, and generate at least one target animation material according to each target attribute information;

[0016] The video generation module is configured to generate a target video corresponding to the video script according to each target animation material, the video subtitle file and the initial video.

[0017] According to a fourth aspect of the embodiments of this specification, a computing device is provided, including:

[0018] memory and processor;

[0019] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned video generation method are implemented.

[0020] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the above-mentioned video generation method are implemented.

[0021] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned video generation method when executed by a processor.

[0022] A video generation method provided by an embodiment of the present specification generates a video subtitle file and an initial video according to an acquired video script. By analyzing the video subtitle file, at least one animation material template is determined from an animation material library. By determining the target attribute information of each animation material template, the synthesis of the target animation material is achieved. The at least one target animation material generated can be integrated with the video subtitle file and the initial video to generate a target video corresponding to the video script, including subtitle text and target animation material. By advancing the manual motion effect editing process after video production to the automation node of video generation, the motion effect generation efficiency of the entire video production is greatly improved, and the efficiency of video motion effect production is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a schematic diagram of a processing process of an existing video generation method provided by an embodiment of this specification;

[0024] Figure 2 This is a scene diagram of a video generation method provided by an embodiment of this specification;

[0025] Figure 3 is a flow chart of a video generation method provided by one embodiment of this specification;

[0026] Figure 4 This is a schematic diagram of a processing process of a video generation method provided by an embodiment of this specification;

[0027] Figure 5 This is a schematic diagram of the structure of a video generation system provided by one embodiment of this specification;

[0028] Figure 6 This is a structural diagram of a video generation device provided by an embodiment of this specification;

[0029] Figure 7 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0030] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0031] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0032] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0033] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0034] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a foundation model. It is pre-trained on a large amount of unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.

[0035] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image description (IC, Image Caption), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0036] First, the terms involved in one or more embodiments of this specification are explained.

[0037] Cloud Editing: This video editing and production service is based on cloud computing and artificial intelligence technologies. It provides users with core functions such as live streaming editing, video editing, template factory, and digital human creation, and can also use AI to assist in editing and production. It can be widely used in industries such as the internet, cultural media, advertising and marketing, education and finance, meeting the needs of enterprises for large-scale, efficient, convenient, and intelligent video content production.

[0038] Motion effect: refers to the dynamic effects in digital media that convey information or enhance user experience through visual changes.

[0039] Lottie: Lottie is an animation file format that allows animations to be displayed consistently across different platforms. It is commonly used for animations in mobile applications.

[0040] Node.js: is an open-source, cross-platform JavaScript runtime environment that allows developers to run JavaScript on the server side.

[0041] The process of video motion effect production is usually to generate a video subtitle file and an initial video through a video script, synthesize an intermediate video through cloud editing, and manually add video motion effects based on the intermediate video to generate the target video.

[0042] For details, please refer to Figure 1 , Figure 1A schematic diagram of the processing process of an existing video generation method provided by an embodiment of the present specification is shown; taking the generation of digital human-related video scenes based on video scripts as an example, when an input video script is received, a digital human video and a video subtitle file are generated based on the video script, the digital human video and the video subtitle file are synthesized through cloud editing, and then video special effects are manually added to the basis of the synthesized video to complete the video production.

[0043] That is, video animation effects are secondary processing products that rely on video synthesis, with high labor costs and low production efficiency.

[0044] To solve the above technical problems, this specification provides a video generation method. This specification also involves a video generation system, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0045] See also Figure 2 , Figure 2 A schematic diagram of a scenario of a video generation method provided according to an embodiment of this specification is shown.

[0046] Specifically, the video processing method is applied to a video generation system, which includes a client 202 and a cloud 204. In response to an interactive operation of a user interface, the client 202 sends a video generation request to the cloud 204, and the video generation request carries a video script.

[0047] When the cloud 204 receives the video generation request sent by the client 202, a video subtitle file and an initial video are generated according to the obtained video script; according to the video subtitle file, at least one animation material template is determined from the animation material library, the target attribute information of each animation material template is determined, and at least one target animation material is generated according to each target attribute information; according to each target animation material, the video subtitle file and the initial video, a target video corresponding to the video script is generated, and the target video is returned to the client 202 so as to be displayed on the client 202.

[0048] The client 202 may include a browser, an application (APP), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program, a type of lightweight application), or a cloud application. The client can be developed based on a software development kit (SDK) for the corresponding service provided by the server, such as a real-time communication (RTC) SDK. The client can be deployed in an electronic device and rely on the device or certain applications within the device to operate. The electronic device may have a display and support information browsing, such as a personal mobile terminal such as a mobile phone, tablet computer, or personal computer. Various other types of applications are also typically configured in the electronic device, such as human-computer interaction applications, model training applications, video generation applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0049] Cloud 204 can be understood as servers that provide various services, including cloud servers. For example, servers providing communication services to multiple clients, servers supporting backend training for models used by clients, and servers processing data sent by clients can be implemented. It should be noted that Cloud 204 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. Cloud 204 can also be a server in a distributed system or a server integrated with blockchain. Cloud 204 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or it can be an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology.

[0050] The video generation method provided in the embodiments of this specification, when generating a video subtitle file and an initial video according to an acquired video script, determines at least one animation material template from an animation material library through analysis of the video subtitle file, and realizes synthesis of target animation materials by determining target attribute information of each animation material template. The generated at least one target animation material can be integrated with the video subtitle file and the initial video to generate a target video corresponding to the video script, including subtitle text and target animation materials; by advancing the manual motion effect editing process after video production to the automation node of video generation, the motion effect generation efficiency of the entire video production is greatly improved, and the efficiency of video motion effect production is greatly improved.

[0051] See also Figure 3 , Figure 3 A flowchart of a video generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0052] Step 302: Generate a video subtitle file and an initial video according to the acquired video script.

[0053] A video script can be understood as text data used to guide the generation and editing of video content. A video subtitle file can be understood as a text data file containing timeline information, used to display text content corresponding to the audio or video images during video playback. An initial video can be understood as a primary video version automatically generated based on the video script. This initial video contains the initially generated video images but does not include subtitle text or animation effects.

[0054] Specifically, when a video script is obtained, a video subtitle file is generated based on the text content in the video script, and the initial video is generated based on the video script. It should be noted that in the cloud editing scenario, the subtitle text in the video subtitle file can be synchronized with the video playback timeline of the initial video.

[0055] Step 304: According to the video subtitle file, at least one animation material template is determined from the animation material library, target attribute information of each animation material template is determined, and at least one target animation material is generated according to each target attribute information.

[0056] An animation material template can be understood as a predefined dynamic graphics design framework used to enhance the visual expressiveness of a video. This animation material template contains a static structure (which includes immutable components such as basic graphics and animation paths) and configurable attribute fields, allowing for batch generation of animation materials with a unified style. Specifically, the animation material template contains configurable attribute fields (such as text content, color, and scaling), and personalized configuration is achieved by assigning different attribute information to the attribute fields. This means that different target animation materials can be generated using the same animation material template with different attribute information.

[0057] The target attribute information can be understood as the actual numerical value or logical rule assigned to the animation material template in a specific application scenario; the target animation material can be understood as the final usable dynamic graphic element generated by injecting specific attribute information into the animation material template. The target animation material is independent of the main video content and is presented as an overlay; the target animation material can be edited independently or exported as an independent media file. For example, the file format of the target animation material is Lottie or GIF, which is not limited here.

[0058] Specifically, based on the video subtitle file, at least one animation material template that matches the subtitle text in the video subtitle file can be determined from the animation material library. The animation material template is a dynamic graphic design prototype that is not bound to specific parameters. The animation material template generates target animation material by receiving external input attribute configuration data (that is, target attribute information, such as specific text content), that is, determining the target attribute information of each animation material template. For each target attribute information, a corresponding target animation material can be determined. The target animation material is a renderable instance after the parameters are solidified.

[0059] The animation material template includes but is not limited to subtitle borders, dynamic icons, etc. Taking the animation material template as a lace subtitle border as an example, the configurable field corresponding to the lace subtitle border includes a text placeholder {text}. The subtitle text in the video subtitle file (for example, the subtitle text is "I am a pediatrician") is assigned to the text placeholder as the specific text content of the subtitle border, and the coordinate position of the subtitle border is configured to (50,80). At this time, the subtitle text and the configured coordinate position are the target attribute information of the subtitle border. According to these configured target attribute information, a subtitle dynamic effect material (i.e., the target animation material) with a lace subtitle border, a coordinate position of (50,80), and a text content of "I am a pediatrician" can be generated.

[0060] In fact, through code logic (such as Node.js scripts), dynamic data can be injected into the configurable parameters in the animation material template to synthesize the target animation material.

[0061] In one or more embodiments of this specification, when determining an animation material template from an animation material library based on a video subtitle file, it is necessary to split the subtitle text in the video subtitle file to obtain individual lines of subtitle text, and then determine the corresponding animation material template based on the line of subtitle text. The specific implementation is as follows:

[0062] The step of determining at least one animation material template from an animation material library according to the video subtitle file includes:

[0063] Splitting the subtitle text in the video subtitle file to obtain at least one line of subtitle text after splitting;

[0064] At least one animation material template is determined from the animation material library according to each line of subtitle text.

[0065] The line subtitle text may be understood as text data obtained by dividing the subtitle text into lines and including line numbering.

[0066] Specifically, when splitting the subtitle text in a video subtitle file, the lines can be divided according to preset rules, such as automatically wrapping at a fixed number of characters (20 characters per line); the lines can also be divided according to the semantics of the text, such as according to punctuation marks (such as periods, commas) or natural language processing (NLP) word segmentation results; the lines can also be divided manually, such as the user forcing a line break through the editing interface. The specific line breaking method can be selected according to the actual situation and is not limited here.

[0067] For example, if the subtitle text is "Pediatricians need to have professional medical knowledge, patient communication skills and rich clinical experience", the subtitle text can be divided into lines through semantic segmentation, and the subtitle text lines that can be obtained are "1. Pediatricians need to have professional medical knowledge", "2. Patient communication skills", and "3. And rich clinical experience".

[0068] In actual applications, when obtaining a video subtitle file through a video script, the video subtitle file containing line subtitle texts can be obtained by splitting the video script, which is not limited here.

[0069] When at least one line of subtitle text is obtained after being split, at least one animation material template is determined from the animation material library according to each line of subtitle text, thereby ensuring accurate matching using each line of subtitle text.

[0070] The video generation method provided in the embodiments of this invention can accurately and dynamically match the animation material template according to the detailed information of each line of subtitle text, while avoiding the time-consuming operation of manually selecting animation materials line by line, by splitting at least one line of subtitle text and determining at least one animation material template from the animation material library based on each line of subtitle text.

[0071] In one or more embodiments of this specification, when obtaining line subtitle text, the animation material template can be dynamically matched based on the semantic information of the line subtitle text. When obtaining the corresponding semantic information of the line subtitle text through the semantic root model, the semantic information determined varies depending on the target analysis type of the prompt word. Therefore, appropriate prompt words can be constructed according to actual needs and input into the semantic analysis model together with the line subtitle text. The specific implementation method is as follows:

[0072] The step of determining at least one animation material template from the animation material library according to each line of subtitle text includes:

[0073] Inputting each line of subtitle text and the prompt word into a semantic analysis model to obtain target semantic information of each line of subtitle text, wherein the prompt word includes a target analysis type, and the target semantic information is semantic information of the target analysis type;

[0074] At least one animation material template is determined from the animation material library according to the target semantic information of each line of subtitle text.

[0075] The semantic analysis model can be understood as a large model designed to parse and understand the deeper meaning of text, including semantic information such as entities, relationships, sentiment, and intent. It can also transform natural language into structured representations. The target analysis type specifies the semantic analysis task the model is intended to perform. For example, semantic analysis tasks can include sentiment analysis, entity recognition, and relationship extraction.

[0076] Specifically, in the embodiment of the specification, the target analysis type is sentiment analysis. For example, the subtitle text lines obtained are "1. Everyone should pay attention", "2. Some patients have high low blood pressure", and "3. It may be related to the medication they are taking". The prompt words are "Please mark the key lines in the subtitles based on the subtitles I provided. And identify his emotional tendency." When each line of subtitle text and the prompt words are input into the semantic analysis model together, the target semantic information obtained for each line of subtitle text is:

[0077] {“emotion”:“Warning”, “lineNumber”:1, “subtitleContent”:“Everyone please pay attention”}

[0078] {"emotion":"neutral","lineNumber":2,"subtitleContent":"Some patients have high low blood pressure"}

[0079] {"emotion":"neutral","lineNumber":3,"subtitleContent":"May be related to medications taken"}

[0080] It should be noted that the prompt word may also include sample data, so that the semantic analysis model can refer to the sample data to better perform semantic analysis and obtain more accurate analysis results.

[0081] When each line of subtitle text is obtained, at least one animation material template may be matched from the animation material library according to target semantic information corresponding to each line of subtitle text.

[0082] The video generation method provided in the embodiment of this description obtains the target semantic information of each line of subtitle text by performing semantic analysis on each line of subtitle text. The target semantic information is the semantic information corresponding to the target analysis type in the prompt word, that is, different semantic information of the line subtitle text can be obtained by the target analysis type in the different prompt words, thereby realizing the extraction of semantic information of different dimensions as needed, thereby matching the animation material template related to the dimension.

[0083] In one or more embodiments of this specification, based on the target semantic information of each line of subtitle text, the animation material template corresponding to each line of subtitle text can be determined from the animation material library. To avoid duplication of animation effects, the determined animation material templates corresponding to each line of subtitle text can be deduplicated. Specific implementation methods are as follows:

[0084] The step of determining at least one animation material template from the animation material library according to the target semantic information of each line of subtitle text comprises:

[0085] Determining, from the animation material library, an animation material template corresponding to each line of subtitle text according to target semantic information of each line of subtitle text, wherein a semantic label of the animation material template corresponds to the target semantic information;

[0086] The at least one animation material template is obtained by performing deduplication processing on the animation material templates corresponding to the lines of subtitle text.

[0087] Specifically, when the target semantic information corresponding to each line of subtitle text is obtained, the corresponding animation material template can be determined from the animation material library for each identified target semantic information. In the case that the target semantic information corresponding to some lines of subtitle text is similar, the animation material templates matched by different lines of subtitle text are the same or similar. In order to avoid the repeated use of the same or similar animation material templates causing visual fatigue to the user and reducing the attractiveness of the content, deduplication processing can be performed on the same or similar animation material templates, thereby ensuring the visual and functional differentiation of the generated target animation material, and the remaining animation material templates after deduplication processing are determined as at least one animation material template determined from the animation material library based on the video subtitle text.

[0088] Of course, if due to functional consistency requirements, it is necessary to add a unified subtitle border to the subtitle text, then deduplication cannot be performed on the unified subtitle border. Therefore, in actual applications, the animation material templates that do not require deduplication can be marked according to actual conditions to meet actual needs.

[0089] The video generation method provided in the embodiments of this specification can determine the corresponding animation material template for each line of subtitle text. In order to avoid duplication of animation effects and improve the freshness of content, the determined animation material template corresponding to each line of subtitle text can be deduplicated.

[0090] In one or more embodiments of this specification, when determining an animation material template that matches each target semantic information from an animation material library based on the target semantic information of a line of subtitle text, the target semantic information is specifically matched with the semantic tag of the animation material template. Specific implementation methods are as follows:

[0091] The step of determining, from the animation material library, the animation material template corresponding to each line of subtitle text according to the target semantic information of each line of subtitle text comprises:

[0092] Determining, from the animation material library, a semantic tag that matches the target semantic information of each line of subtitle text according to the target semantic information of each line of subtitle text;

[0093] An animation material template corresponding to each line of subtitle text is determined from at least one initial material template corresponding to each semantic tag.

[0094] Specifically, a batch of semantically labeled initial material templates are pre-stored in the animation material library, that is, the initial material templates in the animation material library each correspond to a semantic label. Taking a line of subtitle text as an example, when the target semantic information of the line of subtitle text is obtained, the semantic label that matches the target semantic information can be determined, and then an initial material template can be selected from the corresponding multiple initial material templates through the semantic label as the animation material template corresponding to the line of subtitle text.

[0095] In specific implementation, since multiple initial material templates can correspond to the same semantic tag, when selecting an animation material template from the corresponding initial material template based on the semantic tag, the animation material template can be determined by random selection, or the initial material template with the highest selection frequency can be used as the animation material template. There is no limitation here.

[0096] For example, when performing sentiment analysis on a line subtitle text and obtaining that the target semantic information corresponding to the line subtitle text is "angry", it is determined that the semantic label in the animation material library that matches the target semantic information is "angry", and therefore the first initial material template is randomly selected from the first initial material template, the second initial material template, the third initial material template and the fourth initial material template corresponding to the semantic label as the animation material template corresponding to the line subtitle text.

[0097] In fact, the same keyword requires different animation effects in different contexts. For example, the animation effects selected for "explosion" in news videos and variety shows are different. Therefore, the semantic label of the animation material template is a multi-dimensional semantic label. For example, the semantic label of the animation material template is composed of multi-dimensional labels such as domain labels (also called scene labels), emotion labels, and entity labels. In a case where different types of semantic information can be determined based on the different target analysis types in the prompt words for the line subtitle text, multiple types of semantic information corresponding to the line subtitle text can be obtained. Through the multiple types of semantic information, the matching multi-dimensional semantic labels are determined, so that the animation material template can be more accurately determined from at least one initial material template based on the multi-dimensional semantic labels.

[0098] For example, when determining that the emotional semantic information corresponding to a line of subtitle text is "angry" and the domain semantic information is "medicine", in the animation material library, the multi-dimensional label with the domain label of medicine and the emotional label of anger is matched. Specifically, the full label combination can be matched first. If there is no such combination, it is backtracked step by step to determine the animation material template from at least one initial material template corresponding to the multi-dimensional semantic label.

[0099] Of course, when multi-dimensional labels are involved, multiple initial material templates may belong to the same field and express the same emotions, but the weight values ​​of different dimensional labels among them may be different, that is, some initial material templates emphasize the field more (the weight value of the field dimension label is high), and some emphasize emotions more (the weight value of the emotion dimension label is high). Therefore, when selecting an animation material template from at least one initial material template corresponding to the multi-dimensional label, the priority of different dimensional labels can be set, so as to select the initial material template with the most complete label coverage and the highest total weight as the animation material template.

[0100] For example, when the semantic label determined according to the target semantic information of the line subtitle text is [medical, serious], the determined initial material template includes initial material template A and initial material template B, among which the semantic label of initial material template A is [medical: 0.8][serious: 0.8], and the semantic label of initial material template B is [medical: 0.7][serious: 0.9]. When the priority of the emotion dimension label is set higher than the domain dimension label, the initial material template B is selected as the animation material template.

[0101] The video generation method provided in the embodiments of this specification pre-stores semantically labeled initial material templates in an animation material library, and can match corresponding semantic tags from the animation material library according to the target semantic information corresponding to the line subtitle text, thereby determining the animation material template corresponding to the line subtitle text from at least one initial material template corresponding to the semantic tag. When the animation material template is selected using the semantic tag, the standardized source of animation material input is solved through the construction of a semantic animation material library.

[0102] In one or more embodiments of this specification, once an animation material template is obtained, a specific instantiated target animation material can be generated by determining the target attribute information of the animation material template. The target attribute information of the animation material template can be accurately determined based on the target and relative position rules of each animation material template. Specific implementation methods are as follows:

[0103] Determining target attribute information of each animation material template includes:

[0104] Determining the attribute fields to be configured corresponding to the animation material templates according to the action objects of the animation material templates, wherein the action objects include the subtitle text in the video subtitle file and the video object in the initial video;

[0105] By configuring the to-be-configured attribute fields corresponding to the animation material templates, target attribute information of the animation material templates is obtained.

[0106] Among them, the action object can be understood as the video content element that the animation material template needs to be bound to. For example, in the above embodiment, when the animation material template is a subtitle border, the action object of the animation material template is the subtitle text, and when the animation material template is a "small flame" expressing angry emotions, the action object of the animation material template is the video object in the initial video (such as the video character in the video).

[0107] The attribute field to be configured can be understood as the configurable attribute field in the above embodiment. When the action objects corresponding to the animation material template are different, the attribute field to be configured is different.

[0108] In practice, the attribute fields to be configured may include a relative position field. Specifically, the relative position information corresponding to the relative position fields in each animation material template can be determined based on the relative position rules of each animation material template. Relative position rules are used to define the spatial relationship of the animation material template relative to the object it is acting on. For example, relative position rules can include "center display," "10 pixels to the right," "follow the character's movement," and so on.

[0109] Specifically, determine the subtitle text or video object that the animation material template is to act on, and calculate the relative position information of the animation material template according to the preset relative position rule; for example, the animation material template is a subtitle border, and the relative position rule of the subtitle border is that it must be 5 pixels outside the text. Therefore, the relative position information of the animation material template can be determined according to the relative position rule, and the relative position information is filled in the relative position field of the animation material template to obtain the target attribute information.

[0110] The video generation method provided in the embodiment of this specification determines the target attribute information of the animation material template through precise positioning, thereby ensuring that the animation material template is accurately matched with the object of action; by configuring the configuration attribute fields, the efficiency and quality of video production are greatly improved, ensuring that the animation effect is always accurate.

[0111] In one or more embodiments of the present specification, an animation material template whose object is subtitle text in at least one animation material template is determined as a first animation material template, and an animation material template whose object is a video object in at least one animation material template is determined as a second animation material template; the target position coordinates of the first animation material template are determined by determining the text position coordinates of the subtitle text corresponding to the first animation material template and the relative position rules corresponding to the first type of animation material template; the target position coordinates of the second animation material template are determined by determining the object position coordinates of the video object corresponding to the second animation material template and the relative position rules corresponding to the second type of animation material template. Specific implementation methods are as follows:

[0112] The step of obtaining target attribute information of each animation material template by configuring the attribute fields to be configured corresponding to each animation material template includes:

[0113] For a first animation material template among the at least one animation material template, determining that the attribute field to be configured corresponding to each first animation material template includes a text content field, wherein the first animation material template acts on a subtitle text of the video subtitle file;

[0114] configuring the text content field using the subtitle text corresponding to each of the first animation material templates, and determining target attribute information of each of the first animation material templates;

[0115] For a second animation material template in at least one animation material template, determining that the attribute fields to be configured corresponding to each second animation material template include an object binding field, wherein the action object of the second animation material template is a video object in the initial video;

[0116] configuring the object binding field using the video object corresponding to each second animation material template to determine target attribute information of each second animation material template;

[0117] The target attribute information of each animation material template is determined according to the target attribute information of each first animation material template and the target attribute information of each second animation material template.

[0118] Specifically, the animation material template whose object is the subtitle text in at least one animation material template is determined as the first animation material template, and the animation material template whose object is the video object in at least one animation material template is determined as the second animation material template.

[0119] In fact, when the object of the animation material template is different, the attribute fields to be configured corresponding to the animation material template are different. When the object of the animation material template is subtitle text, the attribute fields to be configured corresponding to the animation material template include text content fields. When the object of the animation material template is a video object, the attribute fields to be configured corresponding to the animation material template include object binding fields.

[0120] For the first animation material template, the attribute fields to be configured corresponding to each first animation material template include a text content field. By determining the subtitle text corresponding to the first animation material template as the specific parameter value of the text content field, the configuration of the text content field is realized, thereby obtaining the target attribute information of the first animation material template.

[0121] For the second animation material template, the attribute fields to be configured corresponding to each second animation material template include an object binding field. By determining the video object (object identifier of the video object) corresponding to the second animation material template as the specific parameter value of the object binding field, the object binding field is configured to obtain the target attribute information of the first animation material template.

[0122] For example, the subtitle text is recognized as "This product uses innovative technology", and the video screen corresponding to the subtitle text is a displayed mobile phone product. The animation material template matched with the subtitle text includes a technologically flashing subtitle border (acting on the subtitle text) and a highlight shining line (acting on the video object); when the relative position rule of the subtitle border is 10 pixels below the text, the parameter value of the relative position field corresponding to the subtitle border is determined to be "10 pixels below the text", and the parameter value of the text content field corresponding to the subtitle border is determined to be "This product uses innovative technology", so as to obtain the instantiated (parameterized) subtitle motion material (i.e., the target animation material) by configuring the parameters of the attribute field to be configured of the subtitle border; the parameter value of the object binding field corresponding to the line is determined to be "mobile phone product phone", so as to realize the binding between the line and the mobile phone product.

[0123] The video generation method provided in the embodiments of this specification determines the attribute fields to be configured corresponding to each animation material template for animation material templates applied to different objects, and parameterizes the attribute fields to be configured based on the specific information of the object, thereby accurately obtaining the target attribute information corresponding to each animation material template.

[0124] In one or more embodiments of the present specification, when the subtitle text of a video subtitle file is split to obtain multiple lines of subtitle text, at least two target lines of subtitle text having a target structural relationship are determined from the multiple lines of subtitle text, and a combined animation material template corresponding to the at least two target lines of subtitle text is determined. Specific implementations are as follows:

[0125] Before determining the target attribute information of each animation material template, the following steps are further included:

[0126] Splitting the subtitle text of the video subtitle file to obtain at least one line of subtitle text after splitting;

[0127] In the case where the at least one line of subtitle texts includes a plurality of subtitle texts, determining at least two target line of subtitle texts having a target structural relationship from the plurality of line of subtitle texts;

[0128] Determine the combined animation material template corresponding to the at least two target lines of subtitle text.

[0129] The configuring the text content field using the subtitle text corresponding to each first animation material template to determine target attribute information of each first animation material template includes:

[0130] At least two target lines of subtitle text corresponding to the combined animation material template are used to configure at least two text content fields corresponding to the combined animation material template according to the target structural relationship, and target attribute information of the combined animation material template is determined.

[0131] Among them, the target structural relationship can be understood as a parallel / progressive relationship that is adjacent, structurally similar, and semantically related. For example, a subtitle text is "The operation process is as follows: Step 1, input data; Step 2, run the algorithm; Step 3, analyze the results." When this subtitle text is split, the obtained line subtitle texts are "The operation process is as follows:", "Step 1, input data;", "Step 2, run the algorithm;", and "Step 3, analyze the results."; At this time, "Step 1, input data;", "Step 2, run the algorithm;", and "Step 3, analyze the results." are target line subtitle texts with a target structural relationship.

[0132] In fact, when at least one line of subtitle text is obtained after splitting, an animation material template can correspond to a line of subtitle text. For example, a subtitle text is "Everyone should pay attention, some patients have high low blood pressure, which may be related to the medications they take". After splitting the subtitle text, three line subtitle texts are obtained, namely "Everyone should pay attention", "Some patients have high low blood pressure, "May be related to the medications they take", and two animation material templates are determined to act on the above-mentioned line subtitle texts. Specifically, animation material template A corresponds to the line subtitle text "Everyone should pay attention", and animation material template B corresponds to the line subtitle text "May be related to the medications they take". Therefore, the text content field of animation material template A is configured through the line subtitle text "Everyone should pay attention" to generate the target animation material A, and the text content field of animation material template B is configured through the line subtitle text "May be related to the medications they take" to generate the target animation material B.

[0133] For at least two target lines of subtitle texts with a target structural relationship, a corresponding combined animation material template can be determined. The combined animation material template can strengthen the parallel / progressive relationship between the target lines of subtitle texts through elements such as serial numbers, connecting lines, and cards. Each target line of subtitle text can appear dynamically one by one in the combined animation material template.

[0134] Specifically, the combined animation material template corresponds to at least two text content fields, so the combined animation material template can act on at least two target lines of subtitle texts. By using at least two target lines of subtitle texts as parameter values ​​of at least two text content fields respectively, the parameterized configuration of at least two text content fields in the combined animation material template is realized, and the target attribute information of the combined animation material template is determined.

[0135] Continuing with the above example, when "Step 1, input data;", "Step 2, run the algorithm;", and "Step 3, analyze the results." are determined to be target row subtitle texts with a target structural relationship, the corresponding combined animation material template C is determined. The text content fields of the combined animation material template C include text1, text2, and text3, and text1 is set to appear first and text3 to appear last. Therefore, according to the target structural relationship, "Step 1, input data;" is determined as the parameter value of the text content field text1, "Step 2, run the algorithm;" is determined as the parameter value of the text content field text2, and "Step 3, analyze the results." is determined as the parameter value of the text content field text3 to generate the target combined animation material C, thereby achieving the effect of using the combined animation material template to make the target row subtitle texts with a target structural relationship appear dynamically one by one.

[0136] In fact, when the object of the animation material template is subtitle text and the configurable parameters of the animation material template include text content, the subtitle text is used as the text content of the animation material template, and combined with attribute information such as relative position information, the target animation material can be directly generated; by inserting the subtitle text and the generated target animation material into the initial video, the target video containing subtitles and motion effects is generated.

[0137] The video generation method provided in the embodiments of this specification splits the subtitle text of a video subtitle file to obtain at least one line of subtitle text after the split, and determines at least two target lines of subtitle text having a target structural relationship and a combined animation material template corresponding to the at least two target lines of subtitle text. At least two text content fields of the combined animation material template are configured through the at least two target lines of subtitle text to achieve rich animation display effects.

[0138] Step 306: Generate a target video corresponding to the video script according to each target animation material, the video subtitle file and the initial video.

[0139] Specifically, when a video subtitle file, an initial video, and a target animation material are obtained, the target video corresponding to the video script and including subtitle content and animation effects is obtained by combining the video subtitle file and the target animation material on the basis of the initial video.

[0140] In one or more embodiments of this specification, based on the initial video, the target animation material and the video subtitle file are superimposed on the initial video as independent tracks. Since the target animation material corresponds to a line of subtitle text, the display duration of the target animation material is controlled by the timeline information of the corresponding line of subtitle text. The specific implementation is as follows:

[0141] Generating a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video includes:

[0142] In a case where the video subtitle file includes at least one line of subtitle text and timeline information corresponding to each line of subtitle text, determining the animation timeline information of each target animation material according to the text timeline information corresponding to each line of subtitle text;

[0143] According to the text timeline information corresponding to each line of subtitle text and the animation timeline information of each target animation material, the each line of subtitle text and the each target animation material are synthesized into the initial video to generate the target video corresponding to the video script.

[0144] Specifically, by parsing the video subtitle file, the text timeline information corresponding to each line of subtitle text can be obtained, for example, the start time of the first line of subtitle text is "00:00:01.000" and the end time is "00:00:04.000"; when the target animation material corresponds to the first line of subtitle text, the animation timeline information corresponding to the target animation material can be determined based on the text timeline information corresponding to the first line of subtitle text. For example, when the above-mentioned first line of subtitle text corresponds to the target animation material A, the start time of the target animation material A can be determined to be "00:00:01.000" and the end time is "00:00:04.000".

[0145] Of course, the above situation is that the target animation material appears / disappears at the same time as the associated subtitle text by default. In fact, further settings can be made. For example, by setting the lead amount (the target animation material appears N seconds earlier than the subtitle text) and the lag amount (the target animation material disappears M seconds later than the subtitle text), the animation timeline information of the target animation material can be dynamically determined.

[0146] On the basis of the initial video as the basic video layer, each line of subtitle text and each target animation material is used as a text layer and an animation layer respectively, and is superimposed on the basic video layer according to the text timeline information corresponding to each line of subtitle text and the animation timeline information of each target animation material, that is, synthesized into the initial video; when each line of subtitle text and each target animation material is superimposed on the initial video, their position in the initial video can be adjusted according to actual conditions, so as to generate a target video with better visual perception.

[0147] The video generation method provided in the embodiment of this specification obtains the text timeline information corresponding to each line of subtitle text through analysis, determines the animation timeline information of the target animation material corresponding to the line of subtitle text, and then accurately superimposes each line of subtitle text and each target animation material into the initial video based on the timeline information of each line of subtitle text and each target animation material to complete the production of the target video.

[0148] The video generation method provided in the embodiments of this specification realizes the standardized extraction capability of motion effects in video scripts by semantically labeling animation material templates and matching the semantic tags of animation material templates with the target semantic information recognized by the large model, thereby providing a basis for the batch automated production of video motion effects, and ultimately achieving the timing left shift of the video motion effect production process, moving from the post-processing process of video production to an automated node in the video production pipeline, greatly improving the efficiency of video motion effect production.

[0149] See also Figure 4 , Figure 4 A schematic diagram of the processing process of a video generation method provided by an embodiment of this specification is shown.

[0150] Specifically, a video subtitle file and a digital human video are generated according to the input video script, and the digital human video includes a digital human with a virtual image.

[0151] Construct a material library (i.e., the animation material library in the above embodiment), which stores semantically labeled motion effect material templates (i.e., the animation material templates in the above embodiment), and specific semantic types include but are not limited to summary, question, misunderstanding, shock, etc.

[0152] The video script is intelligently analyzed through the big model (i.e., the semantic analysis model in the above embodiment), and combined with the matching of the material library, the motion effect task corresponding to the video script is generated. Specifically, when the video subtitle file is obtained through the video script, the line subtitle text in the video subtitle file is intelligently semantically analyzed by calling the big model, so as to identify the target semantic information corresponding to the line subtitle text, and determine the motion effect material template corresponding to the line subtitle text by matching the target semantic information of the line subtitle text with the semantic label of the motion effect material template in the material library.

[0153] By calling the Node service, the animation production for the animation material template is realized. Specifically, the animation task is synthesized and output through Node.js, that is, the parameter values ​​of the fields that can be dynamically replaced in the animation material template are filled in with the line subtitle text, thereby generating the target animation material.

[0154] By editing the subtitle file of the video in the cloud, synthesizing the target animation material and the digital human video, the target video is synthesized to produce the target video containing the subtitle content and the motion effect, realizing the automatic insertion of the motion effect in the video cloud editing process.

[0155] The video generation method provided in the embodiments of this specification realizes the automation capability of motion effect production, and advances the manual editing process after video production to the front-end automation node of video cloud editing. The entire process is highly closed-loop, avoiding the inefficiency of manual operation, ensuring the efficient execution of tasks, and greatly improving the motion effect generation efficiency of the entire video production, providing a foundation for large-scale motion effect video production.

[0156] Corresponding to the above method embodiment, this specification also provides a video generation system embodiment, Figure 5 FIG. 1 shows a schematic diagram of the structure of a video generation system provided by an embodiment of this specification. Figure 5 As shown, the video generation system includes: a client 202 and a cloud 204, wherein:

[0157] The client 202 is configured to send a video generation request to the cloud in response to an interactive operation on the user interface, wherein the video generation request carries a video script;

[0158] The cloud 204 is configured to generate a video subtitle file and an initial video based on the acquired video script; determine at least one animation material template from an animation material library based on the video subtitle file, determine target attribute information of each animation material template, and generate at least one target animation material based on each target attribute information; and generate a target video corresponding to the video script based on each target animation material, the video subtitle file, and the initial video;

[0159] The client 202 is further configured to receive the target video returned by the cloud 204 and display the target video on the user interaction interface.

[0160] The interactive operation can be understood as an upload operation of a video script, or an input operation of a video script, which is not limited here. Specific implementation methods can be found in the above embodiments, which will not be described in detail here.

[0161] The above is a schematic scheme of a video generation system of this embodiment. It should be noted that the technical scheme of the video generation system and the technical scheme of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical scheme of the video generation system, please refer to the description of the technical scheme of the above-mentioned video generation method.

[0162] Corresponding to the above method embodiment, this specification also provides a video generation device embodiment, Figure 6 FIG. 1 shows a schematic diagram of the structure of a video generating device provided by an embodiment of this specification. Figure 6 As shown, the video generating device includes:

[0163] The file generation module 602 is configured to generate a video subtitle file and an initial video according to the acquired video script;

[0164] The material generation module 604 is configured to determine at least one animation material template from the animation material library according to the video subtitle file, determine target attribute information of each animation material template, and generate at least one target animation material according to each target attribute information;

[0165] The video generation module 606 is configured to generate a target video corresponding to the video script according to each target animation material, the video subtitle file and the initial video.

[0166] Optionally, the material generation module 604 is further configured to:

[0167] Splitting the subtitle text in the video subtitle file to obtain at least one line of subtitle text after splitting;

[0168] At least one animation material template is determined from the animation material library according to each line of subtitle text.

[0169] Optionally, the material generation module 604 is further configured to:

[0170] Inputting each line of subtitle text and the prompt word into a semantic analysis model to obtain target semantic information of each line of subtitle text, wherein the prompt word includes a target analysis type, and the target semantic information is semantic information of the target analysis type;

[0171] At least one animation material template is determined from the animation material library according to the target semantic information of each line of subtitle text.

[0172] Optionally, the material generation module 604 is further configured to:

[0173] Determining, from the animation material library, an animation material template corresponding to each line of subtitle text according to target semantic information of each line of subtitle text, wherein a semantic label of the animation material template corresponds to the target semantic information;

[0174] The at least one animation material template is obtained by performing deduplication processing on the animation material templates corresponding to the lines of subtitle text.

[0175] Optionally, the material generation module 604 is further configured to:

[0176] Determining, from the animation material library, semantic tags that match the target semantic information of each line of subtitle text;

[0177] An animation material template corresponding to each line of subtitle text is determined from at least one initial material template corresponding to each semantic tag.

[0178] Optionally, the material generation module 604 is further configured to:

[0179] Determining the attribute fields to be configured corresponding to the animation material templates according to the action objects of the animation material templates, wherein the action objects include the subtitle text in the video subtitle file and the video object in the initial video;

[0180] By configuring the to-be-configured attribute fields corresponding to the animation material templates, target attribute information of the animation material templates is obtained.

[0181] Optionally, the material generation module 604 is further configured to:

[0182] For a first animation material template among the at least one animation material template, determining that the attribute field to be configured corresponding to each first animation material template includes a text content field, wherein the first animation material template acts on a subtitle text of the video subtitle file;

[0183] configuring the text content field using the subtitle text corresponding to each of the first animation material templates, and determining target attribute information of each of the first animation material templates;

[0184] For a second animation material template in at least one animation material template, determining that the attribute fields to be configured corresponding to each second animation material template include an object binding field, wherein the action object of the second animation material template is a video object in the initial video;

[0185] configuring the object binding field using the video object corresponding to each second animation material template to determine target attribute information of each second animation material template;

[0186] The target attribute information of each animation material template is determined according to the target attribute information of each first animation material template and the target attribute information of each second animation material template.

[0187] Optionally, the material generation module 604 is further configured to:

[0188] Splitting the subtitle text of the video subtitle file to obtain at least one line of subtitle text after splitting;

[0189] In the case where the at least one line of subtitle texts includes a plurality of subtitle texts, determining at least two target line of subtitle texts having a target structural relationship from the plurality of line of subtitle texts;

[0190] Determine the combined animation material template corresponding to the at least two target lines of subtitle text.

[0191] Optionally, the material generation module 604 is further configured to:

[0192] At least two target lines of subtitle text corresponding to the combined animation material template are used to configure at least two text content fields corresponding to the combined animation material template according to the target structural relationship, and target attribute information of the combined animation material template is determined.

[0193] Optionally, the video generation module 606 is further configured to:

[0194] In a case where the video subtitle file includes at least one line of subtitle text and text timeline information corresponding to each line of subtitle text, determining the animation timeline information of each target animation material according to the text timeline information corresponding to each line of subtitle text;

[0195] According to the text timeline information corresponding to each line of subtitle text and the animation timeline information of each target animation material, the each line of subtitle text and the each target animation material are synthesized into the initial video to generate the target video corresponding to the video script.

[0196] The video generation device provided in the embodiments of this specification, when generating a video subtitle file and an initial video according to an acquired video script, determines at least one animation material template from an animation material library through analysis of the video subtitle file, and realizes synthesis of target animation materials by determining target attribute information of each animation material template. The generated at least one target animation material can be integrated with the video subtitle file and the initial video to generate a target video corresponding to the video script, including subtitle text and target animation materials; by advancing the manual motion effect editing process after video production to the automation node of video generation, the motion effect generation efficiency of the entire video production is greatly improved, and the efficiency of video motion effect production is greatly improved.

[0197] The above is a schematic diagram of a video generation device according to this embodiment. It should be noted that the technical solution of the video generation device and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the video generation device, please refer to the description of the technical solution of the above-mentioned video generation method.

[0198] Figure 77 shows a block diagram of a computing device 700 according to one embodiment of the present disclosure. Components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.

[0199] Computing device 700 also includes an access device 740 that enables computing device 700 to communicate via one or more networks 760. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 740 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0200] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 7 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0201] Computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 700 can also be a mobile or stationary server.

[0202] The processor 720 is configured to execute the following computer program / instruction, which implements the steps of the above-mentioned video generation method when executed by the processor.

[0203] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the computing device embodiment is generally similar to the video generation method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the video generation method embodiment.

[0204] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction. When the computer program / instruction is executed by a processor, the steps of the above-mentioned video generation method are implemented.

[0205] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the computer-readable storage medium embodiment is generally similar to the video generation method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the video generation method embodiment.

[0206] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned video generation method when executed by a processor.

[0207] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-mentioned video generation method.

[0208] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0209] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0210] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0211] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0212] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A video generation method, applied to the cloud, comprising: Generate a video subtitle file and an initial video according to the acquired video script; Determining at least one animation material template from an animation material library based on the video subtitle file, determining target attribute information of each animation material template, and generating at least one target animation material based on each target attribute information, wherein the at least one animation material template is determined based on target semantic information corresponding to at least one line of subtitle text after the video subtitle file is split; A target video corresponding to the video script is generated according to each target animation material, the video subtitle file and the initial video.

2. The video generation method according to claim 1, wherein determining at least one animation material template from an animation material library based on the video subtitle file comprises: Splitting the subtitle text in the video subtitle file to obtain at least one line of subtitle text after splitting; At least one animation material template is determined from the animation material library according to each line of subtitle text.

3. The video generation method according to claim 2, wherein determining at least one animation material template from the animation material library based on each line of subtitle text comprises: Inputting each line of subtitle text and the prompt word into a semantic analysis model to obtain target semantic information of each line of subtitle text, wherein the prompt word includes a target analysis type, and the target semantic information is semantic information of the target analysis type; Determining, from the animation material library, an animation material template corresponding to each line of subtitle text according to target semantic information of each line of subtitle text, wherein a semantic label of the animation material template corresponds to the target semantic information; The at least one animation material template is obtained by performing deduplication processing on the animation material templates corresponding to the lines of subtitle text.

4. The video generation method according to claim 3, wherein determining the animation material template corresponding to each line of subtitle text from the animation material library based on the target semantic information of each line of subtitle text comprises: Determining, from the animation material library, a semantic tag that matches the target semantic information of each line of subtitle text according to the target semantic information of each line of subtitle text; An animation material template corresponding to each line of subtitle text is determined from at least one initial material template corresponding to each semantic tag.

5. The video generation method according to any one of claims 1 to 4, wherein determining target attribute information of each animation material template comprises: Determining the attribute fields to be configured corresponding to the animation material templates according to the action objects of the animation material templates, wherein the action objects include the subtitle text in the video subtitle file and the video object in the initial video; By configuring the to-be-configured attribute fields corresponding to the animation material templates, target attribute information of the animation material templates is obtained.

6. The video generation method according to claim 5, wherein the step of obtaining target attribute information of each animation material template by configuring the to-be-configured attribute fields corresponding to each animation material template comprises: For a first animation material template among the at least one animation material template, determining that the attribute field to be configured corresponding to each first animation material template includes a text content field, wherein the first animation material template acts on a subtitle text of the video subtitle file; configuring the text content field using the subtitle text corresponding to each of the first animation material templates to determine target attribute information of each of the first animation material templates; For a second animation material template in at least one animation material template, determining that the attribute fields to be configured corresponding to each second animation material template include an object binding field, wherein the object of the second animation material template is a video object in the initial video; configuring the object binding field using the video object corresponding to each second animation material template to determine target attribute information of each second animation material template; The target attribute information of each animation material template is determined according to the target attribute information of each first animation material template and the target attribute information of each second animation material template.

7. The video generation method according to claim 6, before determining the target attribute information of each animation material template, further comprising: Splitting the subtitle text of the video subtitle file to obtain at least one line of subtitle text after splitting; In the case where the at least one line of subtitle texts includes a plurality of subtitle texts, determining at least two target line of subtitle texts having a target structural relationship from the plurality of line of subtitle texts; Determining a combined animation material template corresponding to the at least two target lines of subtitle text; The configuring the text content field using the subtitle text corresponding to each first animation material template to determine target attribute information of each first animation material template includes: At least two target lines of subtitle text corresponding to the combined animation material template are used to configure at least two text content fields corresponding to the combined animation material template according to the target structural relationship, and target attribute information of the combined animation material template is determined.

8. The video generation method according to any one of claims 1 to 4, 6 to 7, wherein generating a target video corresponding to the video script based on each target animation material, the video subtitle file, and the initial video comprises: In a case where the video subtitle file includes at least one line of subtitle text and text timeline information corresponding to each line of subtitle text, determining the animation timeline information of each target animation material according to the text timeline information corresponding to each line of subtitle text; According to the text timeline information corresponding to each line of subtitle text and the animation timeline information of each target animation material, the each line of subtitle text and the each target animation material are synthesized into the initial video to generate the target video corresponding to the video script.

9. A video generation system, comprising a client and a cloud, wherein: The client is configured to send a video generation request to the cloud in response to an interactive operation on the user interaction interface, wherein the video generation request carries a video script; The cloud is configured to generate a video subtitle file and an initial video based on the acquired video script; determine at least one animation material template from an animation material library based on the video subtitle file, determine target attribute information for each animation material template, and generate at least one target animation material based on each target attribute information; generate a target video corresponding to the video script based on each target animation material, the video subtitle file, and the initial video, wherein the at least one animation material template is determined based on target semantic information corresponding to at least one line of subtitle text after the video subtitle file is split; The client is further configured to receive the target video returned by the cloud and display the target video on the user interaction interface.

10. A video generation device, comprising: A file generation module is configured to generate a video subtitle file and an initial video according to the acquired video script; a material generation module configured to determine, based on the video subtitle file, at least one animation material template from an animation material library, determine target attribute information of each animation material template, and generate at least one target animation material based on each target attribute information, wherein the at least one animation material template is determined based on target semantic information corresponding to at least one line of subtitle text after the video subtitle file is split; The video generation module is configured to generate a target video corresponding to the video script according to each target animation material, the video subtitle file and the initial video.

11. A computing device comprising: memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer program or instructions are executed by the processor, the steps of the video generation method according to any one of claims 1 to 8 are implemented.

12. A computer-readable storage medium storing a computer program or instruction, wherein the computer program or instruction, when executed by a processor, implements the steps of the video generation method according to any one of claims 1 to 8.

13. A computer program product, comprising a computer program or instructions, which implements the steps of the video generation method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Text-based video generation method and system and related equipment

    CN117041459A