Video generation method, system and device
Generate video subtitle files and animation material templates through video scripts, and automatically integrate and generate target videos, solving the problem of inefficient video production in the existing technology, and achieving automation and high efficiency of video animation production.
Patent Information
- Application Number
- CN202510740407.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the existing video production process, video special effects need to be manually added after video synthesis, resulting in inefficiency.
Generate video subtitle files and initial videos through video scripts, use templates and attribute information in the animation material library to generate target animation materials, and automatically integrate them into target videos.
The manual animation editing process after video production is advanced to the automation node, which significantly improves the efficiency of video animation production.
Smart Images

Figure CN120321432A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of video generation, and in particular to a video generation method. One or more embodiments of this specification also relate to a video generation system, a video generation device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] During the process of video editing, adding appropriate special effects to the current video content can significantly enhance the video. With the development of digital human projects, the demand for special effects in high-quality video production is increasing.
[0003] In the original video special effect production process, in the case of synthesizing videos through cloud editing, secondary processing is dependent on the products after video synthesis. That is, the conventional video production process is to generate digital human videos and video subtitle files according to a video script, synthesize the digital human videos and video subtitle files through cloud editing, and then manually add video special effects on the basis of the synthesized video to complete video production. Therefore, when video special effects need to be manually added to the synthesized video, the efficiency of video production is low. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a video generation method. One or more embodiments of this specification also relate to a video generation system, a video generation device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a video generation method is provided, including: Generating a video subtitle file and an initial video according to the obtained video script; Determining at least one animation material template from an animation material library according to the video subtitle file, determining the target attribute information of each animation material template, and generating at least one target animation material through each target attribute information; Generating a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video.
[0006] According to the second aspect of the embodiments of this specification, a video generation system is provided, including a client and a cloud, where The client is configured to send a video generation request to the cloud in response to an interaction operation on a user interface, where the video generation request carries a video script; The cloud is used to generate a video subtitle file and an initial video according to the obtained video script; determine at least one animation material template from an animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information; generate a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video. The client is further used to receive the target video returned by the cloud and display the target video on the user interaction interface.
[0007] According to a third aspect of the embodiments of the present specification, a video generation device is provided, including: A file generation module configured to generate a video subtitle file and an initial video according to the obtained video script; A material generation module configured to determine at least one animation material template from an animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information; A video generation module configured to generate a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video.
[0008] According to a fourth aspect of the embodiments of the present specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above video generation method are implemented.
[0009] According to a fifth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above video generation method are implemented.
[0010] According to a sixth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above video generation method are implemented.
[0011] The video generation method provided by an embodiment of this specification, when generating a video subtitle file and an initial video according to the obtained video script, determines at least one animation material template from the animation material library through the analysis of the video subtitle file, and realizes the synthesis of the target animation material by determining the target attribute information of each animation material template. The at least one target animation material generated can be integrated with the video subtitle file and the initial video to generate a target video corresponding to the video script and including subtitle text and the target animation material; by advancing the manual dynamic effect editing process after video production to the automated nodes of video generation, the dynamic effect generation efficiency of the entire video production is greatly improved, and the efficiency of video dynamic effect production is significantly enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a schematic diagram of the processing process of an existing video generation method provided by an embodiment of this specification; Figure 2 is a schematic diagram of the scenario of a video generation method provided by an embodiment of this specification; Figure 3 is a flowchart of a video generation method provided by an embodiment of this specification; Figure 4 is a schematic diagram of the processing process of a video generation method provided by an embodiment of this specification; Figure 5 is a schematic diagram of the structure of a video generation system provided by an embodiment of this specification; Figure 6 is a schematic diagram of the structure of a video generation device provided by an embodiment of this specification; Figure 7 is a block diagram of the structure of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0014] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0015] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0017] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a foundation model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability, such as large language models (LLMs), multi-modal pre-training models, etc.
[0018] When the large model is actually applied, it only needs to be fine-tuned with a small number of samples for the pre-trained model to be applied to different tasks. The large model can be widely applied to fields such as natural language processing (NLP, Natural Language Processing), computer vision, etc. Specifically, it can be applied to tasks in the field of computer vision such as visual question answering (VQA, Visual Question Answering), image captioning (IC, Image Caption), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0019] First, explain the noun terms involved in one or more embodiments of this specification.
[0020] Cloud editing: It is a video editing production service based on cloud computing and artificial intelligence technologies, which can provide users with core functions such as live editing, video editing, template factory, digital human production, etc., and can use artificial intelligence to assist in editing production. It can be widely applied to industries such as the Internet, culture and media, advertising and marketing, education and finance, etc., to meet the needs of enterprises for large-scale, efficient, convenient, and intelligent video content production.
[0021] Dynamic effect: It refers to the dynamic effect that conveys information or enhances the user experience through visual changes in digital media.
[0022] Lottie: Lottie is an animation file format that allows animations to be displayed consistently across different platforms. It is usually used for animation display in mobile applications.
[0023] Node.js: It is an open-source, cross-platform JavaScript runtime environment that allows developers to run JavaScript on the server side.
[0024] The process of video dynamic effect production is usually as follows: when generating a video subtitle file and an initial video through a video script, synthesize an intermediate video through cloud editing, and on the basis of the intermediate video, manually add video dynamic effects to generate the target video.
[0025] Specifically, it can be seen in Figure 1 , Figure 1The figure shows a schematic diagram of the processing process of an existing video generation method provided by an embodiment of this specification. Taking the generation of a digital human-related video scene based on a video script as an example, when an input video script is received, a digital human video and a video subtitle file are generated based on the video script, and the digital human video and the video subtitle file are synthesized through cloud editing. Then, video special effects are manually added on the basis of the synthesized video, thus completing video production.
[0026] That is, the video special effects are secondary processing depending on the product after video synthesis, with high labor costs and low production efficiency.
[0027] To solve the above technical problems, in this specification, a video generation method is provided. This specification also relates to a video generation system, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0028] See Figure 2 , Figure 2 The figure shows a schematic diagram of a scenario of a video generation method provided by an embodiment of this specification.
[0029] Specifically, this video processing method is applied to a video generation system, which includes a client 202 and a cloud 204. When the client 202 responds to an interaction operation of the user interface, it sends a video generation request to the cloud 204, and the video generation request carries a video script.
[0030] When the cloud 204 receives the video generation request sent by the client 202, it generates a video subtitle file and an initial video according to the obtained video script; determines at least one animation material template from the animation material library according to the video subtitle file, determines the target attribute information of each animation material template, and generates at least one target animation material through each target attribute information; generates the target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video, and returns the target video to the client 202 to display the target video on the client 202.
[0031] The client 202 may include a browser, an APP (Application), a web application such as an H5 (Hyper Text Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The client can be developed based on the software development kit (SDK) of the corresponding service provided by the server, such as developed based on the real-time communication (RTC) SDK. The client can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device can have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, video generation applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0032] The cloud 204 can be understood as a server that provides various services, including cloud servers, such as a server that provides communication services for multiple clients, or a server for background training that supports the models used on the client, or a server that processes the data sent by the client, etc. It should be noted that the cloud 204 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The cloud 204 can also be a server of a distributed system, or a server combined with a blockchain. The cloud 204 can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0033] In the video generation method provided in the embodiments of this specification, when generating a video subtitle file and an initial video according to the obtained video script, by analyzing the video subtitle file, at least one animation material template is determined from the animation material library. By determining the target attribute information of each animation material template, the synthesis of the target animation material is realized. The at least one generated target animation material can be integrated with the video subtitle file and the initial video to generate a target video corresponding to the video script and including subtitle text and the target animation material. By advancing the manual special effect editing process after video production to the automated nodes of video generation, the special effect generation efficiency of the entire video production is greatly improved, and the efficiency of video special effect production is greatly enhanced.
[0034] See Figure 3 , Figure 3 which shows a flowchart of a video generation method provided according to an embodiment of this specification, specifically including the following steps.
[0035] Step 302: Generate a video subtitle file and an initial video according to the obtained video script.
[0036] Among them, the video script can be understood as a kind of text data used to guide the generation and editing of video content. The video subtitle file can be understood as a text data file containing timeline information, used to display the text content corresponding to the audio or video picture during video playback. The initial video can be understood as a primary video version automatically generated based on the video script, and this initial video contains the preliminarily generated video pictures without subtitle text and special effects.
[0037] Specifically, in the case of obtaining the video script, generate a video subtitle file according to the text content in the video script, and generate an initial video according to the video script. It should be noted that in the cloud editing scenario, the subtitle text in the video subtitle file can be synchronized and matched with the video playback timeline of the initial video.
[0038] Step 304: Determine at least one animation material template from the animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information.
[0039] Among them, the animation material template can be understood as a predefined dynamic graphic design framework used to enhance the visual expressiveness of the video. This animation material template contains a static structure (this static structure contains immutable parts such as basic graphics and animation paths) and configurable attribute fields, used to batch generate animation materials with a unified style. Specifically, the animation material template contains configurable attribute fields (such as text content, color, scaling ratio, etc.), and realizes personalized configuration by configuring different attribute information for the attribute fields, that is, different target animation materials can be generated through the same animation material template with different attribute information.
[0040] The target attribute information can be understood as the actual numerical values or logical rules given to the animation material template in a specific application scenario; the target animation material can be understood as the finally available dynamic graphic elements generated by injecting specific attribute information into the animation material template. This target animation material is independent of the main video content and is presented in the form of an overlay; this target animation material can be independently edited or exported as an independent media file, for example, the file format of the target animation material is Lottie or GIF, which is not limited here.
[0041] Specifically, according to the video subtitle file, at least one animation material template matching the subtitle text in the video subtitle file can be determined from the animation material library. The animation material template is a dynamic graphic design prototype without bound specific parameters. The animation material template generates a target animation material by receiving externally input attribute configuration data (i.e., target attribute information, such as specific text content), that is, determines the target attribute information of each animation material template. For one target attribute information, a corresponding target animation material can be determined. The target animation material is a renderable instance with solidified parameters.
[0042] The animation material template includes, but is not limited to, subtitle borders, dynamic icons, etc. Taking the lace subtitle border as an example of the animation material template, the configurable fields corresponding to the lace subtitle border include the text placeholder {text}. The subtitle text in the video subtitle file (such as the subtitle text is "I am a pediatrician") is assigned as the specific text content of the subtitle border to the text placeholder, and the coordinate position of the subtitle border is configured as (50, 80). At this time, the subtitle text and the configured coordinate position are the target attribute information of the subtitle border. According to these configured target attribute information, a subtitle animation effect material (i.e., the target animation material) with a lace subtitle border, a coordinate position of (50, 80), and a displayed text content of "I am a pediatrician" can be generated.
[0043] In fact, through code logic (such as a Node.js script), dynamic data can be injected into the configurable parameters in the animation material template, thereby synthesizing the target animation material.
[0044] In one or more embodiments of this specification, when determining the animation material template from the animation material library based on the video subtitle file, it is necessary to split the subtitle text in the video subtitle file to obtain individual line subtitle texts, and then determine the corresponding animation material template based on the line subtitle texts. The specific implementation is as follows: Determining at least one animation material template from the animation material library according to the video subtitle file includes: Splitting the subtitle text in the video subtitle file to obtain at least one line subtitle text after splitting; Determining at least one animation material template from the animation material library according to each line subtitle text.
[0045] Among them, the line subtitle text can be understood as text data obtained by splitting the subtitle text into lines and including line numbers.
[0046] Specifically, when splitting the subtitle text in a video subtitle file, it can be line - broken according to preset rules. For example, it can be automatically line - broken by a fixed number of characters (20 characters per line); it can also be line - broken according to the text semantics, such as according to punctuation marks (e.g., full stops, commas) or the results of natural language processing (NLP) word segmentation; it can also be manually line - broken, for example, the user forces a line - break through the editing interface. The specific line - breaking method can be selected according to the actual situation and is not limited here.
[0047] For example, if the subtitle text is "Pediatricians need to possess professional medical knowledge, patient communication skills, and rich clinical experience", by splitting this subtitle text through semantic segmentation, the line subtitle texts obtained can be "1. Pediatricians need to possess professional medical knowledge", "2. Patient communication skills", and "3. And rich clinical experience" respectively.
[0048] In practical applications, when obtaining a video subtitle file from a video script, it is also possible to split the video script to obtain a video subtitle file containing line subtitle texts, which is not limited here.
[0049] When obtaining at least one line subtitle text after splitting, at least one animation material template is determined from the animation material library according to each line subtitle text, so as to ensure accurate matching using each line subtitle text.
[0050] The video generation method provided in the embodiments of this specification, when splitting to obtain at least one line subtitle text and determining at least one animation material template from the animation material library according to each line subtitle text, can accurately and dynamically match the animation material template according to the detailed information of each line subtitle text, and at the same time avoid the time - consuming operation of manually selecting animation materials line by line.
[0051] In one or more embodiments of this specification, when obtaining a line subtitle text, an animation material template can be dynamically matched according to the semantic information of the line subtitle text. When obtaining the semantic information corresponding to the line subtitle text through a semantic root system model, the determined semantic information is different according to the different target analysis types in the prompt words. Therefore, a suitable prompt word can be constructed according to actual needs and input into the semantic analysis model together with the line subtitle text. The specific implementation is as follows: Determining at least one animation material template from the animation material library according to each line subtitle text includes: Inputting each line subtitle text and a prompt word into a semantic analysis model to obtain the target semantic information of each line subtitle text, where the prompt word contains a target analysis type, and the target semantic information is the semantic information of the target analysis type; Determine at least one animation material template from the animation material library according to the target semantic information of each line of subtitle text.
[0052] Among them, the semantic analysis model can be understood as a large model, aiming to analyze and understand the deep meaning of the text, including semantic information such as entities, relationships, emotions, intentions, etc.; through the semantic analysis model, natural language can be converted into a structured representation. The target analysis type is used to clearly indicate the direction of the semantic analysis task that the model needs to perform. For example, the semantic analysis task can be sentiment analysis, entity recognition, relationship extraction, etc.
[0053] Specifically, in the embodiment of the specification, the target analysis type is sentiment analysis. For example, the subtitle texts of each line obtained are respectively "1. Everyone should pay attention", "2. The "low blood pressure high" of some patients", "3. It may be related to the drugs taken"; the prompt word is "Please mark the key lines in the subtitle according to the subtitle I provided. And identify his emotional tendency." When inputting the subtitle texts of each line and the prompt word into the semantic analysis model together, the target semantic information of each line of subtitle text obtained is respectively: {"emotion": "warning", "lineNumber": 1, "subtitleContent": "Everyone should pay attention"} {"emotion": "neutral", "lineNumber": 2, "subtitleContent": "The \"low blood pressure high\" of some patients"} {"emotion": "neutral", "lineNumber": 3, "subtitleContent": "It may be related to the drugs taken"} It should be noted that the prompt word can also include example data, so that the semantic analysis model can refer to the example data to better perform semantic analysis and obtain more accurate analysis results.
[0054] When obtaining the subtitle text of each line, at least one animation material template can be matched from the animation material library according to the target semantic information corresponding to each line of subtitle text.
[0055] The video generation method provided in this specification embodiment obtains the target semantic information of each line of subtitle text through semantic analysis of each line of subtitle text. The target semantic information is the semantic information corresponding to the target analysis type in the prompt word, that is, different semantic information of the subtitle text of the line can be obtained through different target analysis types in the prompt word, and then different dimensions of semantic information can be extracted according to needs, so as to match the animation material template related to this dimension.
[0056] In one or more embodiments of the present specification, according to the target semantic information of each line of subtitle text, an animation material template corresponding to each line of subtitle text can be determined from the animation material library. To avoid duplication of animation effects, duplicate removal processing can be performed on the determined animation material templates corresponding to each line of subtitle text. The specific implementation is as follows: Determining at least one animation material template from the animation material library according to the target semantic information of each line of subtitle text includes: According to the target semantic information of each line of subtitle text, an animation material template corresponding to each line of subtitle text is determined from the animation material library, where the semantic label of the animation material template corresponds to the target semantic information; The at least one animation material template is obtained by performing duplicate removal processing on the animation material templates corresponding to each line of subtitle text.
[0057] Specifically, in the case of obtaining the target semantic information corresponding to each line of subtitle text, a corresponding animation material template can be determined from the animation material library for each identified target semantic information. In the case where the target semantic information corresponding to some lines of subtitle text is similar, there may be the same or similar animation material templates matched by different lines of subtitle text. To avoid the repeated use of the same or similar animation material templates from causing visual fatigue to users and reducing the content attractiveness, duplicate removal processing can be performed on the same or similar animation material templates, so as to ensure the differentiation of the generated target animation materials in terms of vision and function. The remaining animation material templates after duplicate removal processing are determined as at least one animation material template determined from the animation material library based on the video subtitle text.
[0058] Of course, if due to the requirement of function consistency, a unified subtitle border needs to be added to the subtitle text, in this case, duplicate removal processing cannot be performed on the unified subtitle border. Therefore, in practical applications, the animation material templates that do not require duplicate removal processing can be marked according to the actual situation to meet the actual requirements.
[0059] The video generation method provided by the embodiments of the present specification, in the case of determining the corresponding animation material template for each line of subtitle text, to avoid duplicate animation effects and improve the freshness of the content, duplicate removal processing can be performed on the determined animation material templates corresponding to each line of subtitle text.
[0060] In one or more embodiments of the present specification, when determining an animation material template matching each target semantic information from the animation material library according to the target semantic information of the line of subtitle text, specifically, the target semantic information is matched with the semantic label of the animation material template. The specific implementation is as follows: Determining the animation material template corresponding to each line of subtitle text from the animation material library according to the target semantic information of each line of subtitle text includes: Determining a semantic label that matches the target semantic information of each line of subtitle text from the animation material library according to the target semantic information of each line of subtitle text; Determining the animation material template corresponding to each line of subtitle text from at least one initial material template corresponding to each semantic label.
[0061] Specifically, a batch of initially marked material templates are pre-stored in the animation material library, that is, each initial material template in the animation material library corresponds to a semantic label. Taking a line of subtitle text as an example, in the case of obtaining the target semantic information of the line of subtitle text, the semantic label that matches the target semantic information can be determined, so as to select an initial material template from multiple corresponding initial material templates as the animation material template corresponding to the line of subtitle text.
[0062] In specific implementation, since multiple initial material templates can correspond to the same semantic label, when selecting an animation material template from the corresponding initial material templates based on the semantic label, the animation material template can be determined by random selection, or the initial material template with the highest selection frequency can be used as the animation material template, which is not limited here.
[0063] For example, when performing sentiment analysis on a line of subtitle text and obtaining that the target semantic information corresponding to the line of subtitle text is "angry", it is determined that the semantic label that matches the target semantic information in the animation material library is "angry". Therefore, the first initial material template is randomly selected from the first initial material template, the second initial material template, the third initial material template, and the fourth initial material template corresponding to this semantic label as the animation material template corresponding to the line of subtitle text.
[0064] In fact, the same keyword requires different dynamic effects in different contexts. For example, the dynamic effects selected for "explosion" in news videos and variety shows are different. Therefore, the semantic label of the animation material template is a multi-dimensional semantic label. For example, the semantic label of the animation material template consists of multi-dimensional labels such as domain labels (which can also be called scene labels), emotion labels, and entity labels. In the case where multiple types of semantic information corresponding to a line of subtitle text can be obtained according to different target analysis types in the prompt words, multiple types of semantic information corresponding to the line of subtitle text can be obtained. Through the multiple types of semantic information, a matching multi-dimensional semantic label is determined, so as to more accurately determine the animation material template from at least one initial material template based on the multi-dimensional semantic label.
[0065] For example, when it is determined that the emotional semantic information corresponding to a line subtitle text is "angry" and the domain semantic information is "medicine", in the animation material library, multi-dimensional tags that match the domain tag of medicine and the emotional tag of angry are searched. Specifically, the full tag combination can be preferentially matched, and if not, backtracking is performed level by level, so as to determine the animation material template from at least one initial material template corresponding to the multi-dimensional semantic tag.
[0066] Of course, in the case of multi-dimensional tags, multiple initial material templates may belong to the same domain and express the same emotion, but the weight values of different-dimensional tags among them may be different, that is, some of the initial material templates emphasize the domain more (the weight value of the domain-dimensional tag is high), and some emphasize the emotion more (the weight value of the emotional-dimensional tag is high). Therefore, when selecting the animation material template from at least one initial material template corresponding to the multi-dimensional tag, the priority of different-dimensional tags can be set, so as to select the initial material template with the most complete tag coverage and the highest total weight as the animation material template.
[0067] For example, when the semantic tags determined according to the target semantic information of the line subtitle text are [medical, serious], the determined initial material templates include initial material template A and initial material template B. Among them, the semantic tags of initial material template A are [medical: 0.8][serious: 0.8], and the semantic tags of initial material template B are [medical: 0.7][serious: 0.9]. When the priority of the emotion-dimensional tag is set higher than that of the domain-dimensional tag, initial material template B is selected as the animation material template.
[0068] The video generation method provided by the embodiments of this specification can pre-store the initial material templates marked with semantics in the animation material library, match the corresponding semantic tags from the animation material library according to the target semantic information corresponding to the line subtitle text, and thus determine the animation material template corresponding to the line subtitle text from at least one initial material template corresponding to the semantic tags. When using the semantic tags to select the animation material template, through the construction of the semantic animation material library, the standardized source of the animation material input is solved.
[0069] In one or more embodiments of this specification, when the animation material template is obtained, by determining the target attribute information of the animation material template, a specific instantiated target animation material can be generated, and the target attribute information of the animation material template can be accurately determined according to the action object and relative position rule of each animation material template. The specific implementation method is as follows: The determination of the target attribute information of each animation material template includes: Determine the to-be-configured attribute fields corresponding to the respective animation material templates according to the objects of action of the respective animation material templates, where the objects of action include the subtitle texts in the video subtitle file and the video objects in the initial video; Obtain the target attribute information of the respective animation material templates by configuring the to-be-configured attribute fields corresponding to the respective animation material templates.
[0070] Among them, the object of action can be understood as the video content element to which the animation material template needs to be bound. For example, in the case where the animation material template is a subtitle border in the above embodiment, the object of action of this animation material template is the subtitle text, and in the case where the animation material template is a "small flame" expressing anger, the object of action of this animation material template is the video object in the initial video (such as a video character in the video).
[0071] The to-be-configured attribute fields can be understood as the configurable attribute fields in the above embodiment. When the objects of action corresponding to the animation material templates are different, the to-be-configured attribute fields are different.
[0072] Actually, the to-be-configured attribute fields may include relative position fields. Specifically, according to the relative position rules of the respective animation material templates, the relative position information corresponding to the relative position fields in the respective animation material templates can be determined. The relative position rules are used to define the spatial relationship between the animation material template and the object of action. For example, the relative position rules can be "centered display", "10 pixels to the right", "follow the character's movement", etc.
[0073] Specifically, determine the subtitle text or video object to which the animation material template is to act, and calculate the relative position information of the animation material template according to the preset relative position rules; for example, the animation material template is a subtitle border, and the relative position rule of this subtitle border is to be 5 pixels outside the text. Therefore, according to this relative position rule, the relative position information of the animation material template can be determined, and this relative position information is filled into the relative position field of the animation material template to obtain the target attribute information.
[0074] The video generation method provided in the embodiments of this specification determines the target attribute information of the animation material template through precise positioning, thereby ensuring the precise cooperation between the animation material template and the object of action; by configuring the to-be-configured attribute fields, the efficiency and quality of video production are greatly improved, and it is ensured that the animation effects are always accurate and in place.
[0075] In one or more embodiments of the present specification, an animation material template with subtitle text as the object in at least one animation material template is determined as the first animation material template, and an animation material template with a video object as the object in at least one animation material template is determined as the second animation material template; by determining the text position coordinates of the subtitle text corresponding to the first animation material template and the relative position rule corresponding to the first type of animation material template, the target position coordinates of the first animation material template are determined; by determining the object position coordinates of the video object corresponding to the second animation material template and the relative position rule corresponding to the second type of animation material template, the target position coordinates of the second animation material template are determined. The specific implementation is as follows: The obtaining of the target attribute information of each animation material template by configuring the to-be-configured attribute fields corresponding to each animation material template includes: For the first animation material template in at least one animation material template, it is determined that the to-be-configured attribute fields corresponding to each first animation material template include a text content field, where the acting object of the first animation material template is the subtitle text of the video subtitle file; The text content field is configured using the subtitle text corresponding to each first animation material template to determine the target attribute information of each first animation material template; For the second animation material template in at least one animation material template, it is determined that the to-be-configured attribute fields corresponding to each second animation material template include an object binding field, where the acting object of the second animation material template is the video object in the initial video; The object binding field is configured using the video object corresponding to each second animation material template to determine the target attribute information of each second animation material template; According to the target attribute information of each first animation material template and the target attribute information of each second animation material template, the target attribute information of each animation material template is determined.
[0076] Specifically, an animation material template with subtitle text as the object in at least one animation material template is determined as the first animation material template, and an animation material template with a video object as the object in at least one animation material template is determined as the second animation material template.
[0077] In fact, when the acting objects of the animation material templates are different, the to-be-configured attribute fields corresponding to the animation material templates are different. When the acting object of the animation material template is subtitle text, the to-be-configured attribute fields corresponding to the animation material template include a text content field. When the acting object of the animation material template is a video object, the to-be-configured attribute fields corresponding to the animation material template include an object binding field.
[0078] For the first animation material template, the to-be-configured attribute fields corresponding to each first animation material template include a text content field. By determining the subtitle text corresponding to the first animation material template as the specific parameter value of the text content field, the configuration of the text content field is achieved, thereby obtaining the target attribute information of the first animation material template.
[0079] For the second animation material template, the to-be-configured attribute fields corresponding to each second animation material template include an object binding field. By determining the video object (object identifier of the video object) corresponding to the second animation material template as the specific parameter value of the object binding field, the configuration of the object binding field is achieved, thereby obtaining the target attribute information of the first animation material template.
[0080] For example, it is recognized that the subtitle text is "This product adopts innovative technology", and the video picture corresponding to this subtitle text is a mobile phone product being shown. The animation material templates matched for this subtitle text include a subtitle border with a sci-fi flicker (acting on the subtitle text) and a shiny highlight line (acting on the video object); when the relative position rule of this subtitle border is 10 pixels below the text, determine the parameter value of the relative position field corresponding to this subtitle border as "10 pixels below the text", and determine the parameter value of the text content field corresponding to this subtitle border as "This product adopts innovative technology", thereby obtaining an instantiated (parameterized) subtitle animation effect material (i.e., the target animation material) by performing parameter configuration on the to-be-configured attribute fields of the subtitle border; determine the parameter value of the object binding field corresponding to the line as "mobile phone product phone", thereby realizing the binding between the line and the mobile phone product.
[0081] The video generation method provided in the embodiments of this specification determines the to-be-configured attribute fields corresponding to each animation material template for animation material templates applied to different acting objects, and performs parameterized configuration on the to-be-configured attribute fields through the specific information of the acting object, thereby accurately obtaining the target attribute information corresponding to each animation material template.
[0082] In one or more embodiments of this specification, when splitting the subtitle text of the video subtitle file to obtain multiple split line subtitle texts, determine at least two target line subtitle texts having a target structural relationship from the multiple line subtitle texts, and determine the combined animation material template corresponding to the at least two target line subtitle texts. The specific implementation is as follows: Before determining the target attribute information of each animation material template, it further includes: Split the subtitle text of the video subtitle file to obtain at least one line subtitle text after splitting; In the case that the at least one line subtitle text contains multiple ones, determine at least two target line subtitle texts having a target structural relationship from the multiple line subtitle texts; Determine a combined animation material template corresponding to the at least two target line subtitle texts.
[0083] Configuring the text content field by using the subtitle text corresponding to each first animation material template, and determining the target attribute information of each first animation material template, includes: Using the at least two target line subtitle texts corresponding to the combined animation material template, configuring the at least two text content fields corresponding to the combined animation material template according to the target structural relationship, and determining the target attribute information of the combined animation material template.
[0084] Among them, the target structural relationship can be understood as a juxtaposed / progressive relationship that is adjacent, structurally similar, and semantically related. For example, a subtitle text is "The operation process is as follows: First step, input data; Second step, run the algorithm; Third step, analyze the result." When splitting this subtitle text, the obtained line subtitle texts are respectively "The operation process is as follows:", "First step, input data;", "Second step, run the algorithm;", "Third step, analyze the result."; At this time, "First step, input data;", "Second step, run the algorithm;", "Third step, analyze the result." are the target line subtitle texts having a target structural relationship.
[0085] Actually, in the case of obtaining at least one line subtitle text after splitting, an animation material template can correspond to a line subtitle text. For example, a subtitle text is "Everyone should pay attention. Some patients have high diastolic blood pressure, which may be related to the drugs taken." After splitting the subtitle text, 3 line subtitle texts are obtained, which are respectively "Everyone should pay attention.", "Some patients have high diastolic blood pressure,", "which may be related to the drugs taken.", and two animation material templates are determined to act on the above line subtitle texts. Specifically, animation material template A corresponds to the line subtitle text "Everyone should pay attention.", and animation material template B corresponds to the line subtitle text "which may be related to the drugs taken.". Therefore, the text content field of animation material template A is configured by the line subtitle text "Everyone should pay attention." to generate target animation material A, and the text content field of animation material template B is configured by the line subtitle text "which may be related to the drugs taken." to generate target animation material B.
[0086] For at least two target line subtitle texts having a target structural relationship, a corresponding combined animation material template can be determined. The combined animation material template can strengthen the juxtaposed / progressive relationship between the target line subtitle texts through elements such as serial numbers, connecting lines, and cards, and each target line subtitle text can appear dynamically one by one in the combined animation material template.
[0087] Specifically, the combined animation material template corresponds to at least two text content fields. Therefore, the combined animation material template can act on at least two target line subtitle texts. By using at least two target line subtitle texts as the parameter values of at least two text content fields respectively, the parametric configuration of at least two text content fields in the combined animation material template is realized, and the target attribute information of the combined animation material template is determined.
[0088] Continuing with the above example, when determining that "Step 1, input data;", "Step 2, run the algorithm;", and "Step 3, analyze the results." are target line subtitle texts with a target structural relationship, the corresponding combined animation material template C is determined. The text content fields of the combined animation material template C include text1, text2, and text3, and it is set that text1 appears first and text3 appears last. Therefore, according to the target structural relationship, "Step 1, input data;" is determined as the parameter value of the text content field text1, "Step 2, run the algorithm;" is determined as the parameter value of the text content field text2, and "Step 3, analyze the results." is determined as the parameter value of the text content field text3, generating the target combined animation material C, thereby realizing the effect that the target line subtitle texts with a target structural relationship can appear dynamically one by one by using the combined animation material template.
[0089] Actually, when the action object of the animation material template is subtitle text and the configurable parameters of the animation material template include text content, by using the subtitle text as the text content of the animation material template and combining attribute information such as relative position information, the target animation material can be directly generated; by inserting the subtitle text and the generated target animation material into the initial video, the target video including subtitles and special effects is generated.
[0090] The video generation method provided in the embodiments of this specification splits the subtitle text of the video subtitle file to obtain at least one line subtitle text after splitting, determines at least two target line subtitle texts with a target structural relationship and the combined animation material template corresponding to the at least two target line subtitle texts, and configures at least two text content fields of the combined animation material template through the at least two target line subtitle texts, realizing a rich animation display effect.
[0091] Step 306: Generate the target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video.
[0092] Specifically, when obtaining a video subtitle file, an initial video, and target animation materials, based on the initial video, combining the video subtitle file and the target animation materials to obtain a target video corresponding to the video script, which includes subtitle content and special effects.
[0093] In one or more embodiments of this specification, based on the initial video, the target animation materials and the video subtitle file are superimposed on the initial video as independent tracks. Since the target animation materials correspond to the subtitle texts of the corresponding lines, the display duration of the target animation materials is controlled by the timeline information of the subtitle texts of the corresponding lines. The specific implementation is as follows: Generating the target video corresponding to the video script according to the respective target animation materials, the video subtitle file, and the initial video includes: When the video subtitle file includes at least one line of subtitle text and the timeline information corresponding to each line of subtitle text, determining the animation timeline information of each target animation material according to the text timeline information corresponding to each line of subtitle text; According to the text timeline information corresponding to each line of subtitle text and the animation timeline information of each target animation material, synthesizing each line of subtitle text and each target animation material into the initial video to generate the target video corresponding to the video script.
[0094] Specifically, by parsing the video subtitle file, the text timeline information corresponding to each line of subtitle text can be obtained. For example, the start time of the first line of subtitle text is "00:00:01.000" and the end time is "00:00:04.000"; when the target animation material corresponds to the first line of subtitle text, the animation timeline information of the corresponding target animation material can be determined according to the text timeline information corresponding to the first line of subtitle text. For example, when the first line of subtitle text corresponds to target animation material A, the start time of target animation material A can be determined to be "00:00:01.000" and the end time to be "00:00:04.000".
[0095] Of course, in the above case, the target animation material appears and disappears at the same time as the associated subtitle text by default. In fact, further settings can be made. For example, by setting the lead (the target animation material appears N seconds earlier than the subtitle text), the lag (the target animation material disappears M seconds later than the subtitle text), etc., the animation timeline information of the target animation material can be dynamically determined.
[0096] Based on the initial video as the base video layer, each line of subtitle text and each target animation material are used as the text layer and the animation layer respectively. According to the text timeline information corresponding to each line of subtitle text and the animation timeline information of each target animation material, they are superimposed on the base video layer, that is, synthesized into the initial video; when each line of subtitle text and each target animation material are superimposed on the initial video, their positions in the initial video can be adjusted according to the actual situation, so as to generate a target video with better visual perception.
[0097] The video generation method provided in the embodiments of this specification obtains the text timeline information corresponding to each line of subtitle text by parsing, determines the animation timeline information of the target animation material corresponding to the line of subtitle text, and then accurately superimposes each line of subtitle text and each target animation material on the initial video according to the timeline information of each line of subtitle text and each target animation material, completing the production of the target video.
[0098] The video generation method provided in the embodiments of this specification realizes the standardized extraction ability of the dynamic effects in the video script by semantically labeling the animation material template and matching the semantic labels of the animation material template with the target semantic information recognized by the large model, thus providing a basis for the batch automated production of video dynamic effects. Finally, it achieves the time sequence left shift of the video dynamic effect production process, advancing from the post-processing process of video production to an automated node in the video production pipeline, greatly improving the efficiency of video dynamic effect production.
[0099] See Figure 4 , Figure 4 shows a schematic diagram of the processing process of a video generation method provided in an embodiment of this specification.
[0100] Specifically, a video subtitle file and a digital human video are generated according to the input video script, and the digital human video contains a digital human with a virtual image.
[0101] Build a material library (i.e., the animation material library in the above embodiments), which stores dynamically effective material templates (i.e., the animation material templates in the above embodiments) that have been semantically labeled. The specific semantic types include but are not limited to summary, question, misunderstanding, shock, etc.
[0102] Intelligently analyze the video script through a large model (i.e., the semantic analysis model in the above embodiments), and combine the matching of the material library to generate a dynamic effect task corresponding to the video script. Specifically, in the case of obtaining a video subtitle file through the video script, call the large model to perform intelligent semantic analysis on the line subtitle text after separating the lines in the video subtitle file, so as to identify the target semantic information corresponding to the line subtitle text, and determine the dynamically effective material template corresponding to the line subtitle text by matching the target semantic information of the line subtitle text with the semantic labels of the dynamically effective material templates in the material library.
[0103] The dynamic effect production for the dynamic effect material template is realized by calling the Node service. Specifically, the dynamic effect task is synthesized and output through Node.js, that is, the parameter values of the fields that can be dynamically replaced in the dynamic effect material template are filled in combination with the line subtitle text, so as to generate the target animation material.
[0104] The target video is synthesized by the cloud-edited video subtitle file, the synthesized target animation material, and the digital human video, so as to complete the production of the target video containing subtitle content and dynamic effects, and realize the automatic insertion of dynamic effects during the video cloud editing process.
[0105] The video generation method provided by the embodiments of this specification realizes the automation ability of dynamic effect production, advances the manual editing process after video production to the pre-automation node of video cloud editing. The whole process is highly closed-loop, avoiding the inefficiency of manual operations, ensuring the efficient execution of tasks, greatly improving the dynamic effect generation efficiency of the entire video production, and providing a basis for large-scale dynamic effect video production.
[0106] Corresponding to the above method embodiments, this specification also provides embodiments of a video generation system. Figure 5 The structural schematic diagram of a video generation system provided by an embodiment of this specification is shown. As Figure 5 shown, the video generation system includes: a client 202 and a cloud 204. Among them, The client 202 is used to send a video generation request to the cloud in response to an interaction operation on the user interface, where the video generation request carries a video script. The cloud 204 is used to generate a video subtitle file and an initial video according to the obtained video script; determine at least one animation material template from the animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information; generate the target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video. The client 202 is further used to receive the target video returned by the cloud 204 and display the target video on the user interface.
[0107] This interaction operation can be understood as an upload operation of the video script or an input operation of the video script, which is not limited here. For the specific implementation, reference can be made to the above embodiments, which will not be elaborated here.
[0108] The above is a schematic solution of a video generation system according to this embodiment. It should be noted that the technical solution of this video generation system and the technical solution of the above video generation method belong to the same concept. For the details not described in detail in the technical solution of the video generation system, reference can be made to the description of the technical solution of the above video generation method.
[0109] Corresponding to the above method embodiment, this specification also provides an embodiment of a video generation device. Figure 6 The following shows a schematic structural diagram of a video generation device provided by an embodiment of this specification. As Figure 6 shown, this video generation device includes: A file generation module 602, configured to generate a video subtitle file and an initial video according to the obtained video script; A material generation module 604, configured to determine at least one animation material template from an animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information; A video generation module 606, configured to generate a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video.
[0110] Optionally, the material generation module 604 is further configured to: Split the subtitle text in the video subtitle file to obtain at least one line subtitle text after splitting; Determine at least one animation material template from the animation material library according to each line subtitle text.
[0111] Optionally, the material generation module 604 is further configured to: Input each line subtitle text and a prompt word into a semantic analysis model to obtain the target semantic information of each line subtitle text, where the prompt word includes a target analysis type, and the target semantic information is the semantic information of the target analysis type; Determine at least one animation material template from the animation material library according to the target semantic information of each line subtitle text.
[0112] Optionally, the material generation module 604 is further configured to: Determine the animation material template corresponding to each line subtitle text from the animation material library according to the target semantic information of each line subtitle text, where the semantic label of the animation material template corresponds to the target semantic information; Obtain the at least one animation material template by performing a duplicate removal process on the animation material templates corresponding to each line subtitle text.
[0113] Optionally, the material generation module 604 is further configured to: Determine semantic tags matching the target semantic information of each line of subtitle text from the animation material library; Determine the animation material templates corresponding to each line of subtitle text from at least one initial material template corresponding to each semantic tag.
[0114] Optionally, the material generation module 604 is further configured to: Determine the to-be-configured attribute fields corresponding to each animation material template according to the objects of action of each animation material template, where the objects of action include subtitle text in the video subtitle file and video objects in the initial video; Obtain the target attribute information of each animation material template by configuring the to-be-configured attribute fields corresponding to each animation material template.
[0115] Optionally, the material generation module 604 is further configured to: For a first animation material template among at least one animation material template, determine that the to-be-configured attribute fields corresponding to each first animation material template include text content fields, where the object of action of the first animation material template is subtitle text in the video subtitle file; Configure the text content fields with the subtitle text corresponding to each first animation material template to determine the target attribute information of each first animation material template; For a second animation material template among at least one animation material template, determine that the to-be-configured attribute fields corresponding to each second animation material template include object binding fields, where the object of action of the second animation material template is a video object in the initial video; Configure the object binding fields with the video objects corresponding to each second animation material template to determine the target attribute information of each second animation material template; Determine the target attribute information of each animation material template according to the target attribute information of each first animation material template and the target attribute information of each second animation material template.
[0116] Optionally, the material generation module 604 is further configured to: Split the subtitle text of the video subtitle file to obtain at least one line of subtitle text after splitting; In the case where there are multiple lines of subtitle text among the at least one line of subtitle text, determine at least two target lines of subtitle text having a target structural relationship from the multiple lines of subtitle text; Determine the combined animation material template corresponding to the at least two target line subtitle texts.
[0117] Optionally, the material generation module 604 is further configured to: Use the at least two target line subtitle texts corresponding to the combined animation material template to configure the at least two text content fields corresponding to the combined animation material template according to the target structural relationship, and determine the target attribute information of the combined animation material template.
[0118] Optionally, the video generation module 606 is further configured to: When the video subtitle file includes at least one line subtitle text and the text timeline information corresponding to each line subtitle text, determine the animation timeline information of each target animation material according to the text timeline information corresponding to each line subtitle text; According to the text timeline information corresponding to each line subtitle text and the animation timeline information of each target animation material, synthesize each line subtitle text and each target animation material into the initial video to generate the target video corresponding to the video script.
[0119] The video generation device provided in the embodiments of this specification, when generating a video subtitle file and an initial video according to the obtained video script, determines at least one animation material template from the animation material library through the analysis of the video subtitle file, and realizes the synthesis of the target animation material by determining the target attribute information of each animation material template. The at least one target animation material generated can be integrated with the video subtitle file and the initial video to generate a target video corresponding to the video script and including subtitle texts and target animation materials; by advancing the manual dynamic effect editing process after video production to the automated nodes of video generation, the dynamic effect generation efficiency of the entire video production is greatly improved, and the efficiency of video dynamic effect production is greatly improved.
[0120] The above is a schematic solution of a video generation device according to an embodiment of this specification. It should be noted that the technical solution of this video generation device and the technical solution of the above video generation method belong to the same concept. For the details not described in the technical solution of the video generation device, reference can be made to the description of the technical solution of the above video generation method.
[0121] Figure 7 FIG. shows a structural block diagram of a computing device 700 according to an embodiment of this specification. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.
[0122] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interfaces (e.g., network interface controller (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0123] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 7 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0124] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0125] Among them, the processor 720 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above video generation method are implemented.
[0126] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiment of the computing device, since it is basically similar to the embodiment of the video generation method, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the embodiment of the video generation method.
[0127] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned video generation method are implemented.
[0128] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiment of the computer-readable storage medium, since it is basically similar to the embodiment of the video generation method, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the embodiment of the video generation method.
[0129] An embodiment of this specification also provides a computer program product including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned video generation method are implemented.
[0130] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the above-mentioned video generation method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above-mentioned video generation method.
[0131] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments and still achieve the desired result. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0132] The computer instructions include computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0133] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0134] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0135] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A video generation method, applied to the cloud, comprising: Generating a video subtitle file and an initial video according to the obtained video script; Determining at least one animation material template from an animation material library according to the video subtitle file, determining target attribute information of each animation material template, and generating at least one target animation material through the target attribute information; Generating a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video.
2. The video generation method according to claim 1, wherein the determining at least one animation material template from the animation material library according to the video subtitle file comprises: Splitting the subtitle text in the video subtitle file to obtain at least one line subtitle text after splitting; Determining at least one animation material template from the animation material library according to each line subtitle text.
3. The video generation method according to claim 2, wherein the determining at least one animation material template from the animation material library according to each line subtitle text comprises: Inputting each line subtitle text and a prompt word into a semantic analysis model to obtain target semantic information of each line subtitle text, wherein the prompt word contains a target analysis type, and the target semantic information is semantic information of the target analysis type; Determining an animation material template corresponding to each line subtitle text from the animation material library according to the target semantic information of each line subtitle text, wherein a semantic label of the animation material template corresponds to the target semantic information; Obtaining the at least one animation material template by performing duplicate removal processing on the animation material templates corresponding to each line subtitle text.
4. The video generation method according to claim 3, wherein the determining an animation material template corresponding to each line subtitle text from the animation material library according to the target semantic information of each line subtitle text comprises: Determining a semantic label matching the target semantic information of each line subtitle text from the animation material library according to the target semantic information of each line subtitle text; Determining an animation material template corresponding to each line subtitle text from at least one initial material template corresponding to each semantic label.
5. The video generation method according to any one of claims 1-4, wherein the determining target attribute information of each animation material template comprises: Determining a to-be-configured attribute field corresponding to each animation material template according to an action object of each animation material template, wherein the action object comprises subtitle text in the video subtitle file and a video object in the initial video; Obtaining the target attribute information of each animation material template by configuring the to-be-configured attribute field corresponding to each animation material template.
6. The video generation method according to claim 5, wherein the obtaining the target attribute information of each animation material template by configuring the to-be-configured attribute field corresponding to each animation material template comprises: For the first animation material template among at least one animation material template, determine that the to-be-configured attribute fields corresponding to each first animation material template include a text content field, where the object of action of the first animation material template is the subtitle text of the video subtitle file; Configure the text content field with the subtitle text corresponding to each first animation material template to determine the target attribute information of each first animation material template; For the second animation material template among at least one animation material template, determine that the to-be-configured attribute fields corresponding to each second animation material template include an object binding field, where the object of action of the second animation material template is the video object in the initial video; Configure the object binding field with the video object corresponding to each second animation material template to determine the target attribute information of each second animation material template; Determine the target attribute information of each animation material template according to the target attribute information of each first animation material template and the target attribute information of each second animation material template.
7. The video generation method according to claim 6, before determining the target attribute information of each animation material template, further comprising: Split the subtitle text of the video subtitle file to obtain at least one line subtitle text after splitting; In the case that there are multiple ones among the at least one line subtitle text, determine at least two target line subtitle texts having a target structural relationship from the multiple line subtitle texts; Determine the combined animation material template corresponding to the at least two target line subtitle texts; The configuring the text content field with the subtitle text corresponding to each first animation material template to determine the target attribute information of each first animation material template includes: Configure at least two text content fields corresponding to the combined animation material template according to the at least two target line subtitle texts corresponding to the combined animation material template in accordance with the target structural relationship to determine the target attribute information of the combined animation material template.
8. The video generation method according to any one of claims 1-4, 6-7, the generating the target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video includes: In the case that the video subtitle file includes at least one line subtitle text and the text time axis information corresponding to each line subtitle text, determine the animation time axis information of each target animation material according to the text time axis information corresponding to each line subtitle text; Synthesize each line subtitle text and each target animation material into the initial video according to the text time axis information corresponding to each line subtitle text and the animation time axis information of each target animation material to generate the target video corresponding to the video script.
9. A video generation system includes a client and a cloud, where The client is configured to send a video generation request to the cloud in response to an interaction operation of a user interface, where the video generation request carries a video script; The cloud is used to generate a video subtitle file and an initial video according to the obtained video script; determine at least one animation material template from the animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information; generate a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video. The client is further used to receive the target video returned by the cloud and display the target video on the user interaction interface.
10. A video generation device, comprising: A file generation module configured to generate a video subtitle file and an initial video according to the obtained video script; A material generation module configured to determine at least one animation material template from the animation material library according to the video subtitle file, determine the target attribute information of each animation material template, and generate at least one target animation material through each target attribute information; A video generation module configured to generate a target video corresponding to the video script according to each target animation material, the video subtitle file, and the initial video.
11. A computing device, comprising: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the video generation method according to any one of claims 1 to 8 are implemented.
12. A computer-readable storage medium storing computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the video generation method according to any one of claims 1 to 8 are implemented.
13. A computer program product comprising computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the video generation method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Animation make method, device, terminal and medium
CN109300179A
Video generation method and device
CN112866776A
Text-based video generation method and system and related equipment
CN117041459A
Multimedia resource editing method and device, equipment, storage medium and program product
CN118860232A
Video generation method and device, computing equipment, storage medium and program product
CN119299800A
Cited By
Real-time digital human video generation method and device, electronic equipment and storage medium
CN121309905A