Video generation method and device, storage medium and electronic equipment

By processing documents through a storyboard generation model to generate video storyboard information, images, and audio, the problem of low efficiency in manual video generation is solved, and high efficiency in automated video generation is achieved.

CN121815037APending Publication Date: 2026-04-07CHINA TOWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies rely on manual methods to generate videos from documents, resulting in low video generation efficiency.

Method used

The target document is processed by a storyboard generation model to generate storyboard information for multiple video storyboards. Storyboard images and audio are then generated based on the storyboard information, and finally, the target video is synthesized.

Benefits of technology

It achieves automated video generation, improves video generation efficiency, avoids manual intervention, and enhances generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815037A_ABST
    Figure CN121815037A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and device, a storage medium and electronic equipment. Relates to the field of artificial intelligence, and the method comprises the steps: obtaining a target document which comprises information used for describing a to-be-generated target video; processing the target document through a sub-mirror generation model to obtain sub-mirror information of a plurality of video sub-mirrors; for each video sub-mirror, generating a sub-mirror image and a sub-mirror audio according to the sub-mirror information of the video sub-mirror; and generating a target video corresponding to the target document according to the sub-mirror images and the sub-mirror audios of the plurality of video sub-mirrors. Through the method and the device, the problem of low video generation efficiency due to dependence on manual generation of the video according to the document in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to a video generation method and device, a storage medium and an electronic device. BACKGROUND

[0002] In the current era of digital transformation and high popularity of multimedia content, the demand for converting document information into video format is increasing, especially in the fields of education, training, propaganda and information dissemination. This change not only improves the attractiveness and dissemination efficiency of content, but also adapts to the receiving habits of different audiences, especially in the information consumption environment dominated by mobile devices and social media.

[0003] Currently, the related art usually relies on manual generation of videos from documents, resulting in low video generation efficiency. In view of the above problems in the related art, no effective solutions have been proposed so far. SUMMARY

[0004] The main purpose of the present application is to provide a video generation method, device, storage medium and electronic device to solve the problem of low video generation efficiency caused by relying on manual generation of videos from documents in the related art.

[0005] In order to achieve the above purpose, according to one aspect of the present application, a video generation method is provided. The method comprises: obtaining a target document, wherein the target document includes information for describing a target video to be generated; processing the target document through a shot generation model to obtain shot information of a plurality of video shots; for each video shot, generating a shot image and a shot audio according to the shot information of the video shot; and generating a target video corresponding to the target document according to the shot images and the shot audios of the plurality of video shots.

[0006] Optionally, the video generation method further comprises: inputting a first prompt word and the target document into the shot generation model to obtain a document content structure of the target document through processing by the shot generation model, wherein the first prompt word is used to guide the shot generation model to determine the document content structure according to the target document; inputting a second prompt word into the shot generation model to obtain key content elements in the target document through processing by the shot generation model, wherein the second prompt word is used to guide the shot generation model to determine the key content elements according to the target document; and inputting a third prompt word into the shot generation model to obtain the shot information of the plurality of video shots through processing by the shot generation model, wherein the third prompt word is used to guide the shot generation model to generate the shot information of the plurality of video shots according to the target document, the document content structure and the key content elements.

[0007] Optionally, the video generation method further comprises: obtaining a script description text from the script information, and generating a script image according to the script description text; obtaining a script dialogue from the script information, and generating a script audio according to the script dialogue.

[0008] Optionally, the video generation method further comprises: generating an initial script image according to the script description text; determining a script sequence position of a video script to which the script description text belongs in the plurality of video scripts; determining a matching image animation effect of the script sequence position; and determining the script image according to the initial script image and the image animation effect.

[0009] Optionally, the video generation method further comprises: for each video script, performing a synthesis processing on the script image and the script audio of the video script to obtain a script sub-video; and generating the target video according to the script sub-video of each video script.

[0010] Optionally, the video generation method further comprises: determining a plurality of video script pairs according to the plurality of video scripts, wherein each video script pair is composed of two adjacent video scripts in the plurality of video scripts; for each video script pair, determining a content correlation degree between the two video scripts in the video script pair; determining an animation transition effect between the two video scripts in the video script pair according to the content correlation degree; and determining the target video according to the script sub-video of each video script and the animation transition effect.

[0011] Optionally, the video generation method further comprises: generating an initial script audio according to the script dialogue; and performing an audio post-processing on the initial script audio to obtain the script audio, wherein the audio post-processing comprises at least one of the following: noise reduction processing, equalizer processing, and compression processing.

[0012] To achieve the above object, according to another aspect of the present application, a video generation device is provided. The device comprises: an acquisition module configured to acquire a target document, wherein the target document comprises information for describing a target video to be generated; a processing module configured to process the target document by a script generation model to obtain script information of a plurality of video scripts; a first generation module configured to, for each video script, generate a script image and a script audio according to the script information of the video script; and a second generation module configured to generate a target video corresponding to the target document according to the script image and the script audio of each video script.

[0013] Optionally, the processing module further comprises: a first processing submodule configured to input the first prompt word and the target document into the split shot generation model, and obtain a document content structure of the target document by processing the split shot generation model, wherein the first prompt word is used to guide the split shot generation model to determine the document content structure according to the target document; a second processing submodule configured to input the second prompt word into the split shot generation model, and obtain a key content element in the target document by processing the split shot generation model, wherein the second prompt word is used to guide the split shot generation model to determine the key content element according to the target document; and a third processing submodule configured to input the third prompt word into the split shot generation model, and obtain split shot information of the plurality of video split shots by processing the split shot generation model, wherein the third prompt word is used to guide the split shot generation model to generate the split shot information of the plurality of video split shots according to the target document, the document content structure and the key content element.

[0014] Optionally, the first generation module further comprises: a first generation submodule configured to obtain a split shot description text from the split shot information, and generate a split shot image according to the split shot description text; and a second generation submodule configured to obtain a split shot script from the split shot information, and generate a split shot audio according to the split shot script.

[0015] Optionally, the first generation submodule further comprises: a first generation unit configured to generate an initial split shot image according to the split shot description text; a first determination unit configured to determine a split shot sequence position of a video split shot to which the split shot description text belongs in the plurality of video split shots; a second determination unit configured to determine an image animation effect matched with the split shot sequence position; and a third determination unit configured to determine the split shot image according to the initial split shot image and the image animation effect.

[0016] Optionally, the second generation module further comprises: a fourth processing submodule configured to, for each video split shot, synthesize the split shot image and the split shot audio of the video split shot to obtain a split shot sub-video; and a third generation submodule configured to generate the target video according to the split shot sub-video of each of the plurality of video split shots.

[0017] Optionally, the third generation submodule further comprises: a fourth determination unit configured to determine a plurality of video split shot pairs from the plurality of video split shots, wherein each video split shot pair is composed of two adjacent video split shots in the plurality of video split shots; a fifth determination unit configured to determine, for each video split shot pair, a content correlation degree between the two video split shots of the video split shot pair; a sixth determination unit configured to determine an animation transition effect between the two video split shots in the video split shot pair according to the content correlation degree; and a seventh determination unit configured to determine the target video according to the split shot sub-video of each of the plurality of video split shots and the animation transition effect.

[0018] Optionally, the second generating sub-module further comprises: a second generating unit, configured to generate initial storyboard audio according to the storyboard script; and a processing unit, configured to perform audio post-processing on the initial storyboard audio to obtain the storyboard audio, wherein the audio post-processing comprises at least one of the following: noise reduction processing, equalizer processing, and compression processing.

[0019] To achieve the above object, according to another aspect of the present application, a computer readable storage medium is provided, which comprises a stored executable program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the video generation method described above when the executable program is run.

[0020] To achieve the above object, according to another aspect of the present application, an electronic device is provided, which comprises a memory storing an executable program, and a processor configured to run the program, wherein the program is configured to execute the video generation method described above when the program is run.

[0021] To achieve the above object, according to another aspect of the present application, a computer program product is provided, which comprises computer instructions configured to implement the steps of the video generation method described above when executed by a processor.

[0022] In the embodiments of the present application, the target document is processed by using the storyboard generation model to obtain the storyboard information of the plurality of video storyboards, so that the plurality of storyboards required in the to-be-generated video are automatically determined according to the content of the target document. The storyboard image and the storyboard audio are generated according to the storyboard information of the video storyboard, so that the automatic generation of each storyboard content in the video is implemented. The target video corresponding to the target document is generated according to the storyboard image and the storyboard audio of the plurality of video storyboards, so that the video is automatically generated based on the storyboard content, thereby avoiding the dependence on manual video generation and effectively improving the video generation efficiency.

[0023] Therefore, the method provided by the present application achieves the purpose of automatically generating a video according to a document, and achieves the technical effect of improving the video generation efficiency, thereby solving the technical problem of low video generation efficiency caused by the dependence on manual video generation according to a document in the related art. BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application, together with their description, are intended to explain the present application, and are not intended to limit the present application in any manner. In the drawings: Figure 1 is a hardware structure block diagram of a computer terminal provided according to an embodiment of the present application; Figure 2 is a flowchart of a video generation method provided according to an embodiment of the present application; Figure 3This is a schematic diagram of the operation of the target processing system provided in the embodiments of this application; Figure 4 This is a schematic diagram of the operation of the application development platform provided in the embodiments of this application; Figure 5 This is a schematic diagram of the target processing device provided according to an embodiment of this application; Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.

[0028] Example 1

[0029] According to an embodiment of this application, an embodiment of a video generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0030] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a video generation method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor (MCU) or a field-programmable gate array (FPGA), etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output (I / O) interface, a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0031] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0032] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned video generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0033] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0034] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0035] Under the aforementioned operating environment, this application provides the following: Figure 2 The video generation method shown. Figure 2 This is a flowchart of a video generation method according to Embodiment 1 of this application.

[0036] Step S201: Obtain the target document, wherein the target document includes information describing the target video to be generated.

[0037] Optionally, electronic devices, application systems, servers, or other similar devices can be used as the execution subject of this application. In this embodiment, the target processing system is used as the execution subject to execute the video generation method described above.

[0038] The target document is a document containing text content, including at least information describing the target video to be generated, serving as the raw material and basis for target video generation. Users can upload documents to the target processing system via an interface or API. In one optional embodiment, the target processing system directly uses the user-uploaded document as the target document. In another optional embodiment, the target processing system receives and verifies the uploaded document format, while performing a preliminary security check (such as a virus scan), and then uses the user-uploaded document as the target document if the document format meets the requirements and the security check passes.

[0039] Converting target documents into videos can enhance content appeal and dissemination efficiency, such as in education, training, publicity, and information dissemination. The content of the target document varies depending on the specific application scenario; for example, it may be subject teaching texts, exam training texts, or product user manuals. In an optional embodiment, the target document content may include chapters such as a case overview, evidence analysis, handling suggestions, and cautionary education. The case overview describes the basic facts of the case in detail; the evidence analysis organizes and analyzes the collected evidence materials; the handling suggestions provide recommendations for handling measures; and the cautionary education summarizes the lessons learned from the case, proposes preventative measures, and includes cautionary educational content to achieve the purpose of warning and education. Converting target documents into video format can improve information dissemination efficiency and enhance the effectiveness of education and training. It should be emphasized that the content of the target document can also be other content with video conversion requirements; in this embodiment, the content of the target document is not specifically limited.

[0040] For example, the target processing system can employ a chunked upload mechanism to support stable transmission of large files. When a user submits a document, the system first performs file type validation, accepting only files in the target format (e.g., ".docx" format). It then performs virus scanning and security checks to prevent malicious file uploads. Files that pass the security check are stored as target documents in a distributed file system, and a globally unique file identifier is generated. The system records file metadata, including filename, size, and upload time, providing a basis for subsequent processing.

[0041] In an optional embodiment, after a document passes security checks, the target processing system can preprocess the document and then store the preprocessed document as the target document in a distributed file system. For example, during preprocessing, the system can perform a series of standardized operations, including removing redundant spaces and line breaks, standardizing punctuation formats, and converting full-width and half-width characters. For table content in the document, the system processes it in a structured manner to maintain the integrity and readability of the table data. Non-text elements such as images and charts are extracted and stored separately for subsequent processing.

[0042] In an optional embodiment, the target processing system has an intelligent document quality assessment function, which can detect potential problems in documents, such as disordered formatting, unclear structure, and missing content, and generate corresponding assessment reports. For documents of poor quality, the system can provide optimization suggestions to guide users in making modifications and improvements.

[0043] The entire document preprocessing process can employ an asynchronous mechanism to avoid blocking user operations. Processing progress is provided to the user in real time, improving the smoothness of the user experience. The parsed results are saved in a structured format, laying the foundation for subsequent content analysis and video generation.

[0044] Step S202: Process the target document using the storyboard generation model to obtain storyboard information for multiple video storyboards.

[0045] In an optional embodiment, the storyboard generation model is a neural network model. For example, the storyboard generation model is a large language model.

[0046] For example, the system uses a storyboard generation model based on a large language model. This model is domain-specific trained and can understand the specific structure and content of the target document. The system inputs the document content into the model and guides it to produce a video overview and storyboard information for multiple video scenes using specially designed prompts. The storyboard information can include storyboard description text, storyboard dialogue, visual emphasis annotations, duration suggestions, etc. The model's output can be in structured or text format.

[0047] Step S203: For each video segment, generate segment image and segment audio based on the segment information of the video segment.

[0048] In an optional embodiment, the system can utilize a text-based image model to generate storyboard images based on descriptions in the video storyboard information (e.g., storyboard description text), with each storyboard image corresponding to a storyboard description.

[0049] The system can use a text-to-speech model to convert the dialogue text in the video storyboard into speech, generating storyboard audio.

[0050] Step S204: Generate the target video corresponding to the target document based on the segment images and segment audio of each of the multiple video segments.

[0051] In an optional embodiment, the target processing system first combines each segment's image with its corresponding audio to obtain the segment video corresponding to the video segment. Then, the segment videos of multiple video segments are sequentially spliced ​​together to obtain the target video.

[0052] In an optional embodiment, before stitching together the video segments of multiple video scenes, the target processing system can apply appropriate transition effects (such as fade-in / fade-out, wipe animation) to adjacent scenes to achieve a smooth transition between shots. After stitching, opening and closing credits, subtitles, background music, etc., are automatically added as needed to complete the video structure.

[0053] In an optional embodiment, after obtaining the target video, the target video is sent to the user, who is also the user who submitted the target document.

[0054] In this embodiment, by processing the target document using a storyboard generation model, storyboard information for multiple video storyboards is obtained, enabling the automatic determination of multiple storyboards required for the video to be generated based on the content of the target document. By generating storyboard images and audio based on the storyboard information of the video storyboards, automatic generation of each storyboard content in the video is achieved. Furthermore, by generating the target video corresponding to the target document based on the storyboard images and audio of each of the multiple video storyboards, automatic video generation based on storyboard content is realized, thereby avoiding reliance on manual video generation and effectively improving video generation efficiency.

[0055] Therefore, the method provided in this application achieves the goal of automatically generating videos from documents, realizes the technical effect of improving video generation efficiency, and solves the technical problem of low video generation efficiency caused by relying on manual generation of videos from documents in related technologies.

[0056] Optionally, in the video generation method provided in this application embodiment, the target document is processed by a storyboard generation model to obtain storyboard information for multiple video storyboards, including: inputting a first prompt word and the target document into the storyboard generation model, and processing the target document to obtain the document content structure of the target document through the storyboard generation model, wherein the first prompt word is used to guide the storyboard generation model to determine the document content structure based on the target document; inputting a second prompt word into the storyboard generation model, and processing the target document to obtain key content elements through the storyboard generation model, wherein the second prompt word is used to guide the storyboard generation model to determine key content elements based on the target document; inputting a third prompt word into the storyboard generation model, and processing the storyboard generation model to obtain storyboard information for multiple video storyboards through the storyboard generation model, wherein the third prompt word is used to guide the storyboard generation model to generate storyboard information for multiple video storyboards based on the target document, document content structure, and key content elements.

[0057] For example, after receiving the target document, the initial prompt and the target document are input into the storyboard generation model. The model then processes the data to obtain the document's content structure. For instance, if the target document is a case analysis document, the initial prompt could be, "Please analyze the content structure of the following document and identify the chapters for case overview, evidence analysis, handling suggestions, and cautionary education." The initial prompt and the target document are input together into the storyboard generation model. The model uses a text understanding mechanism to identify the main parts of the document and their relationships, and can output the analysis results of the document's content structure. The model can first perform basic text extraction, such as reading paragraph text, table content, and image information from the document. Next, it performs document structure analysis, identifying elements such as heading levels, paragraph relationships, and list structures. Then, it can automatically identify typical chapters such as case overview, evidence analysis, handling suggestions, and cautionary education based on chapter titles, thus facilitating the model's learning of the document content. This step aims to use the model to analyze the overall structure of the document, identify the logical relationships between different parts of the document, and lay the foundation for subsequent detailed processing.

[0058] Following document structure analysis, the system uses a second prompt to further guide the storyboard generation model. This second prompt could be something like, "Please extract key information points from the following document, including but not limited to evidence, event details, and the basis for handling the situation." The second prompt, along with the target document, is input into the storyboard generation model. The model performs a deep scan of the document content, identifying key content elements related to the prompt. The model can then mark these key content elements in the document's content structure analysis results, creating a detailed list of key content elements. The purpose of this step is to enable the model to identify key information (i.e., the content outline) in the document, such as evidence, and highlight them in the video. This helps ensure that the focus of the generated video aligns with the work's priorities, strengthening the communication of key information.

[0059] After identifying the document's content structure and key elements, the system uses a third prompt to guide the model in generating storyboard information. For example, the third prompt could be: "Based on the document's content structure and key elements summarized above, please generate storyboard information for multiple video scenes for the target document. The storyboard information should at least include storyboard description text and storyboard dialogue, and may also include fields such as visual emphasis annotations and duration suggestions." Here, visual emphasis annotations refer to close-up shots or long shots, and duration refers to the length of the video scene.

[0060] In an optional embodiment, to improve the processing performance of the storyboard generation model, the model can be fine-tuned for domain adaptation before generating storyboard information. For example, a large number of target documents, along with their corresponding document content structure, key content elements, and storyboard information, can be used to train the model's professional knowledge understanding ability. During the fine-tuning process, the focus is on optimizing the model's mastery of specific knowledge such as professional terminology, concepts, and case-handling procedures.

[0061] In an optional embodiment, the model employs a structured output mechanism. The model can generate results according to a preset JSON (JavaScript Object Notation) format, and the output content can include a video overview and multiple scene information. Each scene information includes at least a scene description text and scene dialogue, and may also include fields such as visual emphasis annotations and duration suggestions. This structured output enables subsequent processing stages to accurately understand and utilize the generated content.

[0062] In an optional embodiment, to improve generation quality, the system can set up multiple verification mechanisms. For example, firstly, the JSON format output by the model is syntax-validated to verify the correctness of the data structure. Next, a content integrity check is performed to verify that necessary fields are complete, i.e., whether storyboard description text and dialogue are missing, and whether the content is logically sound. Finally, the generated content can be checked to ensure it meets the work specifications.

[0063] In an optional embodiment, the system also supports manual intervention and adjustment of the storyboard content. Users can modify and improve the automatically generated storyboard information, and the system can record the user's adjustment preferences to gradually optimize the generated model.

[0064] It should be noted that the above method achieves a deep understanding and intelligent conversion of the target document content. The first step of document structure analysis provides a clear logical clue for video production, the second step of key element extraction improves the comprehensiveness and importance of the video content, and the third step of storyboard information generation transforms the document content into specific guidance for video production, thereby improving the accuracy of the generated storyboard information.

[0065] In another optional embodiment, the target processing system can directly input the fourth prompt word and the target document into the storyboard generation model. The storyboard generation model processes the data to obtain storyboard information for multiple video scenes. The fourth prompt word guides the storyboard generation model to generate storyboard information for multiple video scenes based on the target document. The storyboard generation model can be trained using a training sample set, where the training samples are sample documents, and the true labels of the training samples are the sample storyboard information for multiple sample video scenes corresponding to the sample documents.

[0066] Optionally, in the video generation method provided in this application embodiment, generating storyboard images and storyboard audio based on the storyboard information of the video storyboard includes: obtaining storyboard description text from the storyboard information and generating storyboard images based on the storyboard description text; obtaining storyboard dialogue from the storyboard information and generating storyboard audio based on the storyboard dialogue.

[0067] The storyboard information includes storyboard description text, which typically contains visual presentation suggestions for the video storyboard, such as scene setting, character actions, and visual focus. Based on these descriptions, the system generates corresponding storyboard images, providing the visual foundation for the video.

[0068] In one optional embodiment, the system uses a text-based image generation model to generate images based on the storyboard description text in the storyboard information. The text-based image model used here can be trained using deep learning and has the ability to generate high-quality images from text descriptions. For example, the system takes the storyboard description text as input to the model, and the model outputs an image reflecting the description content. In one optional embodiment, one storyboard description text corresponds to one storyboard image. In another optional embodiment, one storyboard description text may also correspond to multiple storyboard images.

[0069] For example, a text-based image model employing a stable diffusion architecture exhibits excellent performance in image quality and generation stability. Specifically optimized for the characteristics of the work, the model is fine-tuned using a large amount of professional image data, such as relevant promotional materials and educational materials, to better understand visual expression needs. Another example is the training of the text-based image model, where sample storyboard description text corresponding to sample documents is used as training samples, and sample storyboard images corresponding to the description text are used as ground truth labels to train the text-based image model.

[0070] In an optional embodiment, after obtaining the storyboard images, content security audits can be performed on them. For example, each generated storyboard image undergoes multiple filtering checks, including sensitive content detection, character image verification, and logo / identification verification. Images that may be ambiguous or do not meet specifications are automatically rejected and regenerated by the system to ensure the accuracy and appropriateness of the output content.

[0071] In an optional embodiment, after obtaining the storyboard images, post-processing can be performed on them. The image post-processing stage can include a series of optimization operations. For example, color correction to ensure the image tone meets relevant requirements; sharpening enhancement to improve the clarity of key elements; and size standardization to ensure all images conform to video production specifications. The processed storyboard images are stored in a standardized format and a detailed metadata index is established for easy retrieval and use in subsequent stages.

[0072] In an optional embodiment, the system can employ a combination of objective indicators and subjective evaluation to conduct a comprehensive quality assessment of the generated images. Objective indicators include image clarity, color accuracy, and compositional rationality; subjective evaluation focuses on factors such as the matching degree between the image and text, visual appeal, and professional compliance. Optionally, storyboard images that pass the quality assessment are passed to the next processing stage. If a storyboard image for a particular scene fails the quality assessment, a new storyboard image for that scene is generated.

[0073] In an optional embodiment, the system can employ a batch generation strategy to improve processing efficiency. This involves packaging the image generation tasks for multiple scenes into a single package, fully utilizing the parallel computing capabilities of the GPU (Graphics Processing Unit). Simultaneously, a caching mechanism is introduced to prioritize the use of already generated image resources for identical or similar descriptive content, reducing unnecessary computational overhead.

[0074] The storyboard information includes the storyboard dialogue. This text is the narration or dialogue corresponding to the storyboard image.

[0075] Optionally, the system can convert text information into audio using text-to-speech technology. For example, the dialogue from a storyboard can be input into a speech generation model to obtain the audio. This speech generation model can be a neural network model. In an optional embodiment, during the generation of the audio, the system can adjust parameters such as speech rate, tone, and volume to adapt to the emotional needs and context of different video storyboards.

[0076] In an optional embodiment, to improve the accuracy of the speech generation model, the model is trained based on sample data corresponding to the scene. For example, if the current scene is used to generate video for a target document, the training samples for the speech generation model can be sample documents, and the real labels for the training samples can be the audio corresponding to the sample document. This improves the model's accuracy in pronouncing relevant technical terms.

[0077] In an optional embodiment, the system can achieve personalized voice output through voice parameter adjustment. The system can provide a variety of voice options, including male and female voices of different age groups, to meet the needs of different scenarios. Speech rate adjustment supports precise control; important content can be spoken at a slower pace to enhance emphasis. The audio generation process can employ streaming processing technology, supporting efficient conversion of large-scale text. Long texts are automatically segmented into appropriate paragraphs, generated separately, and then seamlessly spliced ​​together to ensure overall fluency.

[0078] In an optional embodiment, after generating the storyboard audio, the system can perform a comprehensive inspection of the generated audio. Objective indicators include technical parameters such as signal-to-noise ratio, frequency response, and distortion. Optionally, storyboard audio that passes the quality assessment is passed to the next processing stage. If the storyboard audio for a particular scene fails the quality assessment, the storyboard audio for that scene is regenerated. The system can establish an audio sample library and continuously optimize the generated quality through comparative analysis.

[0079] It should be noted that by analyzing the descriptive text and dialogue in the storyboard information, the system can accurately understand the intent and requirements of each part of the video, generating images and audio that highly match the content, thereby improving the accuracy of the generated images and audio. Furthermore, by utilizing the model's automated processing capabilities, manual intervention is greatly reduced, accelerating video production and improving video generation efficiency.

[0080] Optionally, in the video generation method provided in this application embodiment, generating a storyboard image based on the storyboard description text includes: generating an initial storyboard image based on the storyboard description text; determining the storyboard order of the video storyboard to which the storyboard description text belongs among multiple video storyboards; determining the image animation effect matching the storyboard order; and determining the storyboard image based on the initial storyboard image and the image animation effect.

[0081] In an optional embodiment, the system calls a text-based image model, takes the storyboard description text as input to the model, and the model generates a static image (i.e., the initial storyboard image) based on the text description.

[0082] After obtaining the initial storyboard image, the storyboard order of the video storyboard to which the storyboard description text belongs is determined among multiple video storyboards. In an optional embodiment, when the storyboard generation model outputs storyboard information for multiple video storyboards, it generates the information according to the order in which the multiple video storyboards are arranged. Therefore, the target processing system can determine the order of the storyboard information in the output of the storyboard generation model as the aforementioned storyboard order. In another optional embodiment, the storyboard information may include the storyboard order of the video storyboard to which the storyboard information belongs.

[0083] The target processing system can pre-set multiple matching relationships between storyboard positions and image animation effects. Based on these matching relationships, the system can determine the image animation effect matching the video storyboard to which the storyboard description text belongs. For example, the system has a built-in image animation effect library, including basic translation, scaling, and rotation, as well as complex transition effects (such as fade-in / fade-out, wipe, fly-in, etc.). According to the storyboard position determined in the previous step, the system selects the matching image animation effect from the image animation effect library. The design of these image animation effects follows the principles of visual perception. For example, translation animation simulates the effect of a camera panning, creating a sense of space; scaling animation achieves a push-pull shot effect, highlighting important content; and rotation animation adds dynamic changes, avoiding visual monotony.

[0084] In an optional embodiment, the target processing system can directly determine the storyboard image based on the initial storyboard image and the image animation effect. For example, after determining the image animation effect to be used, the determined image animation effect is applied to the aforementioned generated initial storyboard image to generate the final storyboard image, which includes dynamic effects, i.e., a dynamic image. For example, the system can use a multimedia processing framework to combine the image animation effect with the initial storyboard image to generate a final storyboard image with animation effects. This process may involve image preprocessing (such as resizing and color correction), parameter control of the animation effect (such as speed and direction), and post-processing of the effect (such as smooth transitions and noise reduction enhancement).

[0085] In an optional embodiment, the target processing system can determine a first storyboard image based on an initial storyboard image and image animation effects, and then add animation effects to the first storyboard image according to its image content to obtain a new storyboard image. For example, the system analyzes the image content of the first storyboard image using computer vision technology, identifies key areas and important elements, and adds targeted animation effects accordingly. For example, for a first storyboard image containing a person, the system automatically identifies the face position and adds visual guidance animation (e.g., adding a dynamic outline effect); for an image containing text, it adds animation effects for highlighting the text (e.g., a dynamic arrow effect).

[0086] In an optional embodiment, the system may preprocess the initial storyboard images before adding image animation effects. For example, it may focus on optimizing the sharpness and recognizability of key elements in the images and ensuring that all generated initial storyboard images maintain consistency in terms of tone, brightness, and composition.

[0087] In an optional embodiment, the system can employ various techniques to improve processing efficiency. For example, during image preprocessing, image size can be optimized based on the target resolution to reduce unnecessary computational overhead. The animation generation stage supports parallel processing, fully utilizing the computing power of multi-core CPUs (Central Processing Units). A caching mechanism stores intermediate processing results to avoid redundant calculations.

[0088] In an optional embodiment, the system can detect the smoothness of the animation and automatically optimize or regenerate substandard storyboard images. It also provides a preview function, allowing users to view the animation effects and offer adjustments before final compositing.

[0089] It should be noted that the above method enables the effective generation of dynamic storyboard images, enhances the expressiveness of the images, improves the overall viewing experience and information delivery efficiency of the video, and increases the reliability of video generation.

[0090] Optionally, in the video generation method provided in this application embodiment, generating a target video corresponding to a target document based on the segment images and segment audio of each of the multiple video segments includes: for each video segment, performing a synthesis process on the segment image and segment audio of the video segment to obtain a segment video; and generating a target video based on the segment videos of each of the multiple video segments.

[0091] Optionally, the system can synthesize the image and audio content of a single video storyboard to form a video clip containing both visual and auditory information, i.e., a storyboard video. For example, the video duration of the storyboard video can be controlled to the duration indicated in the storyboard information. Alternatively, the video duration can be controlled to the maximum of the following durations: the audio duration of the storyboard audio and the duration of a single animation of the storyboard image. For instance, the target processing system can perform the above synthesis process using a multimedia processing framework to obtain the storyboard video.

[0092] In an optional embodiment, after the compositing of all video segments is completed, the system splices these video segments sequentially to obtain the target video.

[0093] In an optional embodiment, after completing the compositing process of all video storyboards, the system can add animated transition effects between adjacent storyboard videos, and then determine the target video based on the individual storyboard videos of the multiple video storyboards and the animated transition effects.

[0094] In an optional embodiment, the target processing system can add preset elements such as intro videos, outro videos, and background music during the process of generating the target video based on the individual storyboard videos of multiple video scenes, ultimately generating a complete target video. For example, the intro and outro videos can include elements such as project logos and copyright information. The background music system can automatically select music based on the mood of the content, enhancing the video's emotional impact.

[0095] It should be noted that the above method enables the accurate generation of the target video.

[0096] Optionally, in the video generation method provided in this application embodiment, generating a target video based on the individual video segments of multiple video segments includes: determining multiple video segment pairs based on the multiple video segments, wherein each video segment pair consists of two adjacent video segments from the multiple video segments; for each video segment pair, determining the degree of content association between the two video segments in the video segment pair; determining the animation transition effect between the two video segments in the video segment pair based on the degree of content association; and determining the target video based on the individual video segments of the multiple video segments and the animation transition effect.

[0097] Optionally, the system can select two adjacent scenes one by one from multiple video scenes to form scene pairs. For example, if the video contains three scenes A, B, and C, then two scene pairs, AB and BC, will be formed. Or, if the video contains four scenes A, B, C, and D, then three scene pairs, AB, BC, and CD, will be formed.

[0098] For each video shot pair, the system can employ a deep learning model, such as a pre-trained language model (e.g., BERT) or a specially trained large language model, to compare and analyze the descriptive text of the two shots within the pair. This involves calculating their semantic similarity, which is then used to determine the degree of content association between the video shots. This calculation may be based on cosine similarity of word vectors, edit distance, or more complex semantic analysis algorithms.

[0099] The target processing system can preset multiple correlation value ranges and their relationships with animation transition effects. The system can determine the correlation range to which the aforementioned content correlation belongs, and then determine the animation transition effect associated with that correlation range as the animation transition effect between two video scenes in a video scene pair. For example, when the analyzed correlation is low, a more explicit transition may be selected; while when the correlation is high, the system tends to choose a softer transition, such as fade-in or fade-out, to maintain narrative continuity.

[0100] In an optional embodiment, during the synthesis of the target video, the system can detect technical indicators such as audio and video quality and audio level in real time. The final product undergoes automated quality inspection, including format verification, content integrity checks, and playback compatibility testing. The target video that passes the inspection is then provided back to the user who submitted the target document.

[0101] In an optional embodiment, for performance optimization, the system can employ an intelligent encoding strategy to optimize file size while maintaining image quality. It supports multiple resolution outputs to adapt to different usage scenarios. Parallel processing technology fully utilizes system resources to improve video compositing efficiency. An incremental generation mechanism allows for partial modifications to the synthesized video, avoiding complete re-rendering.

[0102] It should be noted that by identifying the content correlation between adjacent shots, the system can accurately select and adjust transition effects, thereby not only ensuring a smooth visual transition, but also enhancing the logic and expressiveness of the content, and improving the effect and reliability of the generated target video.

[0103] Optionally, in the video generation method provided in this application embodiment, generating storyboard audio based on storyboard dialogue includes: generating initial storyboard audio based on storyboard dialogue; performing audio post-processing on the initial storyboard audio to obtain storyboard audio, wherein the audio post-processing includes at least one of the following: noise reduction processing, equalizer processing, and compression processing.

[0104] In an optional embodiment, the system can convert text information into audio using text-to-speech technology. For example, the dialogue from a storyboard can be input into a speech generation model to obtain initial audio for the storyboard.

[0105] After generating the initial storyboard audio, the initial storyboard audio is optimized using at least one of the following methods to improve sound quality and listening experience: noise reduction, equalizer processing, and compression.

[0106] Optionally, during the noise reduction process, the system can use noise suppression algorithms (such as deep learning-based noise filtering) to eliminate background noise, hissing, etc. in the audio, thereby improving the purity of the speech.

[0107] During equalizer processing, the system can optimize the clarity and fullness of speech by adjusting the frequency response of the audio (high and low frequency balance), making key information stand out and enhancing the content delivery effect.

[0108] During the compression process, the system can apply dynamic range compression technology to control volume fluctuations in audio, avoid discomfort caused by sudden volume changes, and optimize audio file size to facilitate video synthesis and subsequent transmission.

[0109] After performing the above audio post-processing on the initial storyboard audio, the storyboard audio is obtained.

[0110] It should be noted that by combining audio post-processing techniques to obtain the storyboard audio, the audio quality can be effectively improved, thereby increasing the accuracy of the generated target video.

[0111] In an optional embodiment, the overall architecture of the target processing system can include four main layers: a presentation layer, a business logic layer, a data persistence layer, and an infrastructure layer. The presentation layer is responsible for interacting with the user and providing a user-friendly interface; the business logic layer encapsulates the core processing flow and implements business rules and algorithms; the data persistence layer manages the storage and access of system data; and the infrastructure layer provides support for computing, network, and storage resources.

[0112] In terms of technology framework selection, the system can adopt a front-end and back-end separation development model. The front-end builds a responsive user interface, using relevant component libraries to maintain a consistent interface style. The back-end uses relevant frameworks to provide API (Application Programming Interface) services, leveraging their auto-configuration and starter dependency features to quickly build production-grade applications. The workflow engine uses an application development platform that provides visual AI (Artificial Intelligence) workflow orchestration capabilities, supports custom code nodes and rich AI model integration. The database system stores structured data.

[0113] The system deployment architecture can utilize containerization technology. This architecture offers excellent horizontal scalability, allowing for dynamic resource allocation adjustments based on business load. For security, the system employs a multi-layered protection strategy, including network isolation, authentication, access control, and data encryption, to enhance data confidentiality and integrity.

[0114] In an optional embodiment, Figure 3 This is a schematic diagram of the operation of the target processing system provided in the embodiments of this application, such as... Figure 3 As shown, the target processing system can be divided into three main parts: (1) Application services: The front end provides a user interface, receives user requests, calls the application development platform API, manages task status and returns results.

[0115] (2) Application Development Platform (AI Workflow Engine): The core processing unit, which includes independent workflows.

[0116] (3) Model Inference Server Cluster: Independently deployed GPU servers provide model inference services for the application development platform workflow through APIs, such as providing text-to-image models, providing storyboard generation models, and providing speech generation models.

[0117] In an optional embodiment, the application service may perform the following functions: (1) Import / Export: You can add, delete, and modify files.

[0118] (2) Upload format prompts: remind users which file formats can be uploaded.

[0119] (3) Queue list: Displays the queuing progress of all users in the video production process and the remaining time for the current task.

[0120] (4) Download page: The export is complete and displays "Completed", shows the size of the video file, and provides a download button for the video file.

[0121] (5) Supported file format for import: docx. And the text must conform to the five-paragraph structure requirement.

[0122] (6) Receive task: Save the video uploaded by the user and record the task type (scaling, merging, etc.).

[0123] (7) Task queuing: Put new tasks into the queue in order and give each task a unique number.

[0124] (8) Processing tasks: Call the tools in sequence to process the video.

[0125] (9) Return status: tells the front end the current queue position, processing progress, remaining time, etc.

[0126] (10) Supported export video (including audio) formats: Default mp4 format (best compatibility). Export video resolution: Default 640. 480 (wide compatibility / low hardware requirements).

[0127] In an optional embodiment, Figure 4 This is a schematic diagram of the application development platform provided in the embodiments of this application, such as... Figure 4 As shown, the application development platform can execute the following processes: Input: The target document uploaded by the user.

[0128] 1. Document parsing: Extracting text from a document.

[0129] 2. Summary and Storyboard: Generate a storyboard model with the prompt (example): "Please summarize the following document content into an overview and divide it into several concise shot descriptions and dialogue." Output structured text (JSON format) containing the overview and storyboard information.

[0130] 3. Voice: Convert the dialogue text into audio files using a voice generation model.

[0131] 4. Still Image Generation: Using the text-based image model, still images are generated for each shot description.

[0132] 5. Image Animation Processing: Add basic dynamic effects to static images.

[0133] 6. Output: Synthesize the sub-videos of each shot and stitch all the synthesized sub-videos together in sequence to form the final target video.

[0134] Therefore, the method provided in this application achieves the goal of automatically generating videos from documents, realizes the technical effect of improving video generation efficiency, and solves the technical problem of low video generation efficiency caused by relying on manual generation of videos from documents in related technologies.

[0135] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0136] Example 2

[0137] This application also provides a video generation apparatus. It should be noted that the video generation apparatus of this application can be used to execute the video generation method provided in this application. The video generation apparatus provided in this application will be described below.

[0138] According to an embodiment of this application, an apparatus for implementing the above-described video generation method is also provided, such as... Figure 5 As shown, the device includes: The acquisition module 501 is used to acquire a target document, wherein the target document includes information describing the target video to be generated; The processing module 502 is used to process the target document through the storyboard generation model to obtain storyboard information for multiple video storyboards; The first generation module 503 is used to generate a storyboard image and a storyboard audio for each video storyboard based on the storyboard information of the video storyboard. The second generation module 504 is used to generate the target video corresponding to the target document based on the storyboard images and audio of each of the multiple video storyboards.

[0139] In this embodiment, by processing the target document using a storyboard generation model, storyboard information for multiple video storyboards is obtained, enabling the automatic determination of multiple storyboards required for the video to be generated based on the content of the target document. By generating storyboard images and audio based on the storyboard information of the video storyboards, automatic generation of each storyboard content in the video is achieved. Furthermore, by generating the target video corresponding to the target document based on the storyboard images and audio of each of the multiple video storyboards, automatic video generation based on storyboard content is realized, thereby avoiding reliance on manual video generation and effectively improving video generation efficiency.

[0140] Therefore, the method provided in this application achieves the goal of automatically generating videos from documents, realizes the technical effect of improving video generation efficiency, and solves the technical problem of low video generation efficiency caused by relying on manual generation of videos from documents in related technologies.

[0141] Optionally, in the video generation apparatus provided in this application embodiment, the processing module further includes: a first processing submodule, used to input a first prompt word and a target document into a storyboard generation model, and obtain the document content structure of the target document through the storyboard generation model, wherein the first prompt word is used to guide the storyboard generation model to determine the document content structure based on the target document; a second processing submodule, used to input a second prompt word into a storyboard generation model, and obtain key content elements in the target document through the storyboard generation model, wherein the second prompt word is used to guide the storyboard generation model to determine key content elements based on the target document; and a third processing submodule, used to input a third prompt word into a storyboard generation model, and obtain storyboard information of multiple video storyboards through the storyboard generation model, wherein the third prompt word is used to guide the storyboard generation model to generate storyboard information of multiple video storyboards based on the target document, the document content structure, and the key content elements.

[0142] Optionally, in the video generation apparatus provided in this application embodiment, the first generation module further includes: a first generation submodule, used to obtain storyboard description text from storyboard information and generate storyboard image based on the storyboard description text; and a second generation submodule, used to obtain storyboard dialogue from storyboard information and generate storyboard audio based on the storyboard dialogue.

[0143] Optionally, in the video generation apparatus provided in this application embodiment, the first generation submodule further includes: a first generation unit, configured to generate an initial storyboard image based on the storyboard description text; a first determination unit, configured to determine the storyboard sequence position of the video storyboard to which the storyboard description text belongs among multiple video storyboards; a second determination unit, configured to determine the image animation effect matching the storyboard sequence position; and a third determination unit, configured to determine the storyboard image based on the initial storyboard image and the image animation effect.

[0144] Optionally, in the video generation apparatus provided in this application embodiment, the second generation module further includes: a fourth processing submodule, used to synthesize the segment image and segment audio of each video segment to obtain a segment video; and a third generation submodule, used to generate a target video based on the segment videos of each of the multiple video segments.

[0145] Optionally, in the video generation apparatus provided in this application embodiment, the third generation submodule further includes: a fourth determining unit, configured to determine multiple video segment pairs based on multiple video segments, wherein a video segment pair consists of two adjacent video segments among the multiple video segments; a fifth determining unit, configured to determine the content association degree between the two video segments in each video segment pair; a sixth determining unit, configured to determine the animation transition effect between the two video segments in the video segment pair based on the content association degree; and a seventh determining unit, configured to determine the target video based on the individual video segments of the multiple video segments and the animation transition effect.

[0146] Optionally, in the video generation apparatus provided in this application embodiment, the second generation submodule further includes: a second generation unit, used to generate initial storyboard audio based on storyboard dialogue; and a processing unit, used to perform audio post-processing on the initial storyboard audio to obtain storyboard audio, wherein the audio post-processing includes at least one of the following: noise reduction processing, equalizer processing, and compression processing.

[0147] It should be noted that the acquisition module 501, processing module 502, first generation module 503, and second generation module 504 mentioned above correspond to steps S201 to S204 in Embodiment 1. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.

[0148] Example 3

[0149] Embodiments of this application may provide an electronic device. Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) Processor 1002, memory 1004, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0150] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0151] The processor can access information and applications stored in memory via a transmission device to perform the following steps: acquiring a target document, wherein the target document includes information describing the target video to be generated; processing the target document using a storyboard generation model to obtain storyboard information for multiple video storyboards; for each video storyboard, generating storyboard images and storyboard audio based on the storyboard information of the video storyboard; and generating the target video corresponding to the target document based on the storyboard images and storyboard audio of each of the multiple video storyboards.

[0152] The processor can also invoke information and applications stored in the memory via a transmission device to perform the following steps: inputting a first prompt word and a target document into the storyboard generation model, and processing the storyboard generation model to obtain the document content structure of the target document, wherein the first prompt word is used to guide the storyboard generation model to determine the document content structure based on the target document; inputting a second prompt word into the storyboard generation model, and processing the storyboard generation model to obtain key content elements in the target document, wherein the second prompt word is used to guide the storyboard generation model to determine key content elements based on the target document; inputting a third prompt word into the storyboard generation model, and processing the storyboard generation model to obtain storyboard information for multiple video storyboards, wherein the third prompt word is used to guide the storyboard generation model to generate storyboard information for multiple video storyboards based on the target document, document content structure, and key content elements.

[0153] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: obtain the storyboard description text from the storyboard information and generate the storyboard image based on the storyboard description text; obtain the storyboard dialogue from the storyboard information and generate the storyboard audio based on the storyboard dialogue.

[0154] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: generating an initial storyboard image based on the storyboard description text; determining the storyboard sequence of the video storyboard to which the storyboard description text belongs among multiple video storyboards; determining the image animation effect that matches the storyboard sequence; and determining the storyboard image based on the initial storyboard image and the image animation effect.

[0155] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: for each video segment, synthesize the segment image and segment audio of the video segment to obtain the segment video; generate the target video based on the segment videos of multiple video segments.

[0156] The processor can also call information and applications stored in the memory via a transmission device to perform the following steps: determining multiple video scene pairs based on multiple video scenes, wherein a video scene pair consists of two adjacent video scenes from the multiple video scenes; for each video scene pair, determining the degree of content association between the two video scenes in the video scene pair; determining the animation transition effect between the two video scenes in the video scene pair based on the degree of content association; and determining the target video based on the individual video scenes of the multiple video scenes and the animation transition effect.

[0157] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: generate initial storyboard audio based on the storyboard dialogue; perform audio post-processing on the initial storyboard audio to obtain storyboard audio, wherein the audio post-processing includes at least one of the following: noise reduction processing, equalizer processing, and compression processing.

[0158] In this embodiment, by processing the target document using a storyboard generation model, storyboard information for multiple video storyboards is obtained, enabling the automatic determination of multiple storyboards required for the video to be generated based on the content of the target document. By generating storyboard images and audio based on the storyboard information of the video storyboards, automatic generation of each storyboard content in the video is achieved. Furthermore, by generating the target video corresponding to the target document based on the storyboard images and audio of each of the multiple video storyboards, automatic video generation based on storyboard content is realized, thereby avoiding reliance on manual video generation and effectively improving video generation efficiency.

[0159] Therefore, the method provided in this application achieves the goal of automatically generating videos from documents, realizes the technical effect of improving video generation efficiency, and solves the technical problem of low video generation efficiency caused by relying on manual generation of videos from documents in related technologies.

[0160] Those skilled in the art will understand that Figure 6The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.

[0161] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0162] Example 4

[0163] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the video generation method provided in Embodiment 1.

[0164] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0165] This application also provides a computer program product that, when executed on a data processing device, is adapted to perform video generation method steps.

[0166] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0167] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0168] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0169] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0170] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0171] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0172] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A video generation method, characterized in that, include: Obtain a target document, wherein the target document includes information describing the target video to be generated; The target document is processed by a storyboard generation model to obtain storyboard information for multiple video scenes; For each video segment, generate a segment image and segment audio based on the segment information of the video segment; The target video corresponding to the target document is generated based on the segment images and segment audio of each of the multiple video segments.

2. The method according to claim 1, characterized in that, The target document is processed using a storyboard generation model to obtain storyboard information for multiple video scenes, including: The first prompt word and the target document are input into the storyboard generation model, and the storyboard generation model processes the target document to obtain the document content structure. The first prompt word is used to guide the storyboard generation model to determine the document content structure based on the target document. The second prompt word is input into the storyboard generation model, and the storyboard generation model processes the key content elements in the target document to obtain the key content elements. The second prompt word is used to guide the storyboard generation model to determine the key content elements based on the target document. The third prompt word is input into the storyboard generation model, and the storyboard generation model processes the storyboard to obtain the storyboard information of the multiple video storyboards. The third prompt word is used to guide the storyboard generation model to generate the storyboard information of the multiple video storyboards based on the target document, the document content structure, and the key content elements.

3. The method according to claim 1, characterized in that, Generate storyboard images and storyboard audio based on the storyboard information of the video storyboard, including: Obtain the storyboard description text from the storyboard information, and generate the storyboard image based on the storyboard description text; The storyboard dialogue is obtained from the storyboard information, and the storyboard audio is generated based on the storyboard dialogue.

4. The method according to claim 3, characterized in that, Generating the storyboard image based on the storyboard description text includes: Generate an initial storyboard image based on the storyboard description text; Determine the sequence number of the video segment to which the segment description text belongs among the multiple video segments; Determine the image animation effect matching the storyboard sequence; The storyboard image is determined based on the initial storyboard image and the image animation effect.

5. The method according to claim 1, characterized in that, Generate the target video corresponding to the target document based on the segment images and audio of each of the multiple video segments, including: For each video segment, the segment image and segment audio of the video segment are synthesized to obtain the segment video; The target video is generated based on the individual storyboard videos of the multiple video segments.

6. The method according to claim 5, characterized in that, The target video is generated based on the individual storyboard videos of the multiple video storyboards, including: Multiple video shot pairs are determined based on the multiple video shots, wherein each video shot pair consists of two adjacent video shots from the multiple video shots; For each video shot pair, determine the degree of content association between the two video shots in the video shot pair; Based on the degree of content correlation, determine the animation transition effect between the two video shots in the video shot pair; The target video is determined based on the individual storyboard videos of the multiple video segments and the animation transition effects.

7. The method according to claim 3, characterized in that, Generating the storyboard audio based on the storyboard dialogue includes: Generate initial storyboard audio based on the storyboard dialogue; The initial storyboard audio is subjected to audio post-processing to obtain the storyboard audio, wherein the audio post-processing includes at least one of the following: noise reduction processing, equalizer processing, and compression processing.

8. A video generation apparatus, characterized in that, include: An acquisition module is used to acquire a target document, wherein the target document includes information describing the target video to be generated; The processing module is used to process the target document through the storyboard generation model to obtain storyboard information for multiple video storyboards; The first generation module is used to generate a storyboard image and storyboard audio for each video storyboard based on the storyboard information of the video storyboard. The second generation module is used to generate the target video corresponding to the target document based on the storyboard images and storyboard audio of each of the multiple video storyboards.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the video generation method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the video generation method according to any one of claims 1 to 7.