Video generation method and device, equipment and storage medium

By using a large language model to split the script and generate character image images, combined with storyboards and audio, a video with a coherent plot is automatically generated. This solves the problem that existing technologies cannot generate complete logical videos, and achieves low-cost and efficient personalized video generation.

CN120640094APending Publication Date: 2025-09-12BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510635086.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing AI-generated video models can only generate dynamic images of a single shot, cannot achieve video interpretation with complete logic and plot, and cannot meet the low-cost, personalized video generation needs of ordinary users.

Method used

The target script is split into shots through a large language model, and the character's characteristic information is extracted to generate a character image. The storyboard script and audio are combined to generate a storyboard video, and finally combined into a target video with a coherent plot, replacing the shooting, editing and post-synthesis steps in the traditional film and television production process.

Benefits of technology

It achieves low-cost and efficient generation of videos with complete story logic, saving production costs and time, ensuring the coherence and consistency of the video, and meeting the personalized needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640094A_ABST
    Figure CN120640094A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a video generation method and device, equipment and a storage medium. The method comprises the following steps: splitting a target script through a large language model to obtain a split script, and extracting feature information of each role from the target script; generating a corresponding role image graph based on the feature information through a generation tool; respectively inputting the role image graph and each corresponding split script into a third language model to obtain a split graph corresponding to each split script; matching a corresponding target audio for each split script; combining each split script with the corresponding split image and the target audio to generate a corresponding split video; and combining all the split videos to obtain a target video. By adopting the method, matched videos with correct logic and plot coherence can be automatically and efficiently generated according to story characters, the individual requirements of users are met, and a large amount of labor cost, money cost and time cost can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for generating a video. Background Art

[0002] With the rapid development of the internet, interpreting literary works through videos has become a popular online communication method. For example, novels are adapted into films and television programs through live-action filming and broadcast on video websites, satisfying users' entertainment needs. The process from script to finished product requires casting, storyboarding, filming, and post-production, which incurs high production costs and a long production cycle. This method is typically only used by film and television companies and is not suitable for ordinary users. Ordinary users prefer low-cost, personalized video production needs. For example, users want to present their own stories in video form and share them on social media.

[0003] With the rapid development of AI (Artificial Intelligence) technology, AI models capable of generating videos have emerged. However, these models can only generate single-shot dynamic images and are unable to achieve video interpretation with complete logic and plot. Therefore, it is necessary to find a low-cost video generation method that can fully present the story logic to meet the personalized needs of users. Summary of the Invention

[0004] In view of this, the present application aims to propose a method, apparatus, device and storage medium for generating videos, so as to achieve low-cost and efficient generation of videos with complete story logic.

[0005] To achieve the above objectives, the technical solutions of this application are as follows:

[0006] A first aspect of an embodiment of the present application provides a method for generating a video, the method comprising:

[0007] Input the target script into the first language model to obtain multiple storyboards;

[0008] Inputting the target script into a second language model to extract feature information of each character from the target script through the second language model;

[0009] Inputting the characteristic information into a generation tool to generate a corresponding character image;

[0010] Inputting the character image and each corresponding storyboard into a third language model to obtain a storyboard corresponding to each storyboard;

[0011] Match the target audio to each storyboard;

[0012] Generate the corresponding storyboard video using the image-to-video tool based on the storyboard images and target audio for each storyboard script;

[0013] Combine all storyboard videos to get the target video.

[0014] According to a second aspect of an embodiment of the present application, a device for generating a video is provided, for implementing the steps of the method provided in the first aspect of the embodiment of the present application, the device comprising:

[0015] The splitting module is used to input the target script into the first language model to obtain multiple storyboards;

[0016] a character construction module, configured to input the target script into a second language model to extract characteristic information of each character from the target script through the second language model; and input the characteristic information into a generation tool to generate a corresponding character image;

[0017] A storyboard generation module is configured to input the character image and each corresponding storyboard script into a third language model to obtain a storyboard image corresponding to each storyboard script;

[0018] Audio generation module, used to match the corresponding target audio for each storyboard;

[0019] A video generation module is used to generate a corresponding storyboard video using a picture-to-video tool based on the storyboard image and target audio corresponding to each storyboard script;

[0020] The fusion module is used to combine all the storyboard videos to obtain the target video.

[0021] According to a third aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the first aspect of the embodiment of the present application are implemented.

[0022] According to the fourth aspect of the embodiments of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the computer program is executed by the processor, the steps in the method provided in the first aspect of the embodiments of the present application are implemented.

[0023] The method for generating a video provided by the present application is used to extract the characteristic information of each character from the target script, and generate a corresponding character image based on the characteristic information of each character. The target script is split by a large language model to obtain multiple storyboards, and corresponding storyboards are generated according to the storyboards. In the process of generating storyboards, the character image images obtained in advance are used as a reference to ensure that the consistency of the characters is maintained between the storyboards generated based on each storyboard script. Based on each storyboard, a corresponding target audio is generated, and based on the target audio and storyboard of each storyboard, a storyboard video corresponding to the storyboard is generated. Finally, all the storyboard videos are combined to generate a target video with a coherent plot.

[0024] The method for generating a video provided by the present application uses a large language model to realize script shot splitting and character information extraction, determines the character image through a generation tool and generates a storyboard in combination with the large language model, replacing the manual storyboard and storyboard drawing process in the early stage of shooting in the traditional way. In addition, the target audio is matched based on the storyboard script, and the production of the storyboard video is realized through the image-generated video tool, and then the storyboard video is combined to obtain the target video, replacing the process of shooting materials, editing, dubbing and post-synthesis in the traditional film and television production process, greatly saving the video production cost and improving the video generation efficiency. In addition, in the process of generating the storyboard video through the storyboard, the pre-obtained character image image is used as a benchmark to ensure that the image of the same character in each generated storyboard video remains consistent, thereby ensuring that the target video generated in the end has coherence on the screen and conforms to normal logic. Compared with the traditional film and television production process, the present application does not require manual participation, and can realize the automatic and efficient generation of matching videos with correct logic and plot coherence based on the story text, meet the personalized needs of users, and greatly save labor costs, money costs and time costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 This is a flowchart of a method for generating a video proposed in one embodiment of the present application;

[0027] Figure 2 This is one of the flow charts for generating a storyboard video in one embodiment of the present application;

[0028] Figure 3 This is the second flow chart of generating a storyboard video in one embodiment of the present application;

[0029] Figure 4 is a schematic diagram of a device for generating a video proposed in one embodiment of the present application;

[0030] Figure 5 FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] It should be understood that references throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout this specification do not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0033] In the various embodiments of the present application, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with certain aspects as detailed herein.

[0035] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.

[0036] With the rapid development of large AI models, various generative models are constantly emerging. Related technologies can generate dynamic graphics from text or images. However, due to the limitations of generative models, both image-based and text-based videos can only produce short, single-shot dynamic videos. They are unable to generate complete videos with coherent plots and correct logic from textual content with storylines.

[0037] The solution of this application uses automated operations to replace manual labor to realize the production process of film and television works from script to finished product, thereby reducing the difficulty and cost of video production and meeting the personalized needs of users.

[0038] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0039] Figure 1 This is a flow chart of a method for generating a video according to an embodiment of the present application. Figure 1 As shown, the method includes:

[0040] S1: Input the target script into the first language model to obtain multiple storyboards;

[0041] S2: inputting the target script into a second language model to extract feature information of each character from the target script through the second language model;

[0042] S3: inputting the feature information into a generation tool to generate a corresponding character image;

[0043] S4: Inputting the character image and each corresponding storyboard into a third language model to obtain a storyboard corresponding to each storyboard;

[0044] S5: Match the corresponding target audio for each storyboard;

[0045] S6: Generate the corresponding storyboard video using the image-to-video tool based on the storyboard images and target audio corresponding to each storyboard script;

[0046] S7: Combine all the storyboard videos to obtain the target video.

[0047] In this embodiment, a large language model is first used to segment the target script into shots, generating multiple storyboards. Feature information for all characters is then extracted from the target script, and a character image corresponding to each character is generated based on that feature information. A text-to-image operation is then performed, generating corresponding storyboards based on the storyboards using the generative model.

[0048] In addition to splitting the shot script, this embodiment of the application also uses a large language model to extract feature information for all characters from the target script. A generative model is then used to generate corresponding character image diagrams based on the feature information for each character, completing the character image design process. In practical applications, generative models that can be used include the image generation tool Midjourney, the Wensheng graph model Flux, and tools such as Ketu.

[0049] Specifically, a character's characteristic information may include gender, age, occupation, physical description, and clothing description. Optionally, the character's characteristic information may also include the name or reference image of a specific public figure (e.g., a well-known actor). For example, the target script states, "Zhang San looks very similar to the well-known actor Zhao Liu." Based on this characteristic information, the large language model obtains an image of Zhao Liu and provides it to the generative model along with Zhang San's other characteristic information. The generative model ultimately generates a character image of Zhang San that resembles Zhao Liu.

[0050] After obtaining the character image diagram, the character image diagram is used to control the generation process of the storyboard diagram. Considering that the target script in this application is a logical and story-telling content, the same character image in each storyboard script should be consistent to ensure that the generated target video has correct logic and coherence. Since it is difficult to ensure consistency when directly generating a video using text, this application adopts a two-stage video generation solution. The first stage generates pictures through text, and the second stage generates videos through pictures.

[0051] In the first stage, a character image is introduced as a baseline reference for the model generation, and then storyboards are generated based on the character image. Because each storyboard is generated based on a predetermined character image, the same character in each storyboard has a consistent image, improving the controllability of the storyboard images and avoiding character image confusion, plot disconnects, and even logical errors. Optionally, when generating each storyboard based on the character image, the IP-Adapter algorithm, InstantID algorithm, or the image generation fine-tuning model LoRA is used to perform the storyboard generation operation.

[0052] In this application, the corresponding target audio is also generated based on the storyboard script. The storyboard image and the target audio are input into the image-generated video model to obtain a storyboard video containing sound. Furthermore, all the storyboard videos are synthesized according to the order of the storyboard scripts to obtain the final target video.

[0053] This application realizes the automatic generation of videos that match the narrative text, saving a large number of manual operation steps, reducing production costs, and improving video generation efficiency. In response to the personalized needs of users, this application can realize the automatic filming of "movies" based on the "scripts" written by users according to their preferences, which greatly reduces the production costs compared to the traditional film and television production process. In addition, by replacing manual work with automation for video generation, the uncontrollable factors of human factors in the process are eliminated, and 7×24 hours of uninterrupted production can be achieved, which is much more efficient than traditional solutions.

[0054] As an implementation of the present application, before "inputting the target script into the first language model" in the above step S1, the following steps are further included:

[0055] Obtaining specified story elements and script formats; the story elements include: theme, atmosphere information, environment information, character characteristics, character relationships, and story outline; the script format includes: chapter information, text format, and word count requirements;

[0056] generating a second prompt word including the story elements and the script format;

[0057] The second prompt word is input into the first language model to generate the target script.

[0058] In one embodiment, a large language model is used to automatically create scripts, replacing the manual script creation process. First, the user-specified story elements and the script format to be created are obtained. The script format is used to control the script format output by the large language model, such as chapter format, scene table format, and general format. Story elements include: theme, atmosphere information, environment information, character characteristics, character relationships, and story outline, etc., as follows:

[0059] (1) Theme: This includes the core background of the story;

[0060] (2) Atmospheric information: including the time (era) and location of the story;

[0061] (3) Environmental information: including descriptions of static and dynamic objects in the specific environment, weather conditions, etc.

[0062] (4) Character characteristics include appearance, clothing style, personality, etc.

[0063] (5) Role relationships: introduce the relationships between characters, such as character relationships;

[0064] (6) Story synopsis: A brief description of the core of the story.

[0065] Based on the acquired story elements and script format, a second prompt word (i.e., prompt) is constructed to control the large language model to write the script as required. Optionally, a prompt word template is pre-built, and the detailed content of the story elements and the script format are respectively filled in the fixed position of the prompt word template to generate a standardized second prompt word of a unified standard, which is convenient for controlling the large language model. In the embodiment of the present application, the large language model for generating the script can be selected from GPT, Kimi, Doubao, Wenxinyiyan, etc., and this application does not limit this. Experimental verification shows that GPT and Kimi have a stronger sense of logic and drama than the scripts generated by other large language models. Further, the script generated by the large language model is used as the target script, and subsequent shot splitting, storyboard generation, audio matching, video generation and other automated processes are carried out.

[0066] In the traditional way, manually writing scripts (such as short stories, novels, etc.) requires a lot of time and has high professional requirements for the creators of the scripts, requiring high literary literacy and a lot of time to create, which is usually time-consuming. Most ordinary users who are not professionals only have a general idea of ​​the story, without professional cultural level and time and energy, and hope to obtain the finished product quickly. Based on this, the solution of the embodiment of the present application can replace the process of manually creating scripts, quickly generate scripts based on the story elements provided by the user and the script format to be created, reduce the requirements for the professional skills of users in script creation, meet the needs of users for low-cost and fast generation of personalized video content that suits their preferences, and improve the user experience.

[0067] As an implementation of the present application, the above step S1 specifically includes:

[0068] S11: Split the target script based on a preset word count to obtain at least one text to be processed;

[0069] S12: Construct a first prompt word; the first prompt word includes the script quantity and shooting parameters; the shooting parameters include: shooting position, shooting angle and shooting distance;

[0070] S13: Input each to-be-processed text and the first prompt word into the first language model respectively to generate a corresponding storyboard script.

[0071] In one embodiment, the target script is split based on a word count limit using a large language model. Different large language models are limited by their own performance in understanding the context of the text content during processing, and therefore have limits on the number of words in the input and output content. When splitting a shot, in order to avoid loss of text input to the model, a preset number of words needs to be set based on the performance of the first large language model, not exceeding the upper limit of the number of words that the model can process at a time. In actual applications, a preset number of words not exceeding the upper limit of the number of words that the selected first large language model can process at a time is set in advance. Then, based on the preset number of words, the target script is split to obtain one or more texts to be processed with a number of words not exceeding the preset number of words. For example, the upper limit of the number of words that the selected first large language model can process at a time is 5,000 words, and the preset number of words is set to 5,000. Then, based on the preset number of words, a text processing tool is used to split the target script into one or more text blocks (texts to be processed) of no more than 5,000 words.

[0072] Optionally, splitting the text to be processed can be implemented through Python code, or using text segmentation tools, command line tools, etc.

[0073] Construct the first prompt word, including the number of storyboards that need to be generated by the large language model and the shooting parameters of each storyboard. Specifically, by setting the number of scripts, the large language model is controlled to generate a specified number of storyboards based on the text content input in one call. For example, if the number of scripts is set to 1-2, the model will output 1-2 storyboards each time based on an input text to be processed. By setting the shooting parameters, the large model is controlled to select shooting parameters that match the currently input text content from the optional shooting parameters. The shooting parameters in the first prompt word include: shooting position (optional flat position, high position, low position, etc.), shooting angle (optional flat shot, overhead shot, upward shot, side shot, etc.), shooting distance (optional close-up, close shot, medium shot, long shot, etc.).

[0074] A single call is made to the large language model, and each time a text to be processed is input into the large language model together with the first prompt word, and a corresponding number of storyboards generated by the model are obtained. The storyboard is extracted from the target script and is a basis for guiding the shooting of the storyboard screen in text form. By generating a storyboard script, the abstract description content in the target script is converted into specific lens instructions, which can improve the efficiency of video generation and the accuracy of the video screen. In the embodiment of the present application, according to the actual input text to be processed, the storyboard includes: environmental information, character description (such as expression, action, status, etc.), narration and other information. For example, an example of a storyboard script is as follows:

[0075] Zhang San is knocked to the ground by Li Si, and the surrounding students laugh. Wang Wu looks on coldly. Many people, school playground exterior, sunny day.

[0076] [Zhang San, male, short hair, white T-shirt, jeans, center of the screen, looking up, close-up, falling to the ground, angry];

[0077] [Li Si, male, short hair, sportswear, right side of the screen, half left side of the face, medium shot, standing, provocative];

[0078] [Wang Wu, female, long hair, school uniform, left side of the screen, half right side of the face, mid-shot, standing, watching with indifferent eyes];

[0079] Flat shooting, flat camera position, standard lens.

[0080] After all the texts to be processed have generated corresponding storyboards, the large language model is used to summarize and integrate all the storyboards in the order in which the texts to be processed were processed.

[0081] As an implementation method of the present application, the target script includes chapter information; the above step S11 specifically includes:

[0082] S111: Based on chapter information and a preset word count, the target script is split to obtain at least one text to be processed; the text to be processed includes texts of an integer number of chapters, and the word count of the text to be processed is less than or equal to the preset word count.

[0083] Chapter information is information specified by the author according to the plot during the creation process of a script (for example, a novel) and is used to distinguish the development of the story at different stages of the plot. In one embodiment, a text processing tool is used to split the target script based on the chapter information, and one or more texts to be processed are obtained by splitting the text into integer chapters. In addition, considering the word limit of the large language model for processing text content, the word count of the text to be processed split by integer chapters cannot be greater than the word limit of the large language model. After the text to be processed is obtained by splitting the text into integer chapters, the word count of the text to be processed is further detected. If the word count of the text to be processed does not exceed the preset word count, the text to be processed is retained. If the word count of the text to be processed is greater than the preset word count, the text to be processed is discarded and the number of chapters is reduced by 1, and the text to be processed is split again. Finally, a text to be processed is obtained whose word count is not greater than the preset word count and whose chapters are complete.

[0084] For example, using a text processing tool, the target script is split into 2 chapters. If the number of words in the text to be processed is greater than the preset number of words 5,000, the text to be processed is discarded, the number of chapters is reduced by 1, and the script is split again according to the updated number of chapters 1 to obtain the text to be processed.

[0085] In an embodiment of the present application, if the target script includes chapter information, the target script is split based on a preset number of chapters and a preset number of words to obtain a text to be processed. Then, a storyboard is generated based on the content of the text to be processed with complete chapters. Because chapter information can accurately reflect the staged development of the plot, splitting the text to be processed by integer chapters helps the large language model generate more accurate and smoother storyboards, thereby improving the quality of the final video.

[0086] As an implementation manner of the present application, after the above step S13, the following is further included:

[0087] S14: Generate corresponding spatial information according to each storyboard using the large language model; the spatial information is used to indicate the positional relationship between different characters in the storyboard;

[0088] S15: Add the generated spatial information to the corresponding storyboard to update the storyboard.

[0089] In one embodiment, since there is no intuitive description of spatial information in the target script (for example, Zhang San stands behind Li Si), the spatial structure of the storyboard generated based on the storyboard script and the front-back relationship of the characters do not match. Based on this, in an embodiment of the present application, after the large language model generates a corresponding number of storyboards based on the currently input text to be processed, the text to be processed is further analyzed to obtain implicit spatial information. Based on pre-set spatial keywords, the text to be processed is detected, and the positional relationship between the characters is judged based on the spatial keywords contained in the text to be processed to obtain spatial information.

[0090] Optionally, spatial keywords include static directional words and dynamic directional words. Specifically, static directional words include: left, right, front, back, above, below, left front, right back, diagonally above, directly below, northwest corner, middle, beside, inside, outside, top, bottom, surface, back, behind, in front of, beside, feet, overhead, behind, back to back, face to face, shoulder to shoulder, etc.; dynamic directional words include: turn around, sideways, lean over, lie on your back, bend over, squat, stride forward, retreat, approach, move away, go around, turn back, turn your head, look up, look down, tilt your head, turn left, turn right, and make a U-turn.

[0091] Add the generated spatial information to each currently generated storyboard to complete the update, ensuring that the spatial information of the storyboards for consecutive shots remains consistent. In practice, you can add an item to the storyboard to fill in the corresponding spatial information. If no spatial information exists, this item is set to empty.

[0092] For example, based on the content of the text to be processed, "After hearing Zhang San's voice, Li Si turned his head and looked behind him, only to see Zhang San approaching with a book," the spatial keywords "turn head, behind" are detected, and Zhang San's position is inferred to be behind Li Si. Based on this, the corresponding spatial information "Li Si is close to the camera, Zhang San is behind Li Si" is generated and added to the generated storyboard. The updated storyboard is as follows:

[0093] After hearing Zhang San's voice, Li Si turned around and saw Zhang San walking towards him with a book. It was a sunny day in the corridor of the teaching building.

[0094] [Zhang San, male, short hair, white T-shirt, jeans, right side of the screen, medium long shot, approaches the camera and speaks];

[0095] [Li Si, male, short hair, sportswear, center of the screen, front view, close-up, standing, turning head];

[0096] Flat shooting, flat camera position, standard lens;

[0097] Spatial information: Li Si is close to the camera, and Zhang San is behind Li Si.

[0098] As an implementation manner of the present application, the target audio at least includes a narrator audio; step S5 specifically includes:

[0099] S51-1: Input each storyboard script into a text processing tool to extract the narrator's words;

[0100] S51-2: Determine the first timbre corresponding to the narrator's words according to the audio configuration information; the audio configuration information at least includes: the first timbre for the narrator;

[0101] S51-3: Generate the narrator audio through a speech synthesis model according to the narrator's words and the first timbre.

[0102] In one embodiment, the corresponding target audio is generated according to each storyboard script, including the narrator audio. The narrator's words in the storyboard script include: all text contents except the character dialogues and shooting parameters. In the embodiments of the present application, a speech synthesis model is used to implement the generation of the narrator audio. In practical applications, the speech synthesis model can be a TTS (text to speech) model.

[0103] Specifically, each sentence in the storyboard script is processed to detect whether there are character tags (such as: character names or character levels) and personal pronouns (such as: I, you, he, etc.) in the storyboard script. If there are character tags or personal pronouns in a sentence, the sentence is determined as a character dialogue. If there are no character tags and personal pronouns in a sentence, the sentence is determined as a narrator's words.

[0104] For example, some sentences of the storyboard script are as follows:

[0105] Sentence 1: After Li Si heard Zhang San's voice, he turned his head and looked behind him, and saw Zhang San walking over with a book.

[0106] Sentence 2: [Zhang San] Class is about to start. Why haven't you gone to the classroom yet?

[0107] Sentence 3: [Li Si] I just came out of the teacher's office and I'm going there now.

[0108] Sentence 4: Then, the two of them walked towards the classroom together.

[0109] The text processing tool detects the above 4 sentences. Character tags "[Zhang San], [Li Si]" and personal pronouns "you, I" are detected in Sentence 2 and Sentence 3, and Sentence 2 and Sentence 3 are determined as character dialogues. Character tags and personal pronouns are not detected in Sentence 1 and Sentence 4, so Sentence 1 and Sentence 4 are determined as the narrator's words.

[0110] In an embodiment of the present application, audio configuration information is pre-set, including: a first timbre for narration. For example, the first timbre specified in the audio configuration information is baritone No. 1, that is, all narration audio uses baritone No. 1. After extracting the narration from the storyboard script, the first timbre for the narration is obtained according to the pre-set audio configuration information. The speech synthesis model TTS is called to generate the corresponding narration audio according to the text and the first timbre of the narration. In addition, the same first timbre needs to be set for the narration in all storyboard scripts of the target script to ensure that the target video finally generated has good coherence in both vision and sound.

[0111] Furthermore, a generative model is used to generate a storyboard video based on the storyboards and the corresponding target audio. In the absence of character dialogue, AI (artificial intelligence) tools such as Luma and Pika can be used to generate a corresponding single-shot dynamic video, i.e., a storyboard video, based on each storyboard and the target audio.

[0112] As an implementation manner of the present application, the storyboard script further includes environmental information; the target audio further includes environmental sound effects; and step S5 further includes:

[0113] S52-1: extracting environmental information from each storyboard, and extracting at least one keyword from the environmental information;

[0114] S52-2: Based on the at least one keyword, obtain at least one matching sound effect from a sound effect library and determine it as a candidate sound effect;

[0115] S52-3: Match each candidate sound effect with the storyboard, and calculate the matching degree of the candidate sound effect;

[0116] S52-4: Filter out candidate sound effects whose number is less than or equal to a first threshold value from all candidate sound effects according to the matching degree from high to low, and determine each of the filtered candidate sound effects as the environmental sound effect corresponding to the storyboard script.

[0117] In one embodiment, the natural language processing model is fine-tuned in advance using an environmental keyword library, and the environmental keywords in the storyboard are identified using the natural language processing model. The environmental keyword library includes:

[0118] Weather keywords, such as wind, rain, snow, thunder and lightning, fog, sunny, cloudy, etc.

[0119] Terrain keywords, such as mountains, rivers, deserts, forests, cities, battlefields, wars, etc.

[0120] Atmosphere keywords, such as dim, silent, noisy, bloody, desolate, etc.

[0121] A natural language processing model identifies environmental keywords from storyboards and extracts context based on their placement within the script, generating environmental information. For example, the keyword "battlefield" is extracted from the storyboard, along with the context "the roar of thousands of soldiers and the neighing of war horses on a darkly clouded ancient battlefield," which is then identified as environmental information.

[0122] Then, the natural language processing model extracts phrases containing sound verbs or onomatopoeia from the context as keywords. For example, "the roar of soldiers" and "the neighing of war horses" are used as keywords.

[0123] Furthermore, the natural language processing model performs associative expansion on nouns in the environmental information to obtain implicit keywords. For example, the associative expansion of the noun "battlefield" in the environmental information yields implicit keywords such as "weapon collision, crowd footsteps"; and the associative expansion of the noun "dark cloud" yields implicit keywords such as "thunder, wind."

[0124] Based on each keyword, one or more matching candidate sound effects are retrieved from the sound effects library. Specifically, each keyword is matched against the labels of each sound effect in the sound effects library to obtain the corresponding candidate sound effects. For example, based on the keywords "soldier's roar, horse's neigh, weapon's collision, crowd's footsteps, thunder, wind," a similarity match is performed in the sound effects library to obtain multiple matching candidate sound effects.

[0125] To ensure a better match between the sound effects and the visuals in the generated storyboard video, candidate sound effects are matched to the storyboard, the matching degree of each candidate sound effect is calculated, and all candidate sound effects are sorted from high to low based on matching degree to select sound effects that better match the visuals. Furthermore, when selecting candidate sound effects, priority can be given to those with higher matching degrees. For example, if the storyboard does not include content such as guns and cannons, the calculated matching degree of the sound of gunfire will be lower after matching with the storyboard. Furthermore, to avoid a cluttered sound effect in the storyboard video, which would reduce the video's enjoyment, a number of candidate sound effects, no greater than a first threshold, is selected based on the ranking results of the candidate sound effects according to their matching degree to serve as the final candidate sound effects for generating the storyboard video.

[0126] As an embodiment of the present application, the storyboard script further includes character dialogue; the target audio further includes dialogue audio; the audio configuration information further includes: second timbres for characters of different levels; the above step S5 further includes:

[0127] S53-1: extracting the dialogue of each character from the storyboard using the text processing tool;

[0128] S53-2: Determine a second timbre corresponding to each character according to the audio configuration information;

[0129] S53-3: Generate dialogue audio according to the dialogue of each character and the corresponding second timbre through the speech synthesis model.

[0130] In one embodiment, some storyboards may also include character dialogues. In order to further enhance the user experience, a speech synthesis model is used in the embodiment of the present application to generate dialogue audio with different tones for each character, so that users can have a better viewing experience when watching the target video.

[0131] Specifically, the audio configuration information includes not only the primary timbre used for narration, but also secondary timbres for characters of different ranks. For example, the primary timbre used for narration is "Timbre M," the secondary timbre used for the male protagonist is "Timbre A," the secondary timbre used for the female protagonist is "Timbre B," the secondary timbre used for the male supporting role is "Timbre C," the secondary timbre used for the female supporting role is "Timbre D," and the secondary timbre used for non-player characters (NPCs) is "Timbre E."

[0132] The dialogues of each character are extracted from each storyboard script, and then the corresponding second timbre of each character's dialogue is determined based on the audio configuration information. The speech synthesis model is further called to generate the dialogue audio of the character based on the character's dialogue and the corresponding second timbre.

[0133] In the case where the storyboard script includes character dialogue, a storyboard video is generated based on the storyboard images corresponding to each storyboard script, the target audio (eg, narration audio), and the dialogue audio of all characters in the storyboard script. Figure 2 This is one of the flowcharts for generating storyboard videos in one embodiment of the present application. Figure 2 The present invention shows a process for generating multiple storyboard videos based on a target script. First, the target script is split to generate multiple storyboard scripts, and spatial information is extracted from the target script to enhance the spatial details of the storyboard script and improve the accuracy of the subsequently generated storyboard images. In addition, the characteristic information of each character is extracted from the target script, and a character image image is generated based on the characteristic information. Then, a corresponding storyboard image is generated based on each storyboard script, and corresponding target audio (including environmental sound effects and narration audio) is matched for each storyboard script. In the process of generating storyboard images based on the storyboard script, the corresponding character image image is used as a reference benchmark to ensure the consistency of the same character image in all storyboard images. Finally, a corresponding storyboard video is generated based on each storyboard image and the matching target audio.

[0134] As an implementation of the present application, the above step S6 includes:

[0135] S61: Segmenting the storyboard to obtain at least two area maps; wherein each area map contains only one character;

[0136] S62: For each region map, based on the dialogue audio of each character, using a video-to-image tool, drives the facial expression and lip shape of the character in the region map to adapt to the dialogue audio, thereby generating a corresponding region video; the lip shape and facial expression of the character in the region video are adapted to the dialogue audio;

[0137] S63: Synthesize all the regional videos corresponding to each storyboard and the target audio except the dialogue audio to obtain the storyboard video corresponding to the storyboard.

[0138] In an embodiment of the present application, when the storyboard script includes dialogues of multiple characters, it is necessary to drive the lip shape and facial expression of the character in the storyboard image for each character's dialogue, so that the lip shape and facial expression of the character in the generated storyboard video match the dialogue audio, so as to improve the authenticity and viewing experience of the video.

[0139] Figure 3 This is the second flow chart of generating storyboard video in one embodiment of the present application. Figure 3 The process of generating a storyboard video based on a storyboard is shown. First, the storyboard is segmented according to the characters in the storyboard for which dialogue audio needs to be generated, and a region map containing only a single character is obtained. For example, in the case where the storyboard script contains dialogues between two characters "A and B", the storyboard is segmented into two region maps, namely: a first region map containing only "A", and a second region map containing only "B". Optionally, the storyboard can be segmented by an image processing model based on deep learning. Then, for each character's region map and the corresponding dialogue audio of the character, the image-generated video tool is used to drive the character's facial expression and lip shape to align, and obtain the region videos corresponding to each character.

[0140] Optionally, the image-driven algorithm EMO can be used to align the lip movements and facial expressions of characters in the storyboard with the audio dialogue. The portrait animation algorithm LivePortrait can also be used to adjust the facial expressions of characters in the storyboard. The image generation tool Act-One can also be used to generate regional videos based on the facial movements and labels of real-person dialogue readings.

[0141] Furthermore, all the regional videos and other target audios (e.g., ambient sound effects, narration audio) except for the dialogue audio are synthesized to obtain a storyboard video. Specifically, first, based on the position of the regional graph corresponding to each regional video in the storyboard, all the regional videos are spliced ​​according to the corresponding position to obtain a video whose screen size matches the storyboard. Then, a new audio track is added to the video and the target audio is added to it to obtain a regional video. For example, the storyboard includes two characters, Zhang San and Li Si. The storyboard is segmented to obtain the region containing only Zhang San. Figure 1 , the area containing only John Figure 2 , and background other than Zhang San and Li Si Figure 3 . For the region Figure 1 ,area Figure 2 , use the EMO algorithm to drive the facial image of the character in the picture and obtain the corresponding two area videos.

[0142] Traditional solutions based on facial expression and lip-syncing to audio can only drive a single-shot shot to generate a close-up video of a single person speaking, and cannot generate a video of a multi-shot dialogue scene. The embodiment of the present application uses a split-first-then-synthesize approach to first segment the storyboard to obtain a region map corresponding to each character. Based on the region map, the facial expression, lip-syncing, and dialogue audio of a single person are aligned. Then, the videos of each region are resynthesized to generate a video of a multi-shot dialogue scene, enriching the video image and further improving the video's viewing experience.

[0143] As an implementation of the present application, the above step S7 includes:

[0144] S71: extracting the plot outline of the target script;

[0145] S72: Obtain a piece of music that matches the story outline and determine it as background music;

[0146] S73: Combine all the storyboard videos and the background music to generate the target video.

[0147] In an embodiment of the present application, the story outline of the target script is extracted by a large language model, and matching background music is obtained based on the story outline. In an embodiment of the present application, the background music can be existing music or accompaniment matched from a music library based on the story outline, or it can be music or accompaniment generated based on the story outline. In one embodiment, the music generation tool Suno is used to generate corresponding music based on the story outline of the target script as the background music for the target video. The Suno tool can generate a complete musical work based on the text description (including: lyrics, style, theme, etc.), including melody, arrangement and vocals. According to the order of each storyboard script, all the storyboard videos are spliced ​​and merged, and background music is added to generate the target video for final release.

[0148] Based on the same inventive concept, an embodiment of the present application provides a device for generating a video. Figure 4 , Figure 4 FIG is a schematic diagram of a video generating apparatus 400 proposed in an embodiment of the present application. Figure 4 As shown, the device includes:

[0149] A splitting module 401 is used to input the target script into the first language model to obtain multiple storyboards;

[0150] The character construction module 402 is configured to input the target script into a second language model to extract characteristic information of each character from the target script through the second language model; input the characteristic information into a generation tool to generate a corresponding character image;

[0151] A storyboard generation module 403 is configured to input the character image and each corresponding storyboard script into a third language model to obtain a storyboard image corresponding to each storyboard script;

[0152] An audio generation module 404 is used to match corresponding target audio for each storyboard script;

[0153] The video generation module 405 is used to generate a corresponding storyboard video using a picture-to-video tool according to the storyboard image and target audio corresponding to each storyboard script;

[0154] The fusion module 406 is used to combine all the storyboard videos to obtain the target video.

[0155] As an implementation manner of the present application, the splitting module 401 is configured to perform the following steps:

[0156] Splitting the target script based on a preset word count to obtain at least one text to be processed;

[0157] Constructing a first prompt word; the first prompt word includes the script quantity and shooting parameters; the shooting parameters include: shooting position, shooting angle and shooting distance;

[0158] Each text to be processed and the first prompt word are respectively input into the first language model to generate a corresponding storyboard script.

[0159] As an implementation manner of the present application, the target script includes chapter information; the splitting module 401 is further configured to perform the following steps:

[0160] Based on chapter information and a preset word count, the target script is split to obtain at least one text to be processed; the text to be processed includes texts of an integer number of chapters, and the word count of the text to be processed is less than or equal to the preset word count.

[0161] As an implementation manner of the present application, the splitting module 401 is further configured to perform the following steps:

[0162] Generate corresponding spatial information according to each storyboard using the large language model; the spatial information is used to indicate the positional relationship between different characters in the storyboard;

[0163] The generated spatial information is added to the corresponding storyboard to update the storyboard.

[0164] As an embodiment of the present application, the target audio includes at least narration audio; the audio generation module 404 is used to match the corresponding target audio for each storyboard script, specifically including:

[0165] Each storyboard was fed into a text processing tool to extract the narration;

[0166] Determining a first timbre corresponding to the narration according to audio configuration information; the audio configuration information at least includes: a first timbre for the narration;

[0167] A narration audio is generated according to the narration and the first timbre through a speech synthesis model.

[0168] As an embodiment of the present application, the storyboard script further includes environmental information; the target audio further includes environmental sound effects; the audio generation module 404 is configured to match the corresponding target audio for each storyboard script, and further includes performing the following operations:

[0169] extracting environmental information from each storyboard, and extracting at least one keyword from the environmental information;

[0170] Based on the at least one keyword, obtaining at least one matching sound effect from a sound effect library and determining it as a candidate sound effect;

[0171] Matching each candidate sound effect with the storyboard, and calculating a matching degree of the candidate sound effect;

[0172] According to the matching degree from high to low, candidate sound effects less than or equal to the first threshold number are screened out from all candidate sound effects, and each screened candidate sound effect is determined as the environmental sound effect corresponding to the storyboard script.

[0173] As an embodiment of the present application, the storyboard script further includes character dialogue; the target audio further includes dialogue audio; the audio configuration information further includes: second timbres for characters of different levels; the audio generation module 404 is used to match the corresponding target audio for each storyboard script, and further includes the following steps:

[0174] Extracting the dialogue of each character from the storyboard using the text processing tool;

[0175] Determining a second timbre corresponding to each character according to the audio configuration information;

[0176] The speech synthesis model generates dialogue audio according to the dialogue of each character and the corresponding second timbre.

[0177] As an embodiment of the present application, the video generation module 405 is used to generate a corresponding storyboard video using a picture-to-video tool according to the storyboard image and target audio corresponding to each storyboard script, including:

[0178] Segmenting the storyboard to obtain a region map of each character in the storyboard;

[0179] For each region map, based on the dialogue audio of each character, the image-to-video tool drives the facial expressions and lip movements of the characters in the region map to adapt to the dialogue audio, generating a corresponding region video;

[0180] All the regional videos corresponding to each storyboard and the target audio except the dialogue audio are synthesized to obtain the storyboard video corresponding to the storyboard.

[0181] As an implementation manner of the present application, the fusion module 406 is used to combine all the storyboard videos to obtain the target video, specifically including:

[0182] Extracting a story synopsis of the target script;

[0183] Obtaining a piece of music that matches the story outline and determining it as background music;

[0184] All the storyboard videos and the background music are combined to generate the target video.

[0185] As an embodiment of the present application, the device further includes a script construction module for performing the following steps:

[0186] Obtaining specified story elements and script formats; the story elements include: theme, atmosphere information, environment information, character characteristics, character relationships, and story outline; the script format includes: chapter information, text format, and word count requirements;

[0187] generating a second prompt word including the story elements and the script format;

[0188] The second prompt word is input into the first language model to generate the target script.

[0189] Based on the same inventive concept, an embodiment of the present application provides a readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for generating a video as described in any of the above embodiments of the present application are implemented.

[0190] Based on the same inventive concept, an embodiment of the present application provides an electronic device, referring to Figure 5 , Figure 5 1 is a schematic diagram of an electronic device according to an embodiment of the present application. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the steps of the method for generating a video as described in any of the above embodiments of the present application.

[0191] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0192] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0193] For the sake of simplicity, the method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and components involved are not necessarily required by this application.

[0194] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0195] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0196] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0198] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the underlying inventive concepts. Therefore, this application is intended to include the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0199] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0200] The above is a detailed introduction to the method, device, equipment and storage medium for generating videos provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for generating a video, characterized in that: include: Input the target script into the first language model to obtain multiple storyboards; Inputting the target script into a second language model to extract feature information of each character from the target script through the second language model; Inputting the characteristic information into a generation tool to generate a corresponding character image; Inputting the character image and each corresponding storyboard into a third language model to obtain a storyboard corresponding to each storyboard; Match the target audio to each storyboard; Generate the corresponding storyboard video using the image-to-video tool based on the storyboard images and target audio for each storyboard script; Combine all storyboard videos to get the target video.

2. The method for generating a video according to claim 1, wherein: The target script is input into the first language model to obtain multiple storyboards, including: Splitting the target script based on a preset word count to obtain at least one text to be processed; Constructing a first prompt word; the first prompt word includes the script quantity and shooting parameters; the shooting parameters include: shooting position, shooting angle and shooting distance; Each text to be processed and the first prompt word are respectively input into the first language model to generate a corresponding storyboard script.

3. The method for generating a video according to claim 2, wherein: The target script includes chapter information; the target script is split based on a preset word count to obtain at least one text to be processed, including: Based on chapter information and a preset word count, the target script is split to obtain at least one text to be processed; the text to be processed includes texts of an integer number of chapters, and the word count of the text to be processed is less than or equal to the preset word count.

4. The method for generating a video according to claim 2 or 3, characterized in that: After inputting each to-be-processed text and the first prompt word into the first language model to generate a corresponding storyboard, the method further includes: Generate corresponding spatial information according to each storyboard using the large language model; the spatial information is used to indicate the positional relationship between different characters in the storyboard; The generated spatial information is added to the corresponding storyboard to update the storyboard.

5. The method for generating a video according to claim 1, wherein: The target audio includes at least narration audio; Matching the corresponding target audio for each storyboard script includes: Each storyboard was fed into a text processing tool to extract the narration; Determining a first timbre corresponding to the narration according to audio configuration information; the audio configuration information at least includes: a first timbre for the narration; A narration audio is generated according to the narration and the first timbre through a speech synthesis model.

6. The method for generating a video according to claim 5, wherein: The storyboard script further includes environmental information; the target audio further includes environmental sound effects; and the step of matching the corresponding target audio for each storyboard script further includes: extracting environmental information from each storyboard, and extracting at least one keyword from the environmental information; Based on the at least one keyword, obtaining at least one matching sound effect from a sound effect library and determining it as a candidate sound effect; Matching each candidate sound effect with the storyboard, and calculating a matching degree of the candidate sound effect; According to the matching degree from high to low, candidate sound effects less than or equal to the first threshold number are screened out from all candidate sound effects, and each screened candidate sound effect is determined as the environmental sound effect corresponding to the storyboard script.

7. The method for generating a video according to claim 5 or 6, characterized in that: The storyboard script also includes character dialogue; the target audio also includes dialogue audio; the audio configuration information also includes: a second timbre for characters of different levels; the method of matching the corresponding target audio for each storyboard script also includes: Extracting the dialogue of each character from the storyboard using the text processing tool; Determining a second timbre corresponding to each character according to the audio configuration information; The speech synthesis model generates dialogue audio according to the dialogue of each character and the corresponding second timbre.

8. The method for generating a video according to claim 7, wherein: Based on the storyboards and target audio for each storyboard script, the corresponding storyboard videos are generated using the image-to-video tool, including: Segmenting the storyboard to obtain at least two area maps; wherein each area map contains only one character; For each region map, based on the dialogue audio of each character, the image-to-video tool drives the facial expressions and lip movements of the characters in the region map to adapt to the dialogue audio, generating a corresponding region video; All the regional videos corresponding to each storyboard and the target audio except the dialogue audio are synthesized to obtain the storyboard video corresponding to the storyboard.

9. The method for generating a video according to claim 1, wherein: Combine all storyboard videos to get the target video, including: Extracting a story synopsis of the target script; Obtaining a piece of music that matches the story outline and determining it as background music; All the storyboard videos and the background music are combined to generate the target video.

10. The method for generating a video according to claim 1, wherein: Before the target script is fed into the first language model, it also includes: Obtaining specified story elements and script formats; the story elements include: theme, atmosphere information, environment information, character characteristics, character relationships, and story outline; the script format includes: chapter information, text format, and word count requirements; generating a second prompt word including the story elements and the script format; The second prompt word is input into the first language model to generate the target script.

11. A device for generating a video, characterized in that: Used to implement the method according to any one of claims 1 to 10, comprising: The splitting module is used to input the target script into the first language model to obtain multiple storyboards; a character construction module, configured to input the target script into a second language model to extract characteristic information of each character from the target script through the second language model; and input the characteristic information into a generation tool to generate a corresponding character image; A storyboard generation module is configured to input the character image and each corresponding storyboard script into a third language model to obtain a storyboard image corresponding to each storyboard script; Audio generation module, used to match the corresponding target audio for each storyboard; A video generation module is used to generate a corresponding storyboard video using a picture-to-video tool based on the storyboard image and target audio corresponding to each storyboard script; The fusion module is used to combine all the storyboard videos to obtain the target video.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 10 are implemented.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps in the method according to any one of claims 1 to 10 are implemented.

Citation Information

Cited By

  • Short episode generation method and system, electronic equipment and storage medium

    CN121194034A

  • AI-based animation sub-mirror script automatic generation and visual preview method and system

    CN121236236A

  • Content generation method and device, computer readable storage medium and program product

    CN121334418A

  • Short drama creation method, device and equipment, storage medium and program product

    CN121509776A

  • Short play video generation method and device, equipment, storage medium and program product

    CN121509781A