Video generation method and device, electronic equipment and storage medium
By acquiring the storyboard description text of the novel and generating storyboard images and audio, the video generation process is automated, solving the problems of low efficiency and high labor costs in existing technologies, and achieving efficient video generation.
Patent Information
- Application Number
- CN202510710510.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies for converting novel texts into videos require human interpretation of the content and drawing of images, resulting in low video generation efficiency and high labor costs.
By obtaining descriptive text for multiple storyboards to be generated based on the initial text, and using image generation models and text-to-speech conversion technology, corresponding storyboard images and audio are generated for each storyboard, and finally, a video is synthesized.
It improves the automation of novel text to video, reduces human intervention, lowers the manual cost of video generation, and significantly improves generation efficiency.
Smart Images

Figure CN121531201A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computers, and particularly relates to a video generation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the development of computer technology, the promotion form of novel texts is increasingly diversified. In order to attract more readers, novel texts often need to be videoized and put on various platforms.
[0003] At present, when novel texts are videoized, it is often necessary to manually understand the content of the novel and draw images according to the content of the novel, and then make a video. Obviously, the video generation efficiency of this method is poor, and requires a very high labor cost. SUMMARY
[0004] The present application provides a video generation method, device, electronic device, and storage medium to solve the problem of high labor cost in video generation.
[0005] To solve the above technical problems, the present application is implemented as follows: In a first aspect, the present application provides a video generation method, characterized in that the method comprises: Based on the initial text, a plurality of description texts corresponding to a plurality of to-be-generated shot scripts are obtained; Based on the description text corresponding to each of the to-be-generated shot scripts, a corresponding shot image is generated for each of the to-be-generated shot scripts through a preset image generation model; Each of the description texts is subjected to text-to-speech conversion processing to obtain a shot voice corresponding to each of the to-be-generated shot scripts; Based on the shot image and the shot voice corresponding to each of the to-be-generated shot scripts, a video is generated for the initial text.
[0006] In a second aspect, the present application provides a video generation device, characterized in that the device comprises: A first obtaining module is configured to obtain a plurality of description texts corresponding to a plurality of to-be-generated shot scripts based on an initial text; A first generating module is configured to generate a corresponding shot image for each of the to-be-generated shot scripts based on the description text corresponding to each of the to-be-generated shot scripts through a preset image generation model; A synthesizing module is configured to perform text-to-speech conversion processing on each of the description texts to obtain a shot voice corresponding to each of the to-be-generated shot scripts; A second generating module is configured to generate a video for the initial text based on the shot image and the shot voice corresponding to each of the to-be-generated shot scripts.
[0007] In a third aspect, the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of the first aspect when executing the program.
[0008] In a fourth aspect, the present application provides a readable storage medium, which enables an electronic device to implement the method of the first aspect when an instruction in the storage medium is executed by a processor of the electronic device.
[0009] The video generation method provided by the embodiment of the present application comprises the steps of: obtaining description texts corresponding to a plurality of to-be-generated shot scripts based on an initial text; generating corresponding shot image for each of the to-be-generated shot scripts based on the description text corresponding to each of the to-be-generated shot scripts through a preset image generation model; performing text-to-speech conversion processing on each of the description texts to obtain shot speech corresponding to each of the to-be-generated shot scripts; and generating a video for the initial text based on the shot image corresponding to each of the to-be-generated shot scripts and the shot speech. In this way, the embodiment of the present application can divide the text into different shot scripts and realize the image of each shot script by obtaining the description texts of a plurality of to-be-generated shot scripts according to the initial text and generating the corresponding shot image and shot speech of each to-be-generated shot script through the image generation technology and text-to-speech conversion. Further, the video can be generated through the shot image and shot speech of each to-be-generated shot script, which can improve the automation degree of text video, reduce the manual participation, and greatly reduce the labor cost of video generation. At the same time, the video can be automatically generated to a certain extent according to the initial text, which greatly improves the video generation efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0011] Figure 1 is a step flowchart of a video generation method provided by the embodiment of the present application; Figure 2 is a structural schematic diagram of a video generation system provided by the embodiment of the present application; Figure 3 is a flowchart of another video generation method provided by the embodiment of the present application; Figure 4 is a flowchart of a script processing provided by the embodiment of the present application; Figure 5This is a schematic diagram of an image generation process provided by an embodiment of the present invention; Figure 6 This is a schematic diagram of a video synthesis process provided in an embodiment of the present invention; Figure 7 This is a structural diagram of a video generation device provided in an embodiment of the present invention; Figure 8 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Figure 1 This is a flowchart of the steps of a video generation method provided in an embodiment of the present invention, as follows: Figure 1 As shown, the method may include the following steps: Step 101: Based on the initial text, obtain the descriptive text corresponding to multiple storyboards to be generated.
[0014] Step 102: Based on the description text corresponding to each of the storyboards to be generated, generate corresponding storyboard images for each of the storyboards to be generated using a preset image generation model.
[0015] Step 103: Perform text-to-speech conversion on each of the described texts to obtain the speech corresponding to each of the storyboards to be generated.
[0016] Step 104: Generate video for the initial text based on the storyboard images corresponding to each of the storyboards to be generated and the storyboard audio.
[0017] In this embodiment of the invention, steps 101 to 104 above can be applied to any reading platform that has a need for novel promotion. Accordingly, the initial text can be the novel text to be promoted, and can be any type of novel content, such as urban, suspense, fantasy, etc. This embodiment of the invention does not limit this.
[0018] Specifically, the initial text mentioned above can be determined by receiving input from relevant personnel, or it can be selected from the text library of the reading platform according to the promotion needs of the reading platform. This embodiment of the invention does not limit this.
[0019] The storyboard refers to a storyboard, and usually refers to various image media such as films, animations, television series, advertisements, music videos, etc. Before actual shooting or drawing, the composition of the image is illustrated in a chart, and the continuous picture is divided into a unit of one lens movement, and the lens movement mode, time length, dialogue, special effect, etc. are marked. The generated storyboard in the embodiment of the application can correspond to different novel plots, so that the novel text can be imaged in units of storyboards, and one storyboard can correspond to one plot or scene. Accordingly, the description text described above for generating the plot content of the storyboard can include information such as characters, scenes, and events.
[0020] Specifically, the embodiment of the application can obtain the novel plot from the initial text, and divide different storyboards according to the novel plot. Specifically, one storyboard can include scene, character, background, time, character dialogue, character action, etc. Alternatively, the embodiment of the application can also divide the storyboard based on the scene in the initial text, for example, the plot of the same scene can correspond to one storyboard, and the storyboard division mode can be set according to actual needs, and the embodiment of the application does not limit this.
[0021] Specifically, the above step 101 can be realized by a pre-trained large language model (Large Language Model, LLM). The LLM model can divide the initial text into multiple storyboards based on the character information and the novel plot in the initial text, and obtain the description text corresponding to each storyboard. Alternatively, the above LLM model can be a large language model (gpt-4o) of Microsoft, and of course it can also be other LLM models, and the embodiment of the application does not limit this.
[0022] Further, the number of the generated storyboards can be set according to the length of the initial text and the actual video generation requirement, and the embodiment of the application does not limit this.
[0023] Further, after step 101, the embodiment of the application can obtain the description text corresponding to the plurality of generated storyboards. The description text of different generated storyboards corresponds to different scenes or different plots, and then the embodiment of the application can generate the corresponding storyboard image and storyboard voice for each generated storyboard. The storyboard image refers to an image for displaying the plot content described by the description text of the generated storyboard. Accordingly, the above storyboard voice refers to audio data for displaying the plot content described by the description text of the generated storyboard, which can include plot description audio and character dialogue audio.
[0024] Specifically, one to-be-generated split shot can correspond to one split shot image, of course, can also correspond to two or more than two split shot images, which can be set according to the richness of the description text of the to-be-generated split shot, and the embodiments of the present application do not limit this.
[0025] Specifically, the above split shot image can be generated by a pre-trained image generation model, for example: a stable diffusion (SD) image generation model, of course, other image generation models can also be used, and the embodiments of the present application do not limit this. Specifically, for any to-be-generated split shot, the character information, scene information, background information and the like contained in the to-be-generated split shot can be input into the SD model, and the SD model can generate an image corresponding to the input information, and then the embodiments of the present application can obtain the split shot image corresponding to the to-be-generated split shot by obtaining the output data of the SD model. The above split shot image can be in a cartoon style, of course, it can also be in a realistic style, which can be set according to actual promotion needs, and the embodiments of the present application do not limit this.
[0026] Specifically, the above split shot voice can be obtained through text-to-speech processing, and specifically, a text-to-speech (TTS) tool or a TTS model can be used to realize it. Since the text in the novel usually contains plot description text for describing the plot and dialogue text of the character role, the above split shot voice can contain dialogue voice and plot description voice. The embodiments of the present application can obtain the description text corresponding to the to-be-generated split shot, and then convert the description text into audio data through TTS, that is, the split shot voice corresponding to the to-be-generated split shot can be obtained.
[0027] Further, after obtaining the split shot image and the split shot voice, the embodiments of the present application can encode the voice and the image, that is, the split shot image and the split shot voice corresponding to the same to-be-generated split shot are merged, so that the voice data is consistent with the image content, and then the split shot image and the split shot voice of different to-be-generated split shots are spliced according to the order in the initial text, that is, the video corresponding to the initial text can be obtained. At this time, the video can convey the novel content to the reader through the image and the audio at the same time, so as to diversify the form of the novel.
[0028] Optionally, since the video encoding process is often susceptible to hardware resource constraints, which can result in low video encoding efficiency, in order to meet the resource requirements of video encoding, the embodiment of the application can also generate a video through a pre-set cloud rendering module, and can further add music, animation and other elements to the video through the cloud rendering module, which can be set according to the actual scene, and the embodiment of the application does not limit this. Specifically, the cloud rendering module can first preprocess the shot image and the shot voice, and can respectively denoise or reduce noise of the image and the voice, and adjust the resolution of the shot image. Further, the shot image can be synchronized on the time axis based on the voice time length, so that the display time of a shot image is consistent with the time length of the corresponding shot voice. Further, the voice and the image synchronized on the time axis can be encoded into a video according to a pre-set encoding format or encoding algorithm. Further, in the encoding process, music or animation in a pre-set template can also be used to add music or animation effects to the video.
[0029] To sum up, the video generation method provided by the embodiment of the application can obtain a plurality of description texts of to-be-generated shots according to the initial text, and can generate corresponding shot images and shot voices for each to-be-generated shot through image generation technology and text-to-speech conversion, so as to divide the text into different shots and realize the imaging of each shot. Further, a video can be generated through the shot image and the shot voice of each to-be-generated shot, which can improve the automation degree of text video, reduce manual participation, and greatly reduce the labor cost of video generation. At the same time, since the video can be automatically generated to a certain extent according to the initial text, the video generation efficiency is greatly improved.
[0030] Optionally, in the embodiment of the application, the operation of obtaining a plurality of description texts corresponding to to-be-generated shots based on the initial text in the above step 102 can specifically include: S21, performing character feature extraction and content extraction on the initial text respectively to obtain character information contained in the initial text and a content summary text corresponding to the initial text.
[0031] S22, inputting a split prompt word to a large language model, and inputting the character information and the content summary text into the large language model, wherein the split prompt word is used to instruct the large language model to perform shot splitting on the content summary text based on the character information to obtain a plurality of description texts corresponding to to-be-generated shots, and each description text of the to-be-generated shot at least contains at least one character in the character information and scene description content related to the character.
[0032] The character information can include character identification and character features. The character identification can be a character name and / or a character code. The character features can be image features of the character, including hairstyle features (e.g., hair color and style), facial features (eyes, eyebrows, nose, mouth, and face shape), and in some cases, other features (e.g., wearing glasses, sitting in a wheelchair, having facial scars, etc.).
[0033] Specifically, the character feature extraction can be achieved by a pre-trained character recognition model, or by the LLM model. Specifically, a prompt can be input into the LLM model. The prompt of the LLM model can be used to prompt the type of information required by the LLM model. Accordingly, when extracting character features, the generated prompt can be used to prompt the LLM model to extract character features, and the prompt can be input into the LLM model to extract character features (e.g., extract character features with distinguishing features). For example, when using a generative pre-trained transformer (gpt-4o) as the LLM model, the prompt can be as follows:
Objective
[0034]
Extraction requirements
[0035] 2. To make each character distinct, the information between characters should be different, especially in terms of hair, eyes, and clothing. Each character should have a unique feature to improve their recognizability, especially for women of the same age or men of the same age. Hair color and style should be used to distinguish them.
[0036] 3. You need to determine the era of the story in the article. Do not use hairstyles, clothing, or external features that do not match the era of the article.
[0037]
Extraction content
[0038] Character code format: Character (1, 2, 3...) + name; Format of the external appearance information of the characters: composed of refined, specific, and detailed adjectives and nouns, with certain differences between each character, which can be reflected in hair and clothing. The external information of the characters includes the following: 1. Era: (ancient China, modern city, etc.).
[0039] 2. Role: (student, president, emperor, doctor, nurse, firefighter, teacher, etc. If not mentioned in the text, remove this indicator and do not show it in the answer. Pay attention to the distinction between different eras).
[0040] 3. Identity: Please choose strictly according to the following labels 【1Infant girl (female infant); 1Infant boy (male infant); 1child girl (female child); 1child boy (male child); 1young girl (female youth); 1young boy (male youth); 1elderly woman (female elderly woman); 1elderly woman (male elderly woman)】.
[0041] 4. Hair (hair color + hairstyle; if not mentioned in the text, infer and guess based on the novel. Pay attention to the differences in hairstyle between characters, such as the difference in hair color or hairstyle between the female protagonist and the female supporting character. For example, the female protagonist has long pink hair, and the female supporting character has short black hair. Colors such as black, pink, brown, yellow, red, and silver can be used to make the characters unique and easily identifiable).
[0042] 5. Facial features (eyes, eyebrows, nose, mouth, face shape; if not mentioned in the text, infer and guess based on the novel. Pay attention to the differences in eye color between characters, such as red eyes, green eyes, and blue eyes).
[0043] 6. External characteristics (wearing glasses, sitting in a wheelchair, having facial scars, etc. If not mentioned in the text, remove this indicator and do not show it in the answer. If there are external characteristics in the middle, they will not appear).
[0044] 7. Personality (cheerful and lively, gloomy, depressed, etc. If not mentioned in the text, infer and guess based on the novel. You can use some novel character settings to describe, such as a domineering president or a heartless and vicious character).
[0045] 8. Clothing (luxurious dress, suit, school uniform, shirt, jeans, etc. If not mentioned in the text, infer and guess based on the novel).
[0046]
Output Format
[0047] Reference output format: { "no": "one", "character": "mulberry branch" "era": "modern city" "role": "student" "character_type": "1young girl", "hair": "long, straight pink hair" "facial_features": { "eyes": "blue pupils" "eyebrows": "sword-like eyebrows, heroic spirit" "nose": "Upright and graceful, with simple and smooth lines" "mouth": "small mouth" "face_shape": "V-shaped face" }, "external_features": "wears glasses" "personality": "sunny and cheerful" Clothing: White shirt, jeans, pink earrings }
[0048] Further, after inputting the above prompt word into the LLM model, the initial text can be input into the LLM model again, and the LLM model can extract the character features of the initial text according to the requirements of the above prompt word, obtain the character information contained in the initial text, and output the character information in the required format. Further, the character information can be obtained by reading the output of the LLM model.
[0049] Further, the content abridged text refers to the reduced content of the initial text. Specifically, since the novel text is often described in detail, directly videoing the initial text may result in a large amount of work, requiring the generation of a large number of images and voice, thereby reducing the video generation efficiency. In order to reduce the workload of video generation without affecting the original meaning of the initial text, the initial text can be reduced, and only important plots and scenes can be reflected in the video. Therefore, the embodiment of the present application can obtain the content abridged text based on the reduced initial text. Specifically, the number of words of the content abridged text can be set according to actual needs, and the embodiment of the present application does not limit this.
[0050] Specifically, the above content extraction can be realized by a pre-trained text abridgment model, or can also be realized by an LLM model, and the embodiment of the present application does not limit this. Specifically, when realizing content extraction, the generated prompt word can be used to prompt the LLM model to extract the concise text, and the content extraction requirements (for example: extracting the plot content with attraction degree to generate the concise text) are input to the LLM model through the prompt word. Illustratively, in the case of using gpt-4o as the LLM model, the prompt word can be input as follows: "You are a professional novel tweet writer. Now I will give you a novel content, and you need to convert this novel content into a blockbuster novel tweet script.
[0051] The tweet script is divided into two structures:
plot hook
tweet script
[0052]
Plot hook
[0053]
Tweet script
[0054] 1. I will give you
content focus
[0055] 2. I will provide you with
blockbuster plot hook reference
blockbuster tweet script reference
[0056] 3. I will provide you with the
output requirements
output format
[0057] 4. I will give you the content of the novel, and you need to process the novel content according to the above requirements.
[0058]
Content focus
[0059] 2. Highlight the protagonist's identity and the complex environment they are in, such as the situation of not being accepted at their parents' home, to arouse readers' curiosity and concern about the protagonist's fate.
[0060] First-person perspective and narration: Using the first-person perspective increases the sense of immersion in the story, making it easier for readers to empathize. First-person, need to have a sense of net, attract users. At the same time, through the protagonist's inner monologue and the contrast with others, the impact of the plot is amplified.
[0061] Example: They all call me the "King of Park Fishing", because I always squat and watch the earthworms turn over while the lawn mower roars. Only I know that the "weeds" I delay pruning are actually a rare wildflower waiting for its once-in-a-decade blooming and pollination.
[0062] Unconventional occupation or identity, distinctive character, attractive setting: Use one or two words that are full of characteristics to shape the protagonist's identity, so that readers can quickly form an impression and emotional resonance. Use unexpected or special occupations or identities to set characters, such as interstellar archaeologists, ecological protectors, etc. These unique settings themselves can attract attention.
[0063] Unconventional, reversal: The subversion of common sense, expectations or morality, or unexpected reversal is particularly eye-catching.
[0064] Example: He has been sweeping leaves for ten years, and everyone thinks he is just a janitor, until an alien species invades and the ecosystem collapses, and people discover that he has been sweeping away "wrong" things in the park, not garbage.
[0065] Element 2: Clearly explain key information, and use strong emotional and suspenseful expressions.
[0066] Emphasize the emotional transformation of the main character: Use strong emotional expressions like "I never thought" and "I will never allow" to quickly capture the reader's heart and make them curious about why the main character has such a strong attitude.
[0067] Strong plot conflict: Strong conflict, directly presenting intense conflict or contradiction, making people want to know more. Such as group conflict, personal struggle, etc.
[0068] Sensibility and emotional connection, intense emotional contrast: There are a lot of emotional expressions in the hook text, so that readers can resonate or be curious on the sensory level with little information. In the description, strong emotional contrast can quickly capture the reader's heart.
[0069] Example: I have been pretending to be dull for ten years, sweeping leaves every day, and listening to them laugh at me as a "snail in a tree hole." But today, when the shiny "Dream Flower Introduction Plan" landed on my desk, my hand shook for the first time - not out of fear, but out of anger. They have no idea that they have just opened the last "ecological lock" of this park.
[0070] Element three: reveal the subsequent plot suspense of "I", highlight "change", "reverse", "suspense", "contrast".
[0071] Reverse and suspense: Create a tense and suspenseful atmosphere to attract readers to want to know more about the subsequent development. The text often reverses and suspends, piques the reader's curiosity, and makes them unable to resist wanting to know what will happen next.
[0072] Various hooks: The ending usually leaves obvious plot hooks, making readers deeply interested in the story's development, forcing them to continue reading.
[0073] Example: When they set up the celebration banquet for introducing "Dream Flower" under the banyan tree, I silently activated the plan I had buried ten years ago. Now, they will finally see that this place, which I have carefully disguised as a park, begins to shed its gentle green and reveal the more authentic ecology beneath.
[0074] By using these elements comprehensively, these titles and opening texts can quickly capture the reader's attention and make them interested in the story, thus becoming a hit.
[0075]
Hit plot hook reference
[0076] [Reference Script for Viral Tweet Posts] Copywriting paraphrasing 1: My work manual states on page 37: If a single plant grows abnormally (defined as exceeding the natural average by 300%) within 24 hours, it should be reported immediately and the area should be isolated.
[0077] The foxtail grass in front of me is now visibly growing its seventh stem—it has grown 28 centimeters overnight, exceeding the standard by 478%.
[0078] I crouched down and, just like I had done with the more than four hundred "excessive standards" cases over the past twelve years, gently touched its spike with my gloved fingertips. The veins on the inner side of the leaves were gleaming with an unnatural metallic sheen.
[0079] "Be careful." I lowered my voice, my fingertip sliding down the stem until I found the tiny implanted chip three centimeters below the soil line—again, the "New World Ecology" company's logo. This was the third time this month.
[0080] The blades of grass obediently curled up, their height dropping back to fifteen centimeters. Just as it disguised itself as an ordinary weed, slow but clear applause rang out behind me.
[0081] Output Format The output format is a compressed, formatted XML data structure string, which does not require additional text interpretation. The field names in the XML are: plot_hook (plot hook) and plot_script (tweet script).
[0082] Refer to the XML structure: <result> <script><plot_hook>< / plot_hook><plot_script>< / plot_script>< / script> < / result> .
[0083] Furthermore, after inputting the above-mentioned content extraction prompts into the LLM model, the initial text can be input into the LLM model, and then the LLM model can output the corresponding reduced text according to the requirements of the above-mentioned content extraction prompts. In this embodiment of the invention, the reduced text output by the LLM model can be used as the content abbreviation text.
[0084] Furthermore, after obtaining the character information and the abbreviated text of the content, the abbreviated text of the content can be split into multiple sub-texts based on the character information as the storyboard to be generated.
[0085] Specifically, the above-mentioned splitting operation can be implemented based on plot or scene transitions in the abbreviated text, with one scene corresponding to one plot or scene. Furthermore, this splitting operation can be implemented using a large language model.
[0086] The large language model refers to an LLM model. Specifically, a split prompt can be input to the LLM model, which can be used to prompt the LLM model to split the content summary text based on the character information, and the split prompt can be used to input a shot extraction requirement to the LLM model (for example, extract a script text with a word count range of 25-100 words as a description text of a shot, and the different shots have differences, and the characters in different shots have certain differences). Correspondingly, after the split prompt is input to the LLM model, the character information and the content summary text can be input to the LLM model, so that the LLM model splits the content summary text into description texts corresponding to different shots according to the split prompt and the character information. Further, to ensure the effectiveness of the obtained description texts, the description text corresponding to each to-be-generated shot can include at least one character in the character information and a plot description content related to the character. The plot description content can include character expressions, character actions, environment descriptions, composition, and scene descriptions.
[0087] In the embodiment of the application, the content summary text is obtained by content extraction on the initial text, and the content summary text is split based on the character information, which can reduce the workload of video generation to a certain extent while ensuring that the original meaning of the initial text is not changed, thereby ensuring the video generation efficiency. Meanwhile, the large language model is used to split the text, which can further reduce the degree of human participation in the video generation process and reduce the labor cost.
[0088] Optionally, in the operation of generating a corresponding shot image for each to-be-generated shot based on the description text corresponding to each to-be-generated shot in step 102, the embodiment of the application can specifically include: S31, from the description text corresponding to each to-be-generated shot, extracting a feature text including at least a shot character, a character action, and a shot scene of the to-be-generated shot, to obtain a prompt for each to-be-generated shot.
[0089] S32, inputting the prompt for each to-be-generated shot into the image generation model to obtain a shot image corresponding to each to-be-generated shot output by the image generation model.
[0090] The above feature text refers to text that can represent the plot features of the generated shot list, including at least the shot list characters, character actions, and shot list scenes of the generated shot list. The above shot list character refers to the role contained in the generated shot list, the above character action refers to the posture and expression of the shot list character in the generated shot list, and the above shot list scene refers to the scene and scene elements of the generated shot list, such as indoor or outdoor, outdoor park or square, indoor bed and curtain elements, etc.
[0091] Specifically, the above feature text can be extracted from the description text corresponding to the generated shot list based on the LLM model. The description text corresponding to each generated shot list can be input into the LLM model, and the LLM model can obtain the feature text of the generated shot list based on the character information identified from the initial text.
[0092] Specifically, the LLM model can be input with prompt words to prompt the LLM model to extract feature text from the input description text. Further, the LLM model can also be input with feature extraction requirements, for example, the extraction requirements can be the content (characters, scene descriptions, etc.) contained in each feature text, and the extraction requirements can also include the word limit of the feature text and the output format, etc. Specifically, the output format can be set to realize the splicing of each feature text. For example, the following feature extraction prompt words can be input to the LLM model, which can include extraction content, output format, etc.:
Extraction content
Character information
Character information
[0093] Extract character actions and expressions: Extract the actions and expressions of each character appearing in each shot according to
Script text
[0094] Expression: Describe the specific exaggerated and simple expressions of the appearing characters. Do not have complex expressions. Reference: laughing, crying, sad, angry, fear, etc. Action: Describe the actions of the appearing characters in the shot, ensuring that the actions are specific and explicit, and have a certain sense of picture. If not mentioned in the text, please make reasonable assumptions and guesses based on the plot. Reference: sitting and reading, standing at the door, squatting and crying, etc.
[0095] Extract the picture scene Extract the picture scene of each shot according to
Script text
[0096] 2. Scene: Describe the specific scene environment and actions. You can infer from the context. Information must be given and cannot be left blank. For example: reading in the classroom, standing at the door, sitting on the sofa.
[0097] 3. Surrounding objects: Describe the surrounding environment. You can infer from the context. Information must be given and cannot be shown as none. Examples: trees and blue sky, shops and carriages, tables, forest, elaborate decorations, lamps.
[0098] 4. Composition and perspective: Describe the composition and perspective of the picture to achieve a good effect. Try to ensure diversity and avoid full-body compositions. Refer to: central composition, rule of thirds, frontal perspective, and side perspective.
[0099] [Task Results] The string uses a compressed, formatted XML data structure and requires no additional text explanation. The field names in the XML are: plot (storyboard script), character (character in the storyboard), action (body movements of the character), expression (facial expression of the character), and scene (scene of the storyboard).
[0100] Refer to the XML structure: <result> <plot>I am the eldest daughter of the Sang family, Sang Ju.< / plot> <character>Sang Ju< / character> <action>Reading in front of the desk< / action> <expression>Laughing< / expression> <scene>Indoor, classroom, students in the classroom, desks and chairs, blackboard, three-quarter composition, side view< / scene> < / result> .
[0101] Based on this, the prompts generated for each storyboard to be generated through the above steps can largely express the main plot content of the storyboard to be generated. Therefore, embodiments of the present invention can generate corresponding storyboard images for each storyboard to be generated based on the prompts. Specifically, a preset image generation model can be used to generate storyboard images. The prompts can be input into the preset image generation model, and the image generation model can generate storyboard images that conform to the content described by the prompts. Optionally, the above image generation model can be an SD model, or other pre-trained image generation models; embodiments of the present invention do not limit this.
[0102] Optionally, the style of the storyboard images can be pre-trained in the image generation model. It can be a comic style, a realistic style, etc., and can be set according to the actual video generation needs. This embodiment of the invention does not limit this.
[0103] In this embodiment of the invention, by extracting feature text and generating prompt words based on the feature text, information that can express the plot content of the storyboard to be generated can be obtained. Then, the prompt words can be used to generate corresponding storyboard images for the storyboard to be generated, ensuring both image generation efficiency and accuracy of the generated images.
[0104] Optionally, the operation of inputting the prompt words of each of the to-be-generated storyboards into the image generation model in step S32 to obtain the storyboard images corresponding to each of the to-be-generated storyboards output by the image generation model can specifically include the following steps. S41, generating a weight coefficient for each prompt word of each of the to-be-generated storyboards.
[0105] S42, inputting the prompt words of each of the to-be-generated storyboards and the weight coefficients into a preset image generation model respectively, so that the image generation model adjusts the intensity of the image elements corresponding to each of the prompt words based on the weight coefficients of the prompt words, to obtain the storyboard images corresponding to each of the to-be-generated storyboards; the intensity is used to represent the proportion and prominence of the image elements.
[0106] The weight coefficient is used to represent the importance of each prompt word in the storyboard, and the value range of the weight coefficient can be (0, 1). Accordingly, the weight coefficient of 0 indicates that the importance of the prompt word in the storyboard is the smallest, and the intensity of the image element corresponding to the prompt word in the generated image is the smallest, and the proportion is the smallest. Accordingly, the weight coefficient of 1 indicates that the importance of the prompt word in the storyboard is the largest, and the intensity of the image element corresponding to the prompt word in the generated image is the largest, and the proportion is the largest. The weight coefficient of each prompt word can be generated according to actual video generation requirements. Specifically, if the video generation requirement is to emphasize the text character, a higher weight coefficient can be generated for the prompt word related to the character. Accordingly, if the video generation requirement is to emphasize the text scene, a higher weight coefficient can be generated for the prompt word related to the scene.
[0107] Further, the prompt words and the weight coefficients of each of the to-be-generated storyboards can be input into a preset image generation model, and the image generation model can generate corresponding images that meet the prompt words and the corresponding weight coefficients based on the prompt words and the weight coefficients. The output data of the image generation model can be used as the storyboard images corresponding to each of the to-be-generated storyboards. Specifically, the image generation model can generate image elements corresponding to each of the prompt words based on the weight coefficients of the prompt words, and can adjust the intensity of the image elements corresponding to each of the prompt words, so that the generated storyboard images meet the prompt words and the corresponding weight coefficients. Specifically, the intensity refers to the proportion and prominence of the image elements. For the prompt word with a larger weight coefficient, the proportion and prominence of the image element corresponding to the prompt word in the storyboard image are larger, which can more easily attract the attention of the user. Accordingly, for the prompt word with a smaller weight coefficient, the proportion and prominence of the image element corresponding to the prompt word in the storyboard image are smaller, which can make the image element not easy to be noticed.
[0108] Optionally, for each to-be-generated sub-shot, the prompt can be assembled based on the character information and the text content included in the to-be-generated sub-shot. For example, for one of the to-be-generated sub-shots, the corresponding feature text can be as follows: <plot>I witnessed him being harassed at a banquet, and to protect him, I chose to intervene.< / plot> <character>Jiang Ci worries< / character> <action>Standing by the table and picking up a wine cup< / action> <expression>Decisive< / expression> <scene>Indoor, banquet, table full of wine cups and dishes, three-quarter composition, side view< / scene> Furthermore, through the sub-shot character "Jiang Ci忧" in the above feature text, the feature information of this character can be found from the character information obtained in the above steps: {"character": "Jiang Ci忧", "era": "modern city", "character_type": "1 young girl", "hair": "brown curly hair", "facial_features": { "eyes": "green pupils", "eyebrows": "willow-leaf eyebrows", "nose": "small and delicate nose", "mouth": "red lips", "face_shape": "oval face" }, "personality": "tough and decisive", "clothing": "black dress" } Furthermore, information such as facial features, hairstyle, clothing, and personality can be extracted from it, and then scene information and weight coefficients can be added based on the character feature information. Furthermore, the following prompt for generating images can be composed: (1 young girl, brown curly hair: 1.2) - eyes: green pupils - eyebrows: willow-leaf eyebrows - nose: small and delicate nose - mouth: red lips - face shape: oval face (black dress: 1.2) (standing beside the table, picking up a wine glass, decisive: 1.0) (Indoors, banquet, table full of wine and dishes, three composition, side view angle: 1.0).
[0109] Further, the prompt words and the weight coefficients can be input into a preset image generation model, and output data of the image generation model is read to obtain the corresponding storyboard image.
[0110] In the embodiments of the present application, the content of the generated storyboard image can be adjusted by configuring the weight coefficients for each prompt word, and the flexibility of video generation can be improved.
[0111] Optionally, the operation of text-to-speech conversion processing each description text in step 103 to obtain the storyboard speech corresponding to each to-be-generated storyboard can specifically include: S51, after normalizing the description text corresponding to each to-be-generated storyboard, the normalized description text is mapped to acoustic features.
[0112] S52, the acoustic features are converted into speech waveforms to obtain the storyboard speech corresponding to each to-be-generated storyboard.
[0113] The description text corresponding to the to-be-generated storyboard often contains scene, plot information and character dialogue information. In order to make the audio data in the generated video consistent with the image picture, the description text corresponding to the to-be-generated storyboard can be directly converted into an audio format, so as to obtain the storyboard speech corresponding to each to-be-generated storyboard.
[0114] Specifically, the format conversion operation can first normalize the description text. The normalization can standardize the description text. Further, the normalized description text can be mapped to acoustic features through a preset acoustic modeling algorithm or acoustic modeling model. For example, the description text can be converted into spectral features by using a Hidden Markov Model (HMM). Further, the acoustic features can be converted into playable speech waveforms by using a vocoder, so as to obtain the storyboard speech.
[0115] Specifically, the operation can also be implemented by using a TTS tool or a TTS model. The text corresponding to each to-be-generated storyboard can be input into the TTS tool or the TTS model, and then the TTS tool or the TTS model is used to obtain and convert the acoustic features, so as to obtain the storyboard speech.
[0116] In the embodiments of the present application, the playable storyboard speech can be obtained according to the description text, so as to ensure the diversity of novel promotion. At the same time, the audio data in the generated video can be consistent with the image picture, so as to ensure the effect of video generation.
[0117] Optionally, the operation of generating a video for the initial text based on the split shot images corresponding to each of the to-be-generated split shots and the split shot audios in step 104 can specifically include the following steps. S61, determining at least one to-be-selected template as a target template from a preset video template library.
[0118] S62, splicing each of the split shot images and each of the split shot audios according to a target order to obtain a video stream and an audio stream; the target order is an arrangement order of each of the to-be-generated split shots.
[0119] S63, setting a playing attribute for the video stream and / or the audio stream according to a video attribute contained in the target template, and merging the video stream and the audio stream to obtain a video corresponding to the initial text.
[0120] The video attribute can include a video accompaniment and / or a video animation, and the video accompaniment and / or the video animation can improve the richness and attractiveness of the generated video, thereby further attracting readers. The preset video template library can be pre-set, and can include a plurality of to-be-selected templates. Different to-be-selected templates can include different video attributes, and accordingly, the effects of the videos generated according to different video templates are different. On this basis, at least one to-be-selected template can be selected as a target template from the preset video template library.
[0121] Specifically, the target template can be randomly selected, or can be selected according to a selection instruction, or can be at least one to-be-selected template with a higher frequency of use in the preset video template library, or can be at least one to-be-selected template with better timeliness in the preset video template library. The selection manner of the target template can be set by itself, and the embodiments of the present application do not limit this.
[0122] Further, the number of the selected target templates can be one or a plurality, which can be set according to actual video generation requirements, and the embodiments of the present application do not limit this. Optionally, the above steps can also be implemented by a cloud rendering module, and the split shot images and the split shot audios can be encoded into a final video according to a preset template.
[0123] The target order refers to an arrangement order of each of the to-be-generated split shots, that is, an arrangement order of the description text corresponding to each of the to-be-generated split shots in the initial text. By splicing the images and the audios according to the target order, the accuracy of video generation can be ensured.
[0124] Further, the video attribute can include a video music and / or a video animation, wherein the video music refers to a background music of the video, and the video animation refers to a switching animation of images in the video, and on this basis, the embodiment of the present application can set a playing attribute for the video stream and / or the audio stream according to the video attribute contained in the target template. Specifically, the volume size, pause interval and other playing attributes can be set for the audio stream according to the video music in the video attribute, and at the same time, the video music can be added to the audio stream. Correspondingly, the switching effect can be set for the video stream according to the video animation in the video attribute, so that the switching effect is consistent with the video animation of the video attribute. Further, the video corresponding to the initial text can be obtained by merging the video stream and the audio stream.
[0125] Specifically, the merging operation can be merging the video stream track and the audio stream track. Alternatively, in the process of merging, the video stream and the audio stream can be further time axis aligned to ensure the audio-visual consistency. Alternatively, the video attribute can further include attributes such as resolution, size, etc., and accordingly, the resolution and size of the video stream can be further set according to the video attribute.
[0126] In the embodiment of the present application, at least one selected template is determined as a target template from a preset video template library; each of the shot images and each of the shot voices is spliced according to a target order to obtain a video stream and an audio stream; the target order is an arrangement order of each of the to-be-generated shots; playing attributes are set for the video stream and / or the audio stream according to video attributes contained in the target template, and the video stream and the audio stream are merged to obtain a video corresponding to the initial text. In this way, different videos with different effects can be generated for the initial text by selecting different target templates, and the flexibility and diversity of video generation are improved.
[0127] Alternatively, the embodiment of the present application can also generate subtitles for each of the to-be-generated shots, and add the text contained in each of the to-be-generated shots as subtitles to the video, to further improve the richness of the generated video.
[0128] Alternatively, the video attribute can further include a playing speed.
[0129] Alternatively, Figure 2 is a structural schematic diagram of a video generation system provided by the embodiment of the present application, as Figure 2 shown, the video generation system can include a novel content input module, a character extraction module, a content rewriting module, a shot splitting module, an SD image generation module, a TTS voice generation module, an assembly module and a cloud rendering module.
[0130] Specifically, the novel content input module is configured to receive a novel text as an initial text, and the character extraction module is configured to obtain character information from the initial text. The content rewriting module is configured to output a shortened text by shortening the initial text. The split shot module is configured to split the shortened text into a plurality of to-be-generated shots based on the character information. Further, the SD image generation module and the TTS speech generation module are configured to generate a shot image and a shot speech for the to-be-generated shots, respectively. Further, the assembling module is configured to assemble the shot image and the shot speech to obtain a video. Further, the cloud rendering module is configured to encode the shot image and the shot speech according to a preset template to obtain a final video.
[0131] Optionally, Figure 3 is a flowchart of another video generation method provided by an embodiment of the present application, as shown in Figure 3 The LLM model can be used to perform content shortening on the initial text to obtain a content shortened text, and to perform role feature extraction (character feature extraction) on the initial text to obtain character information. Thus, the LLM model can be used to generate a description text corresponding to the to-be-generated shots according to the content shortened text and the character information, and the description text can include scenario description and lines, etc. Further, prompt words can be generated for different to-be-generated shots, and the optimization of image generation can be realized by configuring weight coefficients for different prompt words. The prompt words and the weight coefficients can be input into the SD model to generate pictures for different shots. Meanwhile, the TTS can be used to generate speech based on the description text (including lines) of the shots. The shot pictures and the speech can be input into the cloud rendering module for encoding to obtain a video corresponding to the initial text.
[0132] It should be noted that in actual application scenarios, the above video generation method can be realized by cooperation of a front end and a back end. The front end can receive user input, and the back end can be used to realize video generation based on the user input. Meanwhile, the video generation process can be divided into three steps of script processing, picture generation, and video synthesis.
[0133] Optionally, Figure 4 is a flowchart of script processing provided by an embodiment of the present application, as shown in Figure 4As shown, the novel information refers to the novel title, novel content, and the number of script washing received through the front end. Among them, the novel title refers to the name of the novel, the novel content corresponds to the initial text of the novel, and the number of script washing refers to the number of videos required to be generated for the novel. Further, after receiving the initial text, the backend can perform character extraction on the novel to generate character information (character prompt). And the novel can be washed to get the hook script and the script script. Among them, the hook script and the script script refer to the content abbreviated text obtained by abbreviating the initial text, and the hook script corresponds to the abstract text of the initial text. Further, the novel can also be split into shots to obtain a plurality of description texts corresponding to the generated shots.
[0134] Specifically, the characters in the novel content can be extracted by the LLM large model. In this process, AI (Artificial Intelligence) prompt words can be input to assist the character extraction process of the LLM large model. At the same time, the script washing operation can be realized by the gpt large model. Similarly, AI prompt words can also be input to assist the script washing process. Further, through script washing, a script washing script (content abbreviated text) can be obtained, which includes a hook script and a script script. Through character extraction, character information in the novel can be obtained, which can include character coding or character name and character image characteristics (identity, hair, facial features, appearance characteristics, and clothing).
[0135] Further, based on the obtained abbreviated text and character information, a picture generation operation can be performed.
[0136] Optionally, Figure 5 is a flowchart of a picture generation process provided by an embodiment of the present application, as Figure 5 As shown, the front end can receive the drawing style required by the user, which can include picture style and interrelated configuration character style. The picture style can be trained by the checkpoint model, and the character style can be trained by the Low-Rank Adaptation (lora) model and the style fine-tuning model. Further, the picture style can include animation and ancient style. Accordingly, when the picture style is animation, the modern male lora and the modern female lora can be trained accordingly. Accordingly, when the picture style is ancient style, the ancient male lora and the ancient female lora can be trained accordingly.
[0137] Further, after obtaining the shot list based on the script processing, the feature information (feature text) in the shot list can be obtained, the feature information includes the character matched by the shot list, the action of each shot of the character, the scene, and the shot. The feature information is input into the LLM model, and the AI prompt word is used to assist the LLM model, so that the LLM model outputs the script of the split shot and the keywords (prompt words) of each shot based on the feature information. Further, the shot keywords can include the character matched by each shot, the scene, and the shot. At the same time, the user can also modify the shot keywords according to the actual needs.
[0138] Further, based on the keywords of each shot, the keyword text of each shot can be formatted by code, which is processed into the prompt word of the image generation model (for example: comfy ui). Optionally, during the format processing, it can be implemented based on the label library of the image generation model. Specifically, the label library can include the prompt words preset by the image generation model, and the label library can include different types of prompt words, for example, it can include multiple prompt words corresponding to the identity of the character, multiple prompt words corresponding to the hairstyle of the character, and multiple prompt words corresponding to the age of the character. On this basis, the keywords of each to-be-generated shot can be selected from the prompt words of each type in the label library to achieve the formatting processing.
[0139] Further, based on the obtained prompt words and shot text, video synthesis operations can be performed.
[0140] Optionally, Figure 6 is a flowchart of a video synthesis process provided by an embodiment of the present application, as shown in Figure 6 First, the video properties such as commentary dubbing, speed processing, background music (bgm), and film ratio can be configured in the front end. Specifically, the commentary dubbing can be automatically matched by selecting the general dubbing from the preset dubbing library. The speed processing can be defaulted to 1.2 times, but any speed between 1 and 1.5 times can also be set, and the present embodiment does not limit this. The bgm music can also be automatically matched according to the novel classification from the preset music library, and the mapping label can be performed according to the novel classification and the music style, so that different labels of novels correspond to different labels of music, thereby realizing automatic matching of music according to novel classification. The film ratio refers to the picture ratio, which can be pre-set to a fixed value to output the video film according to the fixed film ratio. The film ratio can be set to 16:9, 4:3, or 1:1, and can be set according to the actual promotion needs, and the present embodiment does not limit this.
[0141] Further, the script text generated by the picture can be audited and the auditing words can be removed. Specifically, many video platforms currently have an auditing mechanism. If a video uploaded has risky auditing words, the video will be removed. To avoid this situation, the embodiment of the present application can remove the auditing words in the script text before generating the script voice and the letter text. After removing the auditing words, the foreign language text in the text is translated to obtain the final script subtitle text.
[0142] Meanwhile, the script prompt obtained after the picture is generated can be realized by an image generation model (for example, a comfyui text generation model). Among them, script 1 and script n represent different to-be-generated scripts, and correspondingly, AI drawing Figure 1 、 2 represents the script image corresponding to the to-be-generated script. Specifically, one script can correspond to one or more script images.
[0143] Further, after obtaining the script subtitle text and the script image, video synthesis can be performed. Specifically, the script subtitle text can obtain the subtitle lines and the subtitle commentary dubbing, and the script image can obtain the picture under the script. One to-be-generated script corresponds to a group of subtitle lines, subtitle commentary dubbing, and picture. Further, the subtitle lines, subtitle commentary dubbing, and picture corresponding to the to-be-generated script can be encoded by the video properties configured in the front end to obtain a video containing bgm music, title, and drop version. Meanwhile, the video properties can also contain picture moving mode, that is, video animation, which is used to indicate the switching effect between different images, such as fade-in and fade-out, so that the encoding process can be performed according to the picture moving mode in the properties, so that the generated video meets the video properties.
[0144] Figure 7 is a structural diagram of a video generation device provided by the embodiment of the present application, as shown in Figure 7 The device 20 can include: A first acquisition module 201 is configured to acquire description texts corresponding to a plurality of to-be-generated scripts based on an initial text. A first generation module 202 is configured to generate script images corresponding to the to-be-generated scripts by a preset image generation model based on the description texts corresponding to the to-be-generated scripts. A synthesis module 203 is configured to perform text-to-speech processing on the description texts to obtain script voices corresponding to the to-be-generated scripts. A second generation module 204 is configured to generate a video for the initial text based on the script images corresponding to the to-be-generated scripts and the script voices.
[0145] Optionally, the first acquisition module includes: a first extraction submodule configured to extract character features and content from the initial text, to obtain character information contained in the initial text and a content summary text corresponding to the initial text; a splitting submodule configured to input a splitting prompt word to a large language model and input the character information and the content summary text into the large language model, the splitting prompt word being used to instruct the large language model to split the content summary text into multiple description texts corresponding to to-be-generated split shots based on the character information, each of the description texts corresponding to the to-be-generated split shots containing at least one character in the character information and a plot description related to the character.
[0146] Optionally, the first generation module comprises: a second extraction submodule configured to extract a feature text including at least a split shot character, a character action and a split shot scene of each of the to-be-generated split shots from each of the description texts corresponding to the to-be-generated split shots, to obtain a prompt word of each of the to-be-generated split shots; a first input submodule configured to input the prompt word of each of the to-be-generated split shots into the image generation model, to obtain a split shot image corresponding to each of the to-be-generated split shots output by the image generation model.
[0147] Optionally, the first input submodule is specifically configured to: generate a weight coefficient for each of the prompt words of each of the to-be-generated split shots; input the prompt word of each of the to-be-generated split shots and the weight coefficient into a preset image generation model, so that the image generation model adjusts intensity of an image element corresponding to each of the prompt words based on the weight coefficient of each of the prompt words, to obtain the split shot image corresponding to each of the to-be-generated split shots; the intensity is used to represent a proportion and a degree of prominence of the image element.
[0148] Optionally, the synthesis module further comprises: a normalization submodule configured to perform normalization processing on each of the description texts corresponding to the to-be-generated split shots, and map the normalized description text into an acoustic feature; a conversion submodule configured to convert the acoustic feature into a voice waveform, to obtain a split shot voice corresponding to each of the to-be-generated split shots.
[0149] Optionally, the second generation module comprises: a determination submodule configured to determine at least one to-be-selected template as a target template from a preset video template library; The splicing sub-module is configured to splice the split shot images and the split shot audios according to a target order to obtain a video stream and an audio stream, respectively, wherein the target order is an arrangement order of the to-be-generated split shots. The setting sub-module is configured to set a playing attribute for the video stream and / or the audio stream according to a video attribute contained in the target template, and combine the video stream and the audio stream to obtain a video corresponding to the initial text.
[0150] Optionally, the video attribute includes a video music and / or a video animation.
[0151] To sum up, the embodiment of the present application can obtain a plurality of description texts of to-be-generated split shots according to an initial text, and generate corresponding split shot images and split shot audios for each to-be-generated split shot through image generation technology and text-to-speech conversion, so as to divide the text into different split shots and realize imageization of each split shot. Further, a video can be generated through the split shot images and the split shot audios of each to-be-generated split shot, so as to improve the automation degree of text videoization, reduce manual participation, and greatly reduce the labor cost of video generation. At the same time, since the video can be automatically generated to a certain extent according to the initial text, the video generation efficiency is greatly improved.
[0152] The present application also provides an electronic device, referring to Figure 8 , comprising a processor 301, a memory 302, and a computer program 3021 stored in the memory and executable on the processor, wherein the processor implements the video generation method of the foregoing embodiments when executing the program.
[0153] The present application also provides a readable storage medium, when the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute the video generation method of the foregoing embodiments.
[0154] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts are described in the part of the method embodiment.
[0155] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with these teachings, with the structure for a variety of such systems will be apparent from the description above. In addition, the present application is not intended to be limited to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the teachings of the application described herein, and any references below to specific languages are provided for disclosure of enablement only.
[0156] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0157] Similarly, it is to be understood that the embodiments of the present application can be adapted to practice one or more of the various inventive aspects in the description of the exemplary embodiments of the present application above, individual features of the present application are sometimes grouped together in a single embodiment, figure, or description of related features. However, the manner of
[0158] Those skilled in the art will appreciate that the modules in the apparatus of the embodiments can be adapted and placed in one or more apparatuses other than the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and further can be divided into more sub-modules or sub-units or sub-components. Any combination of all the features disclosed in the specification (including the accompanying claims, abstract and drawings), and any method or of the apparatuses so disclosed, can be made, except that at least some of such features and / or processes or units are mutually exclusive. Each feature disclosed in the specification (including the accompanying claims, abstract and drawings) can be replaced by alternative features serving the same, equivalent or a similar purpose, unless otherwise expressly stated.
[0159] Embodiments of the various components of the present application can be implemented in hardware, or as software modules running in one or more processors, or combinations thereof. Those skilled in the art will appreciate that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functionality of some or all of the components in the sequencing apparatus according to the present application. The present application can also be implemented as a program for executing part or all of the methods described herein on a device or apparatus. Such a program can be stored on a computer readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier medium, or in any other form.
[0160] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0161] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0163] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A video generation method, characterized in that, The method includes: Based on the initial text, obtain the descriptive text corresponding to multiple storyboards to be generated; Based on the description text corresponding to each of the storyboards to be generated, a corresponding storyboard image is generated for each of the storyboards to be generated using a preset image generation model; The descriptive texts are processed by text-to-speech conversion to obtain the speech of each storyboard to be generated. Based on the storyboard images corresponding to each of the storyboards to be generated and the storyboard audio, a video is generated for the initial text.
2. The method according to claim 1, characterized in that, The process of obtaining descriptive text corresponding to multiple storyboards to be generated based on the initial text includes: The initial text is subjected to character feature extraction and content extraction to obtain the character information contained in the initial text and the corresponding abbreviated text of the content; Input splitting prompt words into the large language model, and input the character information and the content abbreviated text into the large language model. The splitting prompt words are used to instruct the large language model to split the content abbreviated text into shots based on the character information to obtain descriptive texts corresponding to multiple shots to be generated. Each descriptive text corresponding to a shot to be generated contains at least one character from the character information and plot descriptions related to the character.
3. The method according to claim 1, characterized in that, The step of generating corresponding storyboard images for each of the storyboards to be generated based on the descriptive text corresponding to each of the storyboards to be generated, using a preset image generation model, includes: From the descriptive text corresponding to each of the storyboards to be generated, extract feature text that includes at least the storyboard characters, character actions, and storyboard scenes of the storyboards to be generated, and obtain the prompt words for each of the storyboards to be generated. The prompts for each of the storyboards to be generated are input into the image generation model to obtain the storyboard images corresponding to each of the storyboards to be generated, which are output by the image generation model.
4. The method according to claim 3, characterized in that, The step of inputting the prompt words for each of the storyboards to be generated into the image generation model to obtain the storyboard images corresponding to each of the storyboards to be generated output by the image generation model includes: Generate weight coefficients for each prompt word in each of the storyboards to be generated; The prompts for each storyboard to be generated and the weight coefficients are respectively input into a preset image generation model, so that the image generation model adjusts the intensity of the image elements corresponding to each prompt based on the weight coefficients of each prompt, thereby obtaining the storyboard image corresponding to each storyboard to be generated; the intensity is used to characterize the proportion and salience of the image elements.
5. The method according to claim 1, characterized in that, The step of performing text-to-speech conversion on each of the described texts to obtain the storyboard audio corresponding to each of the storyboards to be generated includes: After normalizing the descriptive text corresponding to each of the storyboards to be generated, the normalized descriptive text is mapped to acoustic features. The acoustic features are converted into speech waveforms to obtain the speech of each of the storyboards to be generated.
6. The method according to claim 1, characterized in that, The step of generating video from the initial text based on the storyboard images corresponding to each of the storyboards to be generated and the storyboard audio includes: Select at least one candidate template from the preset video template library as the target template; The video stream and audio stream are obtained by splicing the storyboard images and audio segments of each storyboard according to a target order; the target order is the arrangement order of the storyboards to be generated. According to the video attributes contained in the target template, set the playback attributes for the video stream and / or the audio stream, and merge the video stream and the audio stream to obtain the video corresponding to the initial text.
7. The method according to claim 6, characterized in that, The video attributes include background music and / or video animation.
8. A video generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire description texts corresponding to multiple storyboards to be generated based on the initial text; The first generation module is used to generate corresponding storyboard images for each of the storyboards to be generated based on the description text corresponding to each of the storyboards to be generated, using a preset image generation model. The synthesis module is used to perform text-to-speech conversion processing on each of the described texts to obtain the speech corresponding to each of the storyboards to be generated. The second generation module is used to generate a video for the initial text based on the storyboard images corresponding to each of the storyboards to be generated and the storyboard audio.
9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, implements the method as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method of any one of claims 1-7.