Video generation method and device based on large model, equipment and storage medium

By building a character knowledge graph and a large language model to generate story mirror pictures, the problems of cumbersome and low quality of manual operations in the existing technology are solved, efficient and automated video generation is achieved, and video quality and user experience are improved.

CN120264097APending Publication Date: 2025-07-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510346902.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing video-based technical solutions for literary novels require a lot of manual operation, the process is cumbersome, the efficiency is low, and it is easily affected by human judgment deviations, resulting in poor video quality and inability to accurately convey the original artistic conception and emotions.

Method used

By building a character knowledge graph, using a large language model (LLM) to extract role information and relationships, generate multi-frame storyboard pictures, and combine audio and video synthesis technology to automatically generate target videos.

Benefits of technology

Improve the efficiency and quality of video generation, ensure character consistency and plot consistency, and improve user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264097A_ABST
    Figure CN120264097A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device based on a large model, equipment and a storage medium, and relates to the field of artificial intelligence, in particular to the fields of large models, text maps and audios and videos. According to the specific implementation scheme, a to-be-converted text is obtained, and a knowledge graph of roles in the to-be-converted text is constructed; wherein the knowledge graph represents role information of roles in the text to be converted; according to the to-be-converted text and the knowledge graph of each role, determining a plurality of frames of split pictures; wherein the split picture represents a part of text content in the to-be-converted text; and generating a target video according to the multiple frames of split pictures. The automation level of the video generation process is improved, the labor cost is reduced, the video quality is improved, and the watching experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to large models, text-to-image, and audio-video fields in the field of artificial intelligence, and particularly relates to a video generation method, apparatus, device, and storage medium based on a large model. Background Art

[0002] With the development of artificial intelligence technology, the entertainment functions of mobile devices have been increasingly enhanced. Adapting literary works into film and television works has become a popular way of cross-media integration.

[0003] Currently existing text-to-novel video technology solutions usually require a large amount of manual operations. The video generation process is cumbersome and inefficient, and is easily affected by human judgment biases, resulting in low video quality and affecting the user's viewing experience. Summary of the Invention

[0004] The present disclosure provides a video generation method, apparatus, device, and storage medium based on a large model.

[0005] According to a first aspect of the present disclosure, there is provided a video generation method based on a large model, including:

[0006] Obtaining a text to be converted, and constructing a knowledge graph of the characters in the text to be converted; wherein, the knowledge graph represents the character information of the characters in the text to be converted;

[0007] Determining multiple frames of storyboard pictures according to the text to be converted and the knowledge graphs of the respective characters; wherein, the storyboard pictures represent partial text content in the text to be converted;

[0008] Generating a target video according to the multiple frames of storyboard pictures.

[0009] According to a second aspect of the present disclosure, there is provided a video generation apparatus based on a large model, including:

[0010] A graph construction unit, configured to obtain a text to be converted, and construct a knowledge graph of the characters in the text to be converted; wherein, the knowledge graph represents the character information of the characters in the text to be converted;

[0011] A picture determination unit, configured to determine multiple frames of storyboard pictures according to the text to be converted and the knowledge graphs of the respective characters; wherein, the storyboard pictures represent partial text content in the text to be converted;

[0012] A video generation unit, configured to generate a target video according to the multiple frames of storyboard pictures.

[0013] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor;

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect of the present disclosure.

[0017] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect of the present disclosure.

[0018] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of the method described in the first aspect of the present disclosure.

[0019] According to the technology of the present disclosure, the efficiency and quality of video generation are improved, and the viewing experience of users is enhanced.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0022] Figure 1 is a schematic flowchart of a method for generating a video based on a large model provided by an embodiment of the present disclosure;

[0023] Figure 2 is a schematic diagram of a knowledge graph of a character provided by an embodiment of the present disclosure;

[0024] Figure 3 is a schematic diagram of a character relationship graph provided by an embodiment of the present disclosure;

[0025] Figure 4 is a schematic flowchart of a method for generating a video based on a large model provided by an embodiment of the present disclosure;

[0026] Figure 5 is a schematic flowchart of a method for generating a video based on a large model provided by an embodiment of the present disclosure;

[0027] Figure 6 is a schematic diagram of the process of novel film and television adaptation provided by an embodiment of the present disclosure;

[0028] Figure 7It is a structural block diagram of a video generation device based on a large model provided according to an embodiment of the present disclosure;

[0029] Figure 8 It is a structural block diagram of a video generation device based on a large model provided according to an embodiment of the present disclosure;

[0030] Figure 9 is a block diagram of an electronic device for implementing the large model-based video generation method of an embodiment of the present disclosure;

[0031] Figure 10 It is a block diagram of an electronic device used to implement the large model-based video generation method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] With the increasing entertainment functions of smartphones and other mobile devices, especially the improvement of video playback capabilities, adapting literary works into AI film and television works has become a popular cross-media integration method.

[0034] The current technical solutions for converting literary novels into videos can be to manually classify the historical background of literary works, summarize the characters in literary works, adapt the original novels into AI (Artificial Intelligence) film and television scripts, then draw pictures through the literary graph model, manually correct deformed pictures, and finally synthesize the pictures into a video.

[0035] However, a large amount of manual participation not only makes the entire process cumbersome and inefficient, but is also susceptible to human judgment bias, which in turn affects the quality of the video and the accuracy of the content. The generated deformed images may cause the video screen to be distorted, damage the visual performance, and fail to accurately convey the artistic conception and emotions of the original text. This not only reduces the artistic expression of the video work, but may also have a long-term negative impact on the brand image of the producer. In addition, the static images output by the model lack sufficient expressiveness, resulting in poor quality of the generated video, which cannot fully show the rich emotions and details of the original text.

[0036] The present disclosure provides a large model-based video generation method, apparatus, device, and storage medium, which are applied to large models, text-to-image, and audio-video fields in the field of artificial intelligence to improve the efficiency and quality of video generation and enhance the viewing experience of users.

[0037] It should be noted that the model in this embodiment is not a model for a specific user and does not reflect the personal information of a specific user. It should be noted that the data in this embodiment comes from a public dataset.

[0038] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0039] To enable readers to more deeply understand the implementation principle of the present disclosure, the following Figures 1 - 10 is used to further refine the embodiments.

[0040] Figure 1 As shown in the flowchart of a large model-based video generation method provided according to an embodiment of the present disclosure, this method can be executed by a large model-based video generation apparatus. As Figure 1 shown, the method includes the following steps:

[0041] S101. Obtain the text to be converted and construct a knowledge graph of the characters in the text to be converted; wherein, the knowledge graph represents the character information of the characters in the text to be converted.

[0042] Exemplarily, the text to be converted is a literary work, for example, it can be a novel. The user can upload the original text of the text to be converted or crawl the literary work to be filmed online as the text to be converted. In this embodiment, when the text to be converted is filmed, the obtained video can be an animated video.

[0043] The text to be converted is a literary work containing plots, characters, environments, etc. After obtaining the text to be converted, the characters contained in the text to be converted can be extracted. For each character, a knowledge graph can be constructed. The knowledge graphs of each character can be independent of each other, and the knowledge graph can be in the form of a topological graph. In each knowledge graph, there can be multiple nodes, and there can be edge connections between different nodes. Each node can represent a piece of character information of the character, and the entire knowledge graph represents all the character information of the character. For each character, the character information of the character can be extracted from the text to be converted, and a knowledge graph can be constructed according to the extracted character information.

[0044] Semantic recognition processing can be performed on the text to be converted to determine the roles and role information from the text to be converted. For example, the powerful information extraction ability of the LLM (Large Language Model) can be utilized to accurately extract role information such as the role's main body, alias, body shape, face shape, hairstyle, clothing, etc. in the novel, as well as the detailed descriptions of the roles in the novel, so as to comprehensively and deeply understand the roles in the novel. In this embodiment, the model structure of the LLM is not specifically limited.

[0045] In this embodiment, the role information includes the main body information and the attribute information. The main body information represents the identifier of the role, and the attribute information represents the situation related to the role; a knowledge graph of the roles in the text to be converted is constructed, including: extracting the main body information and the attribute information of the roles from the text to be converted; connecting the nodes corresponding to each attribute information to the node corresponding to the main body information respectively to obtain the knowledge graph of the role.

[0046] Specifically, the role information may include the main body information and the attribute information of the role. The main body information can represent the identifier of the role, and each role corresponds to a unique main body information. For example, the main body information can be the name of the role. The attribute information can represent various information related to the role. For example, it may include the alias, hairstyle, body shape, clothing, accessories, gender, age, face shape, hair color, eye color, type and color of clothing, etc. of the role. The LLM can be used to extract the main body information and the attribute information of the role from the text to be converted. For example, the role name can be recognized from the text to be converted, and the information related to the role described after the role name is determined as the attribute information of the role.

[0047] For each role, the main body information is a node, and each attribute information is also a node. Connect the nodes corresponding to each attribute information to the node corresponding to the main body information with edges to obtain the knowledge graph of the role. The knowledge graphs of all roles constitute the role archive.

[0048] Figure 2 It is a schematic diagram of the knowledge graph of the role. Figure 2 The node corresponding to the main body information in it is the node of "Role A", and each node connected to the "Role A" node represents the attribute information of Role A.

[0049] The beneficial effect of such a setting is that the character consistency of AI film and television works is the cornerstone of being faithful to the original work, which involves the coherence of image and behavior as well as the reactions and interactions in multiple scenarios. Character coherence refers to the logical consistency of character personality and development throughout the plot. Extract different character information of the characters, center on the main information, construct a topological graph, and obtain a character archive library for the whole book dimension, which details the characteristics, living environment, personality traits, behavior patterns, and psychological development trajectories of each character, etc. This facilitates the subsequent generation of pictures according to the topological graph, improves the depiction accuracy of the characters in the pictures, and further improves the generation accuracy of the video.

[0050] In this embodiment, it further includes: extracting the relationship information between different characters from the text to be converted; connecting the nodes corresponding to the main information of different characters according to the relationship information between different characters to obtain a character relationship graph; wherein, the character relationship graph represents the relationship between characters.

[0051] Specifically, an LLM can be used to extract the relationship information between different characters from the text to be converted. The relationship information represents the relationship between different characters. For example, keywords representing relationships such as "father" and "mother" can be identified from the text to be converted to obtain the relationship information between different characters. For example, if character A says to character B "Teacher, this question XXXX", it can be obtained that the relationship between character A and character B is a teacher-student relationship.

[0052] According to the relationship information between different characters, a character relationship graph can be constructed. The character relationship graph can include multiple nodes, and each node represents the main information of a character. If there is a certain relationship between two characters, the nodes of the main information of these two characters can be connected by an edge. Figure 3 It is a schematic diagram of the character relationship graph. Figure 3 As can be seen from, for character A, there are relationships with character B, character C, character F, and character G.

[0053] It is also possible to group the characters in the text to be converted according to the relationship information between the characters, and divide some characters into the same group. For example, the characters in the same class can be divided into a group, and the characters in the same family can be divided into a group.

[0054] The beneficial effect of such a setting is that it creates a relationship graph between characters, accurately depicts the multi-dimensional interactions and mutual relationships between characters, analyzes the interactions between characters, optimizes the presentation of character relationships, thereby strengthening the emotional depth and coherence of the story in the video and improving the video quality.

[0055] S102. Determine multiple frames of storyboard pictures according to the text to be converted and the knowledge graphs of each character; wherein, the storyboard pictures represent part of the text content in the text to be converted.

[0056] Exemplarily, multiple images can be generated as storyboard pictures according to the full text of the text to be converted and the knowledge graphs of all characters. Each frame of the storyboard picture can represent part of the text content in the text to be converted, and the text content represented by different storyboard pictures is different.

[0057] For example, the text to be converted can be divided into multiple segments from front to back, and the characters appearing in the text content of each segment are determined. According to the text content of each segment, a picture is generated so that the generated picture can display the text content. Then, according to the knowledge graph of the characters in the text, the characters in the picture are depicted to obtain the image corresponding to the text content of this segment, that is, the storyboard picture is obtained. A storyboard picture can be generated for each segment of text content, so as to obtain multiple storyboard pictures. All the storyboard pictures can express the complete text to be converted.

[0058] In this embodiment, a text-to-image model can be preset. The text-to-image model can be an LLM model. When the text is input into the text-to-image model, pictures can be obtained. In this embodiment, the model structure of the text-to-image model is not specifically limited. For example, the text-to-image model can be the model structure of SD (Stable Diffusion).

[0059] S103. Generate a target video according to multiple frames of storyboard pictures.

[0060] Exemplarily, after obtaining multiple frames of storyboard pictures, the storyboard pictures can be sorted according to the order of the text content corresponding to the storyboard pictures in the text to be converted. According to the sorting result, these storyboard pictures are synthesized into a video as the target video. That is, the target video represents all the text content in the text to be converted. For example, each frame of the storyboard picture can be played in sequence at a preset frequency to obtain the target video.

[0061] Background music, camera movement, transition effects, etc. can be added to the video to enhance the viewing immersion and visual impact. Using audio-video synthesis technology, an AI film and television work with strong visual impact and rich auditory levels can be created. In this embodiment, the audio-video synthesis technology is not specifically limited.

[0062] In the embodiments of the present disclosure, role information of each role is extracted from the text to be converted, a knowledge graph of the role is constructed, and based on the original text of the text to be converted and the knowledge graph of each role, text-to-image operations can be performed, that is, multiple frames of storyboard pictures are generated. Each frame of storyboard picture can represent a part of the text content in the text to be converted. The multiple frames of storyboard pictures are synthesized to obtain the final target video, enabling users to understand the content of the text to be converted by watching the target video. By combining the original text and the knowledge graph of the role, it is possible to avoid missing plot or role information in the storyboard pictures, ensure the consistency of role presentation, reduce labor costs, improve the efficiency and accuracy of text-to-video conversion, and enhance the user's viewing experience.

[0063] Figure 4 It is a schematic flow chart of a video generation method based on a large model provided by an embodiment of the present disclosure.

[0064] In this embodiment, multiple frames of storyboard pictures are determined according to the text to be converted and the knowledge graph of each role, including: splitting the text to be converted to obtain multiple storyboard information; where the storyboard information represents a part of the text content in the text to be converted; for each piece of storyboard information, determine the roles included in the storyboard information; and determine the storyboard picture of the storyboard information according to the storyboard information and the knowledge graph of the roles included in the storyboard information.

[0065] This embodiment is based on the above embodiment, as Figure 4 shown, the method includes the following steps:

[0066] S401. Obtain the text to be converted and construct a knowledge graph of the roles in the text to be converted; where the knowledge graph represents the role information of the roles in the text to be converted, and the knowledge graph corresponds to the roles one by one.

[0067] Exemplarily, this step can refer to the above step S101 and will not be elaborated.

[0068] S402. Split the text to be converted to obtain multiple storyboard information; where the storyboard information represents a part of the text content in the text to be converted.

[0069] Exemplarily, the text to be converted can be a literary work with a large number of words, such as a long or short novel. After obtaining the text to be converted, the text to be converted can be split, and the text to be converted can be decomposed into multiple parts, and each part can be used as a piece of storyboard information. That is, the text to be converted can correspond to multiple pieces of storyboard information. For example, the content of each chapter in the text to be converted can be used as a piece of storyboard information, and the number of chapters in the text to be converted is the number of pieces of storyboard information.

[0070] In this embodiment, the text to be converted is split to obtain a plurality of storyboard information, including: identifying chapter titles from the text to be converted, and splitting the text to be converted into a plurality of text blocks according to the chapter titles; wherein, the text blocks correspond to the chapter titles one by one; for each text block, a plurality of storyboard information is obtained according to the text content of the text block.

[0071] Specifically, each chapter in the text to be converted can have a corresponding title. Perform title recognition processing on the text to be converted to determine the positions of the chapter titles in the text to be converted. For example, each title is located in front of the corresponding chapter, and the font is different from the font of the text content of the chapter. By recognizing the font, the chapter title can be recognized. Or, there is a word count limit for the chapter title, and short sentences with this word count limit are recognized from the text to be converted as the chapter titles. In this embodiment, the recognition method of the chapter titles is not specifically limited.

[0072] According to the position of the chapter title, the text content corresponding to the chapter title can be determined. For example, the content between two chapter titles can be used as the text content of the previous chapter title among these two chapter titles. The text content corresponding to each chapter title is used as a text block. That is, according to the chapter titles, the text to be converted is split into a plurality of text blocks, and each chapter title corresponds to a text block.

[0073] For each text block, the text content of the text block can be further split to obtain a plurality of parts that make up the text block. Each part corresponds to a storyboard information, that is, a text block can correspond to a plurality of storyboard information. For example, it can be split according to the paragraphing of the content in the text block, and each paragraph is split into a storyboard information. Or, perform semantic recognition on the text block, and for each recognized plot, the content of the plot is determined as a storyboard information.

[0074] The beneficial effect of such a setting is that the original text is disassembled according to the chapter titles, and a plurality of storyboard information can be further split in each chapter, realizing a fine-grained division of the text, so that the video can display more content in the text to be converted and improve the generation accuracy of the video.

[0075] In this embodiment, a plurality of storyboard information is obtained according to the text content of the text block, including: splitting the text content of the text block into a plurality of storyboard segments; wherein, the storyboard segments represent part of the text content in the text block; determining the subtitle information of the storyboard segments according to the text content of the storyboard segments; wherein, the subtitle information is the text displayed on the storyboard pictures; determining the text content and subtitle information of the storyboard segments as the storyboard information of the storyboard segments.

[0076] Specifically, for each text block, the text content of the text block can be split to obtain multiple parts that make up the text block, and each part is a storyboard segment, that is, multiple storyboard segments can be obtained. The storyboard segment is the original content in the text block.

[0077] Subtitles are usually displayed in high-quality videos. For each storyboard segment, the corresponding subtitle information can be determined. For example, through semantic recognition or symbol recognition of quotation marks, the words spoken by each character can be identified from the content of the text block, and the words spoken by each character are used as the subtitle information for the storyboard segment.

[0078] The text content and subtitle information of the storyboard segment are determined as the storyboard information of the storyboard segment, that is, each storyboard information includes the original text and subtitle information in the corresponding storyboard segment. Each storyboard information corresponds to a frame of storyboard picture, and the subtitle information is the text that needs to be displayed on the storyboard picture.

[0079] The beneficial effect of such a setting is that each chapter can be split into multiple storyboard segments, and the original text and subtitles of the storyboard segments can be used as storyboard information, so that the storyboard picture corresponding to the storyboard information can represent the meaning of the original text and display subtitles, improving the video quality.

[0080] In this embodiment, determining the subtitle information of the storyboard segment according to the text content of the storyboard segment includes: splitting the text content of the storyboard segment into multiple text segments; inputting the multiple text segments into a preset subtitle generation model to obtain the subtitle information of the storyboard segment; wherein, the preset subtitle generation model is a pre-constructed artificial intelligence model for adjusting the received text into subtitles.

[0081] Specifically, a subtitle generation model is pre-constructed. The subtitle generation model can be an LLM model, and the LLM can adjust the text content into subtitles suitable for display on pictures. That is to say, the subtitle information and the original text content of the storyboard segment can be different. In this embodiment, the model structure of the subtitle generation model is not specifically limited.

[0082] The text content of the storyboard segment can be split first to obtain multiple text segments. For example, every time a full stop is recognized, a text segment is split. The split text segments are sequentially input into a preset subtitle generation model, and the subtitle generation model adjusts the text segments. Each text segment can correspond to one or more sentences of subtitles, so as to obtain the subtitle information of the storyboard segment. Some text segments may not have corresponding subtitles. For example, if a text segment only describes the actions of a character without any background or dialogue, then this text segment may not have corresponding subtitles. That is, the storyboard information can include the text content and subtitle information of the storyboard segment, or only the text content of the storyboard segment.

[0083] The beneficial effect of such a setting is that when obtaining subtitle information, the original text of the storyboard segment can be segmented into multiple small segments by using a preset segmentation rule, and each segment is input into the AI model to automatically output subtitle information. The model comprehensively considers the content and context of the storyboard segment, re-conceives and rewrites it into a more suitable subtitle text, generates a more natural, coherent and easy-to-understand subtitle for the audience while retaining the meaning of the original text, improves the user's viewing experience, reduces manual operations, and improves the efficiency and accuracy of video generation.

[0084] S403. For each storyboard information, determine the characters contained in the storyboard information.

[0085] Exemplarily, each storyboard information may contain characters or no characters may appear. For each storyboard information, it can be checked whether the storyboard information contains characters. For example, it can be determined whether a character name appears in the storyboard information or whether there is a dialogue, etc. If so, it is determined that the storyboard information contains characters. By performing character recognition on the storyboard information, all the characters contained in the storyboard information can be determined.

[0086] S404. Determine the storyboard picture of the storyboard information according to the storyboard information and the knowledge graph of the characters contained in the storyboard information.

[0087] Exemplarily, if the storyboard information does not contain characters, the corresponding storyboard picture is determined according to the storyboard information. For example, the storyboard information can be input into a preset text-to-image model, and the text-to-image model outputs the storyboard picture.

[0088] If the storyboard information contains characters, for each character in the storyboard information, obtain the knowledge graph of the character. Determine the storyboard picture of the storyboard information according to the storyboard information and the knowledge graphs of all the characters contained in the storyboard information. For example, the storyboard information can be first input into a preset text-to-image model, and the picture output by the text-to-image model can be the initial picture. Then, according to the knowledge graph of the character, the character image in the initial picture is adjusted to obtain the storyboard picture. For example, the hairstyle, clothing, etc. of the character in the initial picture can be changed.

[0089] In this embodiment, the original text is split into multiple storyboard information, each storyboard information corresponds to a storyboard picture, and each storyboard picture sequentially displays the overall content to be converted into text. Combining with the knowledge graph of the characters, the storyboard picture is generated. It does not require manual participation, avoids being affected by the deviation of subjective judgment, and improves the efficiency and accuracy of video generation.

[0090] In this embodiment, determining the storyboard picture of the storyboard information according to the storyboard information and the knowledge graph of the characters included in the storyboard information includes: determining the scene information of the storyboard picture and the performance information of the characters included in the storyboard information according to the storyboard information; wherein, the scene information represents the scene shown in the storyboard picture, and the performance information represents the performance of the characters in the scene; determining the storyboard picture of the storyboard information according to the scene information of the storyboard picture, the performance information of the characters included in the storyboard information, and the knowledge graph of the characters included in the storyboard information.

[0091] Specifically, each storyboard information corresponds to a storyboard picture, and the storyboard information can be the information that needs to be represented in the storyboard picture. Perform semantic recognition processing on the storyboard information, and extract the information required for the storyboard picture from the storyboard information. For example, a preset large model can be used for information extraction. The information extracted from the storyboard information can include the scene information of the storyboard picture and the performance information of the characters included in the storyboard information, and the characters included in the storyboard information are also the characters included in the storyboard picture. The scene information can represent the scene shown in the storyboard picture, including the shot type, the background environment, time, weather, etc. of the storyboard picture. The shot type refers to the perspective and angle of the storyboard picture. For example, close-up, long shot, front view, side view, etc. The performance information can represent the performance of the characters in the scene, including the perspective of the camera shooting the characters, the current actions and expressions of the characters, etc.

[0092] Combine the scene information of the storyboard picture, the performance information of each character in the storyboard information, and the knowledge graph to generate the storyboard picture of this storyboard information. For example, according to the scene information of the storyboard picture, generate a picture with only the scene, and then determine the number of characters that need to appear in the picture, as well as the expressions, actions, etc. of each character according to the performance information of each character in the storyboard information. Use the picture obtained from the scene information of the storyboard picture and the performance information of each character in the storyboard information as the initial picture. Then, according to the knowledge graph of each character in the storyboard information, more precisely depict the corresponding characters in the initial picture. For example, the hairstyle, clothing, etc. of the characters can be supplemented to obtain the storyboard picture.

[0093] The beneficial effect of such a setting is that according to the storyboard information, determine the shot type, time, weather, current actions, expressions, etc. of the storyboard, and generate the storyboard picture according to this information and the character knowledge graph, improving the generation accuracy of the storyboard picture, and further improving the generation accuracy of the video.

[0094] In this embodiment, determining the storyboard picture of the storyboard information based on the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information includes: determining the prompt information of the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information; wherein, the prompt information represents the storyboard picture in text form; inputting the prompt information into a preset text-to-image model to obtain the storyboard picture; wherein, the text-to-image model is a pre-constructed artificial intelligence model for converting text into pictures.

[0095] Specifically, a text-to-image model is pre-constructed. The text-to-image model can be a large model, and the large model can cooperate with the prompt when used to improve the understanding ability of the large model. Before generating each storyboard picture, the prompt information of the storyboard picture can be generated first, and then the prompt information is input into the large model, that is, the text-to-image model, to obtain the storyboard picture output by the model. The prompt information is data in text form, which can represent the content of the storyboard picture. The prompt information of the storyboard picture can be called the storyboard SD prompt.

[0096] A prompt template is preset. According to the prompt template, the required information is extracted from the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information, and the extracted information is added to the corresponding position in the template to obtain the prompt information of the storyboard picture. For example, the prompt template can be "You are a painter, and you need to draw a picture from the perspective of XXX. The weather in the picture is XXX. The picture includes XXX characters, and the characters are doing XXX".

[0097] The beneficial effect of such a setting is to combine various information to generate the prompt of the storyboard picture and use the AI model to improve the accuracy and efficiency of text-to-image.

[0098] In this embodiment, determining the prompt information of the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information includes: determining the necessary information and prohibited information in the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information; wherein, the necessary information is the information that needs to be displayed in the storyboard picture, and the prohibited information is the information that cannot be displayed in the storyboard picture; determining the positive prompt according to the necessary information, and determining the negative prompt according to the prohibited information; wherein, the positive prompt is the prompt information representing the necessary information, and the negative prompt is the prompt information representing the prohibited information.

[0099] Specifically, the prompt information may include positive prompts and / or negative prompts. Positive prompts represent the content that needs to appear in the storyboard pictures, and negative prompts represent the content that cannot appear in the storyboard pictures. Based on the scene information of the storyboard pictures, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information, the necessary information in the storyboard pictures is determined. The necessary information is the content that needs to appear in the storyboard pictures. According to the necessary information and the preset positive prompt template, positive prompts can be obtained. For example, the necessary information can be filled in the preset positions in the positive prompt template to obtain positive prompts.

[0100] It is also possible to determine the prohibited information in the storyboard pictures based on the scene information of the storyboard pictures, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information. The prohibited information is the content that cannot appear in the storyboard pictures. According to the prohibited information and the preset negative prompt template, negative prompts can be obtained. For example, if the scene information indicates that the time shown in the storyboard picture is noon, it means that stars cannot appear in the storyboard picture, that is, the prohibited information can include stars.

[0101] The beneficial effect of such a setting is that by determining positive prompts and negative prompts, the comprehensiveness of the prompt information can be improved, thereby improving the generation accuracy of the storyboard pictures.

[0102] In this embodiment, inputting the prompt information into a preset text-to-image model to obtain storyboard pictures includes: determining the work category corresponding to the text to be converted; according to the preset association relationship, determining the text-to-image model corresponding to the work category of the text to be converted as the target model; where the preset association relationship represents the association relationship between the work category and the text-to-image model; inputting the prompt information into the target model to obtain storyboard pictures.

[0103] Specifically, the text to be converted corresponds to its own work category. For example, if the text to be converted is a science fiction novel, it can be determined that the work category of this novel is science fiction. The powerful semantic understanding ability of the LLM can be used to analyze the key elements in the text to be converted, such as characters, time, place, events, era background, and other specific details and attributes, to judge the work category to which the text to be converted belongs, which helps to set an appropriate visual and emotional tone for AI film and television works.

[0104] Different work categories can be associated with different text-to-image models. According to the preset association relationship, the text-to-image model corresponding to the work category of the text to be converted can be found as the target model. Inputting the prompt information into the target model to obtain the storyboard pictures output by the model.

[0105] The beneficial effect of such a setting is that by selecting a text-to-image model that matches the work category, precise image generation for different work categories can be achieved, improving the pertinence and accuracy of image generation.

[0106] After obtaining the work category, the prompt information for the storyboard images can be determined based on the work category of the text to be converted, the scene information of the storyboard images, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information. That is, the work category can be included in the prompt information, improving the comprehensiveness of the prompt information and further enhancing the accuracy of image generation.

[0107] S405. Generate a target video based on multiple frames of storyboard images; wherein, the target video represents the text content in the text to be converted.

[0108] Exemplarily, this step can refer to the above step S103 and will not be elaborated here.

[0109] In the embodiments of the present disclosure, the role information of each character is extracted from the text to be converted, and a knowledge graph of the characters is constructed. Based on the original text of the text to be converted and the knowledge graph of each character, text-to-image operations can be performed, that is, multiple frames of storyboard images are generated. Each frame of storyboard image can represent a part of the text content in the text to be converted. The multiple frames of storyboard images are synthesized to obtain the final target video, enabling users to understand the content of the text to be converted by watching the target video. By combining the original text and the knowledge graph of the characters, it is possible to avoid missing plot or character information in the storyboard images, ensure the consistency of character presentation, reduce labor costs, improve the efficiency and accuracy of text-to-video conversion, and enhance the user's viewing experience.

[0110] Figure 5 It is a schematic flowchart of a video generation method based on a large model provided by the embodiments of the present disclosure.

[0111] In this embodiment, the scene information representing the storyboard images in the storyboard information and the performance information of the characters contained in the storyboard information, where the scene information represents the scene shown in the storyboard images, and the performance information represents the performance of the characters in the scene; generating a target video based on multiple frames of storyboard images includes: for each frame of storyboard image, converting the storyboard image from a static form to a dynamic form according to the scene information of the storyboard image and the performance information of the characters contained in the storyboard information; obtaining the target video based on the storyboard images in dynamic form.

[0112] This embodiment is based on the above embodiments, as Figure 5 shown, the method includes the following steps:

[0113] S501. Obtain the text to be converted and construct a knowledge graph for the characters in the text to be converted; wherein, the knowledge graph represents the character information of the characters in the text to be converted, and the knowledge graph corresponds to the characters one by one.

[0114] Exemplarily, this step can refer to the above step S101 and will not be elaborated here.

[0115] S502. Determine multiple frames of storyboard pictures according to the text to be converted and the knowledge graphs of each character; wherein, the storyboard pictures represent partial text content in the text to be converted, and the content represented by different storyboard pictures is different.

[0116] Exemplarily, this step can refer to the above step S102 and will not be elaborated here.

[0117] S503. For each frame of the storyboard picture, convert the storyboard picture from a static form to a dynamic form according to the scene information of the storyboard picture and the performance information of the characters contained in the storyboard information.

[0118] Exemplarily, the storyboard picture is a static picture. After obtaining multiple frames of storyboard pictures, for each frame of the storyboard picture, the static picture can be converted into a dynamic picture, that is, a storyboard picture in the form of a moving picture is obtained. The scene information of the storyboard picture and the performance information of the characters contained in the storyboard information can be understood and analyzed to determine the actions to be presented in the storyboard picture. For example, if the scene information of the storyboard picture is a playground scene and the performance information of the character is a running race, the action of the character running can be converted into a dynamic one.

[0119] The technology of converting a preset static picture into a dynamic picture can be used to convert the static picture into a dynamic picture, accurately restoring the scene of the storyboard picture and the actions of the characters in the storyboard. In this embodiment, the technology of converting a preset static picture into a dynamic picture is not specifically limited. For example, image processing and animation technologies can be used to inject dynamic elements such as zooming, panning, rotating or simulating camera movement into the static picture to create a more vivid visual effect.

[0120] S504. Obtain the target video according to each storyboard picture in the dynamic form.

[0121] Exemplarily, the storyboard picture in the dynamic form is determined as a storyboard moving picture. The storyboard pictures are sorted according to the order of the storyboard information in the entire text to be converted, and the storyboard moving pictures are connected in the order of the storyboard pictures, thereby synthesizing the target video.

[0122] In this embodiment, multiple moving pictures are synthesized into a complete video, improving the fluency between videos, avoiding action jerks between static pictures, realizing automated film and television creation, and improving the generation accuracy and quality of the video.

[0123] In this embodiment, the storyboard information includes subtitle information; obtaining the target video based on the storyboard pictures in various dynamic forms includes: determining the subtitle information corresponding to the storyboard pictures, and converting the subtitle information into voice information; obtaining the target video based on the storyboard pictures in various dynamic forms and the corresponding voice information, based on preset video special effects.

[0124] Specifically, the storyboard information may contain the subtitle information of the storyboard pictures. If there is no subtitle information in the storyboard information, no subtitles will appear in the corresponding storyboard pictures; if there is subtitle information in the storyboard information, the subtitle information corresponding to the storyboard pictures can be determined, and the subtitle information can be embedded into the corresponding animated storyboard pictures to achieve seamless combination of visual content and text information. For example, the subtitle information can be added to the lower half of the animated storyboard pictures.

[0125] The subtitle information can also be converted into voice information. For example, through TTS (Text To Speech) technology, appropriate voice-over, sound effects, speech rate and other sound parameters can be selected to convert the subtitle information into synchronous voice information. When playing the target video, the subtitle information is displayed on the animated storyboard pictures, and the corresponding voice information is played simultaneously to ensure the perfect integration of vision and audition.

[0126] Background music and carefully designed special effects such as camera movement and transitions can also be preset to enhance the viewing immersion and visual impact. For example, dissolve, wipe, slide, etc., to smoothly transition between each video clip, while enhancing the visual impact and rhythm. Narrative techniques can also be used, supplemented by sound effects, text and graphic design, to inject a story context and emotional depth into static pictures and enhance the narrative effect. Interactive elements such as click actions that trigger dynamic effects can be introduced in the video to provide the audience with more sense of participation and rich interactive experiences. Through audio-visual synthesis technology, an AI film and television work with both strong visual impact and rich auditory levels is finally created.

[0127] The beneficial effect of such a setting is that the subtitle information is converted into voice information and added to the video, which improves the richness of the video form and enhances the user's viewing experience.

[0128] Figure 6 It is a schematic diagram of the process of novel film and television adaptation. The text to be converted is the original novel text. For the original novel text, the powerful semantic understanding ability of the LLM can be utilized to judge the work category of the novel by analyzing the key elements in the novel text, such as characters, time, place, events, era background, and other specific details and attributes.

[0129] Utilize the powerful information extraction ability of the LLM to accurately extract the character subjects, aliases, relationships between characters, and detailed descriptions of characters in the novel, comprehensively and deeply understand the character composition in the novel, and obtain character information. Then, through the LLM, deeply analyze the relationships between characters and the character description information, and precisely construct a detailed static image description of each character. For character roles, it includes various information such as the character's gender, age, body shape, face shape, hairstyle and color, eye color, and the type and color of clothing, etc., so as to provide necessary data support for comprehensively understanding and accurately depicting each character role, and create a knowledge graph of characters. It is also possible to create a character relationship graph based on the relationships between characters to accurately depict the multi-dimensional interactions and mutual relationships between characters, analyze the interactions between characters, optimize the presentation of character relationships, and construct a character archive library for the entire book dimension.

[0130] Adopt preset rules and combine them with the LLM to split the entire novel into chapters, obtaining several relatively independent text blocks, and each text block can be a chapter. For each chapter, extract the character names that appear in the chapter to effectively match with the relevant image information in the character archive library. For each chapter, adopt a method that combines rule matching and the large model to subdivide the main text content of each chapter into multiple non-repeating storyboard segments, and obtain information such as the content and subtitles of the storyboard segments as the storyboard information of the storyboard segments. For parts with dense plots or highlight plots, the storyboard can be refined to enhance the visual and narrative effects. Ensure the integrity of individual storyboard information and the coherence and fluency of the plot and vision between storyboards, so that each storyboard transition is logically self-consistent, the visual narrative has no breaks or repetitions, and the audience can smoothly follow the development of the story context and clearly understand the relationships and plot developments between each scene without additional explanations.

[0131] By analyzing the storyboard information of each storyboard segment, combining the relevant information of the novel category and characters, generate corresponding prompt information, that is, the storyboard SD prompt. The prompt information can include positive prompts and negative prompts to accurately guide the text-to-image model to generate pictures that fit the storyboard information, improving the picture quality and performance effects.

[0132] Generate high-quality static pictures through a preset text-to-image model, or first select a text-to-image model that matches the storyboard according to the novel category information in the storyboard SD prompt, and then generate corresponding storyboard pictures based on the storyboard SD prompt using the selected text-to-image model.

[0133] Based on each storyboard picture, combined with the scene information of the storyboard segment and the information of specific characters in the current scene, use the technology of converting static pictures to dynamic pictures to convert the static pictures into dynamic pictures, accurately restoring the storyboard scene and the actions of the characters in the storyboard segment.

[0134] Use audio - video synthesis technology to create the final AI film and television works. During the video synthesis process, the subtitle information of each storyboard can be embedded into the corresponding dynamic graph to achieve seamless integration of visual content and text information. Subsequently, through TTS technology, appropriate dubbing, sound effects, speech rate, etc. are selected to convert the subtitle information into synchronized voice information to ensure perfect integration of vision and hearing. Background music and carefully designed special effects such as camera movements and transitions can also be added to enhance the viewing immersion and visual impact.

[0135] In the embodiments of the present disclosure, character information of each character is extracted from the text to be converted, and a knowledge graph of the character is constructed. According to the original text of the text to be converted and the knowledge graphs of each character, text - to - image operations can be performed, that is, multiple frames of storyboard pictures are generated. Each frame of the storyboard picture can represent a part of the text content in the text to be converted. The multiple frames of storyboard pictures are synthesized to obtain the final target video, enabling users to understand the content of the text to be converted by watching the target video. By combining the original text and the knowledge graph of the character, it is possible to avoid missing plot or character information in the storyboard pictures, ensure the consistency of character presentation, reduce labor costs, improve the efficiency and accuracy of text - to - video conversion, and enhance the user's viewing experience.

[0136] Figure 7 The following is a structural block diagram of a video generation device based on a large - model provided by the embodiments of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Referring to Figure 7 Figure, the video generation device 700 based on a large - model includes: a graph construction unit 701, a picture determination unit 702, and a video generation unit 703.

[0137] The graph construction unit 701 is used to obtain the text to be converted and construct a knowledge graph of the characters in the text to be converted; wherein, the knowledge graph represents the character information of the characters in the text to be converted;

[0138] The picture determination unit 702 is used to determine multiple frames of storyboard pictures according to the text to be converted and the knowledge graphs of each character; wherein, the storyboard pictures represent part of the text content in the text to be converted;

[0139] The video generation unit 703 is used to generate a target video according to the multiple frames of storyboard pictures.

[0140] Figure 8 The following is a structural block diagram of a video generation device based on a large - model provided by the embodiments of the present disclosure, as shown in Figure 8As shown in the figure, the large model-based video generation device 800 includes a graph construction unit 801, a picture determination unit 802, and a video generation unit 803. Among them, the picture determination unit 802 includes a text splitting module 8021, a role determination module 8022, and a picture determination module 8023.

[0141] The text splitting module 8021 is used to split the text to be converted to obtain a plurality of storyboard information; wherein, the storyboard information represents part of the text content in the text to be converted;

[0142] The role determination module 8022 is used to determine the roles included in each storyboard information for each storyboard information;

[0143] The picture determination module 8023 is used to determine the storyboard pictures of the storyboard information according to the storyboard information and the knowledge graph of the roles included in the storyboard information.

[0144] In one example, the text splitting module 8021 includes:

[0145] The text splitting sub-module is used to identify the chapter titles from the text to be converted and split the text to be converted into a plurality of text blocks according to the chapter titles; wherein, the text blocks correspond to the chapter titles one by one;

[0146] The information obtaining sub-module is used to obtain a plurality of storyboard information for each text block according to the text content of the text block.

[0147] In one example, the information obtaining sub-module is specifically used for:

[0148] Split the text content of the text block into a plurality of storyboard segments; wherein, the storyboard segments represent part of the text content in the text block;

[0149] Determine the subtitle information of the storyboard segment according to the text content of the storyboard segment; wherein, the subtitle information is the text displayed on the storyboard picture;

[0150] Determine the storyboard information of the storyboard segment by using the text content and subtitle information of the storyboard segment.

[0151] In one example, the information obtaining sub-module is specifically used for:

[0152] Split the text content of the storyboard segment into a plurality of text segments;

[0153] Input the plurality of text segments into a preset subtitle generation model to obtain the subtitle information of the storyboard segment; wherein, the preset subtitle generation model is a pre-constructed artificial intelligence model used to adjust the received text into subtitles.

[0154] In one example, the picture determination module 8023 includes:

[0155] An information determination sub-module, configured to determine the scene information of the storyboard picture and the performance information of the characters included in the storyboard information according to the storyboard information; wherein, the scene information represents the scene shown in the storyboard picture, and the performance information represents the performance of the characters in the scene;

[0156] A picture determination sub-module, configured to determine the storyboard picture of the storyboard information according to the scene information of the storyboard picture, the performance information of the characters included in the storyboard information, and the knowledge graph of the characters included in the storyboard information.

[0157] In one example, the picture determination sub-module is specifically configured to:

[0158] Determine the prompt information of the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters included in the storyboard information, and the knowledge graph of the characters included in the storyboard information; wherein, the prompt information represents the storyboard picture in text form.

[0159] Input the prompt information into a preset text-to-image model to obtain the storyboard picture; wherein, the text-to-image model is a pre-constructed artificial intelligence model for converting text into pictures.

[0160] In one example, the picture determination sub-module is specifically configured to:

[0161] Determine the necessary information and prohibited information in the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters included in the storyboard information, and the knowledge graph of the characters included in the storyboard information; wherein, the necessary information is the information that needs to be shown in the storyboard picture, and the prohibited information is the information that cannot be shown in the storyboard picture.

[0162] Determine a positive prompt according to the necessary information, and determine a negative prompt according to the prohibited information; wherein, the positive prompt is the prompt information representing the necessary information, and the negative prompt is the prompt information representing the prohibited information.

[0163] In one example, the picture determination sub-module is specifically configured to:

[0164] Determine the work category corresponding to the text to be converted;

[0165] Determine, according to a preset association relationship, the text-to-image model corresponding to the work category of the text to be converted as the target model; wherein, the preset association relationship represents the association relationship between the work category and the text-to-image model.

[0166] Input the prompt word information into the target model to obtain the storyboard picture.

[0167] In one example, the scene information representing the storyboard picture and the performance information of the characters contained in the storyboard information in the storyboard information, the scene information represents the scene shown in the storyboard picture, and the performance information represents the performance of the characters in the scene; the video generation unit 803 includes:

[0168] A picture conversion module, for each frame of the storyboard picture, convert the storyboard picture from a static form to a dynamic form according to the scene information of the storyboard picture and the performance information of the characters contained in the storyboard information;

[0169] A video obtaining module, for obtaining the target video according to the storyboard pictures in dynamic form.

[0170] In one example, the storyboard information includes subtitle information; the video obtaining module includes:

[0171] A subtitle conversion sub-module, for determining the subtitle information corresponding to the storyboard picture and converting the subtitle information into voice information;

[0172] A video obtaining sub-module, for obtaining the target video based on the preset video special effects according to the storyboard pictures in dynamic form and the corresponding voice information.

[0173] In one example, the character information includes subject information and attribute information, the subject information represents the identifier of the character, and the attribute information represents the situation related to the character; the knowledge graph construction unit 801 includes:

[0174] An information extraction module, for extracting the subject information and attribute information of the characters from the text to be converted;

[0175] A node connection module, for connecting the nodes corresponding to the respective attribute information to the nodes corresponding to the subject information to obtain the knowledge graph of the character.

[0176] In one example, it further includes:

[0177] A relationship extraction unit, for extracting the relationship information between different characters from the text to be converted;

[0178] A relationship construction unit, for connecting the nodes corresponding to the subject information of different characters according to the relationship information between different characters to obtain a character relationship graph; wherein, the character relationship graph represents the relationship between characters.

[0179] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device.

[0180] Figure 9 The structural block diagram of an electronic device provided by an embodiment of the present disclosure is shown as Figure 9 shown. The electronic device 900 includes: at least one processor 902; and a memory 901 communicatively connected to the at least one processor 902. Wherein, the memory stores instructions executable by the at least one processor 902, and the instructions are executed by the at least one processor 902 to enable the at least one processor 902 to execute the video generation method based on a large model of the present disclosure.

[0181] The electronic device 900 further includes a receiver 903 and a transmitter 904. The receiver 903 is used to receive instructions and data sent by other devices, and the transmitter 904 is used to send instructions and data to external devices.

[0182] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0183] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product. The computer program product includes: a computer program. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to enable the electronic device to execute the solution provided in any of the above embodiments.

[0184] Figure 10 The schematic block diagram of an example electronic device 1000 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0185] As Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0186] Multiple components in device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0187] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the large model-based video generation method. For example, in some embodiments, the large model-based video generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the large model-based video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the large model-based video generation method by any other appropriate means (e.g., by means of firmware).

[0188] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0189] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0190] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0191] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0192] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0193] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0194] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.

[0195] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A video generation method based on a large model, comprising: Obtaining the text to be converted and constructing a knowledge graph of the characters in the text to be converted; wherein, the knowledge graph represents the character information of the characters in the text to be converted; Determining multiple storyboard pictures according to the text to be converted and the knowledge graphs of each character; wherein, the storyboard pictures represent partial text content in the text to be converted; Generating a target video according to the multiple storyboard pictures.

2. The method according to claim 1, wherein The determining multiple storyboard pictures according to the text to be converted and the knowledge graphs of each character includes: Performing a splitting process on the text to be converted to obtain multiple storyboard information; wherein, the storyboard information represents partial text content in the text to be converted; For each storyboard information, determining the characters included in the storyboard information; Determining the storyboard picture of the storyboard information according to the storyboard information and the knowledge graph of the characters included in the storyboard information.

3. The method according to claim 2, wherein, The performing a splitting process on the text to be converted to obtain multiple storyboard information includes: Identifying chapter titles from the text to be converted and splitting the text to be converted into multiple text blocks according to the chapter titles; wherein, the text blocks correspond one-to-one with the chapter titles; For each text block, obtaining multiple storyboard information according to the text content of the text block.

4. The method according to claim 3, wherein The obtaining multiple storyboard information according to the text content of the text block includes: Splitting the text content of the text block into multiple storyboard segments; wherein, the storyboard segments represent partial text content in the text block; Determining the subtitle information of the storyboard segment according to the text content of the storyboard segment; wherein, the subtitle information is the text displayed on the storyboard picture; Determining the text content and subtitle information of the storyboard segment as the storyboard information of the storyboard segment.

5. The method according to claim 4, wherein The determining the subtitle information of the storyboard segment according to the text content of the storyboard segment includes: Splitting the text content of the storyboard segment into multiple text segments; Inputting the multiple text segments into a preset subtitle generation model to obtain the subtitle information of the storyboard segment; wherein, the preset subtitle generation model is a pre-constructed artificial intelligence model for adjusting the received text into subtitles.

6. The method according to any one of claims 2-5, wherein, The determining the storyboard picture of the storyboard information according to the storyboard information and the knowledge graph of the characters included in the storyboard information includes: Determining the scene information of the storyboard picture and the performance information of the characters included in the storyboard information according to the storyboard information; wherein, the scene information represents the scene shown in the storyboard picture, and the performance information represents the performance of the characters in the scene; Determining the storyboard picture of the storyboard information according to the scene information of the storyboard picture, the performance information of the characters included in the storyboard information, and the knowledge graph of the characters included in the storyboard information.

7. The method according to claim 6, wherein The determining the storyboard picture of the storyboard information according to the scene information of the storyboard picture, the performance information of the characters included in the storyboard information, and the knowledge graph of the characters included in the storyboard information includes: Determine the prompt information of the storyboard picture based on the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information; wherein, the prompt information represents the storyboard picture in text form. Input the prompt information into a preset text-to-image model to obtain the storyboard picture; wherein, the text-to-image model is a pre-constructed artificial intelligence model for converting text into pictures.

8. The method according to claim 7, wherein The step of determining the prompt information of the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information includes: Determine the necessary information and prohibited information in the storyboard picture according to the scene information of the storyboard picture, the performance information of the characters contained in the storyboard information, and the knowledge graph of the characters contained in the storyboard information; wherein, the necessary information is the information that needs to be displayed in the storyboard picture, and the prohibited information is the information that cannot be displayed in the storyboard picture. Determine the positive prompt words according to the necessary information, and determine the negative prompt words according to the prohibited information; wherein, the positive prompt words are the prompt information representing the necessary information, and the negative prompt words are the prompt information representing the prohibited information.

9. The method according to claim 7 or 8, wherein The step of inputting the prompt information into a preset text-to-image model to obtain the storyboard picture includes: Determine the work category corresponding to the text to be converted. According to a preset association relationship, determine the text-to-image model corresponding to the work category of the text to be converted as the target model; wherein, the preset association relationship represents the association relationship between the work category and the text-to-image model. Input the prompt information into the target model to obtain the storyboard picture.

10. The method according to any one of claims 1-9, wherein, The scene information representing the storyboard picture and the performance information of the characters contained in the storyboard information in the storyboard information, the scene information represents the scene shown in the storyboard picture, and the performance information represents the performance of the characters in the scene. The step of generating a target video according to the multiple frames of storyboard pictures includes: For each frame of the storyboard picture, convert the storyboard picture from a static form to a dynamic form according to the scene information of the storyboard picture and the performance information of the characters contained in the storyboard information. Obtain the target video according to the storyboard pictures in dynamic form.

11. The method according to claim 10, wherein, The storyboard information includes subtitle information. The step of obtaining the target video according to the storyboard pictures in dynamic form includes: Determine the subtitle information corresponding to the storyboard picture, and convert the subtitle information into voice information. Based on preset video effects, obtain the target video according to the storyboard pictures in dynamic form and the corresponding voice information.

12. The method according to any one of claims 1-11, wherein, The character information includes subject information and attribute information, the subject information represents the identifier of the character, and the attribute information represents the situation related to the character. The step of constructing the knowledge graph of the characters in the text to be converted includes: Extract the subject information and attribute information of the characters from the text to be converted. Connect the nodes corresponding to each attribute information to the nodes corresponding to the subject information respectively to obtain the knowledge graph of the role.

13. The method according to claim 12, further comprising: Extract the relationship information between different roles from the text to be converted; Connect the nodes corresponding to the subject information of different roles according to the relationship information between different roles to obtain a role relationship graph; wherein, the role relationship graph represents the relationship between roles.

14. A video generation device based on a large model, comprising: A graph construction unit, configured to obtain the text to be converted and construct a knowledge graph of the roles in the text to be converted; wherein, the knowledge graph represents the role information of the roles in the text to be converted; A picture determination unit, configured to determine multiple frames of storyboard pictures according to the text to be converted and the knowledge graphs of each role; wherein, the storyboard pictures represent partial text content in the text to be converted; A video generation unit, configured to generate a target video according to the multiple frames of storyboard pictures.

15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-13.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.

17. A computer program product, wherein, Including a computer program, which when executed by a processor implements the steps of the method according to any one of claims 1-13.

Citation Information

Cited By

  • Video generation method and device, terminal and computer readable storage medium

    CN121000947A