Video generation method and device based on AIGC and storage medium
Through the AIGC-based video generation method, song lyrics information is extracted and videos that fit the artistic conception are generated, which solves the problems of low efficiency and inconsistent artistic conception in the existing technology, and achieves efficient and high-quality song video generation.
Patent Information
- Application Number
- CN202311843227.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology has a lot of work in the production of song videos, which cannot effectively match the artistic conception of the lyrics, and the label matching technology makes it easy for different song videos to have the same picture problems, affecting the user experience.
AIGC-based video generation method is adopted to extract the information in the song lyrics, generate a video story summary, generate a storyline and story shot in segments, and generate images/videos based on the story shot, and finally process and generate song videos in chronological order.
It improves the efficiency and quality of video production, makes the video content better fit the artistic conception of the lyrics, reduces the possibility of the same picture appearing, and improves the user experience.
Smart Images

Figure CN120238693A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multimedia technology, and in particular to a video generation method, device and storage medium based on AIGC. Background Art
[0002] A song video is a video made for a song and played synchronously with the song. In the prior art, song videos are mainly made manually or based on tag matching technology. Manually made song videos mainly involve manually making videos for each song, which is labor-intensive, time-consuming and labor-intensive. It is difficult to meet the needs of rapid music production, and the production effect depends entirely on personal production experience. When making a song video using tag matching technology, first label each material in the material library, and then use the material labels to automatically assemble the required materials to obtain a song video. Although the use of tag matching to make song videos is more efficient, the produced song videos are not sufficiently consistent with the artistic conception of the lyrics, and cannot fully express the emotions of the lyrics. In addition, due to the shared material library, it is easy to have problems with different song videos matching the same screen, affecting the user experience.
[0003] Large Language Model (LLM), also known as Large Language Model, is a natural language processing technology with a deep learning model. It is based on neural networks and is trained using a large amount of text data, enabling it to perform well on natural language processing tasks.
[0004] So far, no research has been found on using large language models to generate song videos with high quality. Summary of the invention
[0005] In view of the above problems, the present application provides a song video generation method, which is used to solve the technical problems existing in the above song video production, such as large workload and inability to effectively match the mood of the lyrics.
[0006] To achieve the above object, the inventor provides a video generation method based on AIGC, comprising the following steps:
[0007] Extract the information expressed in the lyrics of the song;
[0008] generating a story summary of the video based on the information;
[0009] Dividing the lyrics into a plurality of lyrics segments, inputting the story summary and each of the lyrics segments into an artificial intelligence model, and generating a storyline corresponding to each of the lyrics segments;
[0010] Input each of the lyric segments and the corresponding storylines into the artificial intelligence model to generate the storyboards corresponding to each lyric segment, and then generate the corresponding images / videos based on each storyboard;
[0011] Sort and process the images / videos generated from each storyboard in chronological order to obtain the song video corresponding to the song.
[0012] In some technical solutions, the information expressed in the lyrics includes the main idea of the lyrics and the images that appear repeatedly in the lyrics. The images include any one or more of the following: scenery, items, concepts, and characters.
[0013] In some technical solutions, generating the storylines corresponding to each lyric segment further includes: generating the basic environment corresponding to each storyline, and the basic environment includes any one or more of the following: main color tone, location, time, season, and emotional tone.
[0014] In some technical solutions, when inputting each of the lyric segments and the corresponding storylines into the artificial intelligence model to generate the storyboards corresponding to each lyric segment, one lyric segment generates more than one storyboard, and the multiple storyboards corresponding to the same lyric segment maintain the same basic environment.
[0015] In some technical solutions, generating the corresponding images / videos based on each storyboard includes: generating the prompts for generating the corresponding images / videos according to the storyboard; inputting the prompts into the image or video generation intelligent model to generate the images / videos corresponding to each storyboard.
[0016] In some technical solutions, inputting the prompts into the image or video generation intelligent model to generate the images / videos corresponding to each storyboard includes:
[0017] Using the text-to-image technology of Stable Diffusion, generate the corresponding images according to the prompts of each storyboard;
[0018] Or use the text-to-video technology of Imagen to generate the corresponding videos according to the prompts of each storyboard;
[0019] Or based on the images already generated by the Stable Diffusion technology, further generate videos using the Stable Video Diffusion video generation technology.
[0020] In some technical solutions, inputting the prompts into the image or video generation intelligent model to generate the images / videos corresponding to each storyboard includes the steps:
[0021] Select a 3D scene and a character model that meet the requirements of the prompt from the multimedia database according to the prompt;
[0022] Load the character model into the 3D scene, and generate an action sequence for the character model according to the story content in the storyboard or the prompt;
[0023] Control the character model to execute the action sequence in the 3D scene and render it into a video clip.
[0024] In some technical solutions, the information expressed in the lyrics further includes a character description;
[0025] The story summary of generating a video according to the information includes: inputting two or more of the character description, the main idea, and the imagery into an artificial intelligence model, and the artificial intelligence model uses natural language processing technology to perform a text generation task to obtain the story summary.
[0026] In some technical solutions, the prompt words include any two or more of camera movement, characters, character feature description, character action description, location, location scenery description, location item description, lighting atmosphere, and time.
[0027] In some technical solutions, in generating the storyboard corresponding to each lyric segment, it includes more than one atmosphere rendering shot, and the atmosphere rendering shot includes a special effect shot, a close-up shot of an object, or an empty shot.
[0028] In some technical solutions, the steps for extracting the information expressed in the lyrics of a song include:
[0029] Input the lyrics into an artificial intelligence model, and through performing text summarization and information extraction tasks in natural language processing technology, obtain the main idea and imagery to be expressed in the lyrics.
[0030] To solve the above technical problems, the present application also provides another technical solution:
[0031] A video generation device based on AIGC, including:
[0032] An information extraction module, configured to extract the information expressed in the lyrics of a song, and generate a story summary of a video according to the information;
[0033] A segmentation module, which divides the lyrics into multiple lyric segments, inputs the story summary and each lyric segment into an artificial intelligence model, and generates a storyline corresponding to each lyric segment;
[0034] The shot generation module inputs each of the lyric segments and the corresponding storylines into the artificial intelligence model to generate the shots corresponding to each lyric segment, and then generates the corresponding images / videos based on each shot.
[0035] The song video generation module sorts and processes the images / videos generated from each shot in chronological order to obtain the song video corresponding to the song.
[0036] To solve the above technical problems, the present application also provides another technical solution:
[0037] A computer-readable storage medium stores a computer program, and when the computer program is run, it executes the AIGC-based video generation method described in any one of the above technical solutions.
[0038] Different from the prior art, the above technical solution divides the production of the song video into multiple steps, such as lyric information extraction, generating the story outline of the video, generating the shots of the lyric segments, and generating images / videos from the shots, etc., and combines the information of the lyrics in each step, so as to create a video exclusive to the song and make the video content better fit the artistic conception of the lyrics. And in the above technical solution, at least when generating the shots and generating the images / videos corresponding to each shot, an artificial intelligence model is used, such as: large language model, image generation model, video generation model, etc., thus greatly improving the efficiency and quality of video production.
[0039] The above relevant description of the invention content is only an overview of the technical solution of the present application. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of the present application, and then can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of the present application more easily understood, the following is described in conjunction with the specific embodiments of the present application and the drawings. Description of the Drawings
[0040] The drawings are only used to show the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation to the present application.
[0041] In the drawings of the specification:
[0042] Figure 1 It is a flowchart of the AIGC-based video generation method described in the specific embodiment;
[0043] Figure 2 It is a flowchart of extracting lyric information described in the specific embodiment;
[0044] Figure 3 It is a flowchart of generating a 3D video described in the specific embodiment;
[0045] Figure 4 It is a block diagram of the AIGC-based video generation device described in the specific implementation manners;
[0046] Figure 5 It is a schematic diagram of the computer-readable storage medium described in the specific implementation manners;
[0047] The descriptions of the reference numerals involved in the above-mentioned respective drawings are as follows:
[0048] 400, AIGC-based video generation device; 401, Information extraction module; 402, Segmentation module;
[0049] 403, Sub-shot generation module; 404, Song video generation module;
[0050] 500, Computer-readable storage medium; Specific implementation manners
[0051] To illustrate in detail the possible application scenarios, technical principles, specific implementable solutions, achievable purposes and effects, etc. of the present application, the following is described in detail with reference to the specific examples listed and in conjunction with the drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present application, and thus are only examples and cannot be used to limit the protection scope of the present application.
[0052] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The term "embodiment" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it particularly limited to the independence or relevance to other embodiments. In principle, in the present application, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0053] Unless otherwise defined, the meanings of the technical terms used herein are the same as those commonly understood by those skilled in the technical field to which the present application belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit the present application.
[0054] In the description of the present application, the phrase "and / or" is an expression used to describe the logical relationship between objects, indicating that three relationships may exist. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " herein generally represents an "or" logical relationship between the associated objects before and after.
[0055] In this application, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantitative, primary-secondary, or sequential relationship between these entities or operations.
[0056] Without further limitation, in this application, the open-ended expressions such as "including", "comprising", "having", or other similar expressions used in a statement are intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method, or product that includes the said elements. Thus, in a process, method, or product that includes a series of elements, it can include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such a process, method, or product.
[0057] Similar to the understanding in the "Examination Guidelines", in this application, expressions such as "greater than", "less than", "exceeding", etc. are understood not to include the recited number; expressions such as "above", "below", "within", etc. are understood to include the recited number. In addition, in the description of the embodiments of this application, the meaning of "a plurality of" is two or more (including two). Similar expressions related to "many", such as "multiple groups", "multiple times", etc., are understood in the same way, unless otherwise specifically defined.
[0058] In the description of the embodiments of this application, the spatially related expressions used, such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiment or the accompanying drawings. This is only for the convenience of describing the specific embodiments of this application or for the reader's understanding, and does not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, it should not be construed as a limitation on the embodiments of this application.
[0059] Unless otherwise clearly specified or limited, in the description of the embodiments of this application, the terms such as "installed", "connected", "joined", "fixed", "set", etc. should be understood in a broad sense. For example, the said "connection" can be a fixed connection, a detachable connection, or an integral setting; it can be a mechanical connection, an electrical connection, or a communication connection; it can be directly connected, or indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art to which this application pertains, the specific meaning of the above terms in the embodiments of this application can be understood according to specific circumstances.
[0060] This embodiment provides a video generation method based on AIGC. The video generation based on AIGC can be used to automatically produce a song video (i.e., MV) for a song with lyrics. In this song video generation method, based on AIGC (i.e., generative artificial intelligence technology), the production of the song video is divided into multiple sub-steps such as generating a story outline for the video, generating a storyline, generating shot sequences, converting the shot sequences into images / videos, and editing to generate the song video. In each step, the main idea and imagery of the lyrics are combined to create a video exclusive to the song, enabling the video content to better fit the mood of the lyrics.
[0061] Please refer to Figure 1 , a video generation method based on AIGC, comprising the following steps:
[0062] S101. Extract the information expressed in the lyrics of the song;
[0063] S102. Generate a story outline for the video according to the information;
[0064] S103. Divide the lyrics into multiple lyric segments, and input the story outline and each lyric segment into an artificial intelligence model to generate the storyline corresponding to each lyric segment;
[0065] S104. Input each lyric segment and the corresponding storyline into the artificial intelligence model to generate the shot sequences corresponding to each lyric segment, and then generate the corresponding images / videos according to each shot sequence;
[0066] S105. Sort and process the images / videos generated by each shot sequence in chronological order to obtain the song video corresponding to the song.
[0067] In some preferred solutions, the information extracted in step S101 may include the main idea of the lyrics and the imagery that appears multiple times in the lyrics (i.e., the main imagery of the lyrics or the song). Among them, the imagery can be a tangible and specific object, such as the imagery being any one or more of the following: scenery, items, characters; the imagery can also be an abstract object without a specific shape, such as the imagery being an intangible object such as a concept (distance, freedom), etc.
[0068] And in step S101, information such as the main idea and imagery can be extracted through an artificial intelligence model. Of course, this application does not exclude extracting the information expressed in the lyrics of the song through manual extraction.
[0069] Extracting lyrics information through an artificial intelligence model includes the steps of: inputting the lyrics into the artificial intelligence model, and obtaining the main ideas and images to be expressed in the lyrics by executing the text summary and information extraction tasks in the natural language processing technology. Among them, the artificial intelligence model includes an NLP (natural language processing) module, so the lyrics of the song are input into the artificial intelligence model, and the artificial intelligence model executes the NLP module to obtain the main ideas and images of the lyrics. The main idea of the lyrics is the core meaning expressed by the lyrics. For example, the main idea of some lyrics is the longing for relatives, such as "Difference" and "Really Love You"; the main idea of some lyrics is the reluctance to lovers, such as "The Moon Represents My Heart" and "Love You for Ten Thousand Years"; the main idea of some lyrics is the yearning for beautiful things such as peace, such as "Peace and Love" and "Imagine". In order to extract the main ideas and images of the lyrics more accurately and quickly, a prompt template can also be made in advance, and the prompt template includes the main ideas and images of various lyrics as examples (referred to as fewshot technology in the field of machine learning), and the prompt template is used to guide the artificial intelligence model to quickly determine the main ideas and images of the lyrics.
[0070] In the above steps S101 to S104, LLM can be used to extract lyric information, generate storylines, and generate storyboards corresponding to the lyrics. Among them, LLM can use ChatGLM2-6B. ChatGLM2-6B is based on the Transformer architecture. After a large amount of corpus training, it can input a piece of text, encode it into the feature space (often called embedding), and output the answer to the input text according to human preferences in the decoder stage. The text of its answer is based on the compressed information of the training corpus and the statistical probability processing of this information. Therefore, ChatGLM2-6B can better complete the tasks of NLP generation and information extraction in steps S101 to S103.
[0071] When implementing steps S101 to S104, a general un-fine-tuned open source ChatGLM2-6B can be used to perform the corresponding tasks, or a model can be fine-tuned and trained for each segmented task.
[0072] As an example and an improvement, ChatGLM2-6B can also be fine-tuned for a specific subtask to obtain higher accuracy. For example, in order to obtain more accurate "key images", the following fine-tuning method can be adopted: First, prepare a number of samples by manually writing samples. A sample is a set of "question-answer pairs", such as "Question: What is the key image of the following lyrics? 'The cold ice rain slapped randomly on my face...'. Answer: Ice rain". Then, train according to the fine-tuning guide of ChatGLM2-6B to obtain the fine-tuned model; finally, when performing the task of "extracting key images", use the fine-tuned model for extraction.
[0073] Similarly, models such as "central idea" model and "storyboard" model can be obtained through fine-tuning.
[0074] In step S102, the "main idea" and "images" of the input song are used, and the LLM executes the NLP text generation task to obtain the "story summary". The story summary is generated based on the main idea and images, and can be understood as the content summary of the song video to be generated. It can formulate the main line of the video content in one sentence (or two to three sentences). For example, the story summary of the video of "Imagine" is: The world is full of various diseases and disasters, and more attention needs to be paid to vulnerable groups to create a better world together. In step S102, when generating the story summary, combine the main idea of the lyrics and the repeatedly appearing images (i.e., key images), and classic plot elements of novels and movies can be applied to design a short story with only two / two groups of characters: for example, the male protagonist recalls the female protagonist while looking at various objects; for example, the story of a hero defeating a monster.
[0075] When generating the corresponding images / videos according to each storyboard in step S104, StableDiffusion can be used. StableDiffusion is an AI painting generation tool. StableDiffusion is trained based on a large number of text-image pairs. It can input a specified text, encode it into the feature space by CLIP (commonly referred to as text embedding). This space has a certain text-image matching ability. Then, in Unet, a denoising process is performed on the pure noise latent, and the attention mechanism is introduced. The text embedding of the specified text guides the denoising direction. Finally, a latent strongly associated with the text is obtained, and a decode action from latent to image is performed on this latent to obtain the target image.
[0076] From the principle of StableDiffusion, it can be seen that inputting the storyboard text into this model can obtain images that match the storyboard.
[0077] At the same time, we can use fine-tuning techniques such as Lora and Textual Inversion to train specified character images. When a specific character appears in the storyboard text, the fine-tuned model is applied to generate stable and consistent character images in different storyboards.
[0078] In step S103, the lyrics can be automatically segmented using an artificial intelligence model or a computer program. When segmenting, the lyrics can be divided into several lyrics segments according to the set number of sentences, for example, every 6 lyrics are divided into a lyrics segment; or the LLM can be used to divide the paragraphs according to the understanding of the lyrics. When segmenting the lyrics, the number of lyrics sentences in each lyrics segment can be equal or unequal. After the lyrics segments are segmented, the corresponding storyline can be generated according to each lyrics segment.
[0079] Among them, the storyline is generated according to the story summary and the corresponding lyrics. Specifically, the story summary and the lyrics segment can be input into the artificial intelligence model, and the artificial intelligence model performs the NLP text generation task (i.e., the text is generated using natural language processing technology) to generate the storyline corresponding to each lyric segment. Among them, the storyline is not only the detailed content of the above story summary, but also closely relies on the lyrics content to generate, so that the generated storyline can fit the emotions and artistic conception to be expressed by the lyrics. When generating the storyline, the basic environment in which the storyline is located can be included. The basic environment can be directly extracted from the lyrics, or it can be associated with the information of the lyrics. For example, the basic environment of "Ningxia" can be formulated as: summer, evening, lakeside, etc. In other embodiments, the basic environment can also be generated when the storyboard corresponding to the lyrics segment is generated in step S104. In step S103, each storyline should be generated according to the story summary and with reference to the content expressed by the lyrics segment, so that each generated storyline follows the design idea of the story summary, and each storyline is linked together, with a time and logical sequence relationship.
[0080] In step S104, the corresponding image / video may be generated according to each shot:
[0081] Using image generation technologies such as Stable Diffusion, corresponding images are generated according to the prompts of each shot (i.e., the prompts are input into the Stable Diffusion intelligent model to generate corresponding images); or using text-to-video technologies such as Imagen, corresponding videos are directly generated according to the prompts of each shot; or based on the Stable Diffusion technology, the already generated images are further used with the Stable Video Diffusion video generation technology to generate videos. Stable Video Diffusion (SVD) is a new AI model launched by the artificial intelligence startup Stability AI. SVD is an open-source video generation model based on Stable Diffusion. SVD supports image-to-video, text-to-video, etc. By adding a temporal generation layer and a temporal attention layer to the Stable Diffusion model, this model enhances the ability to construct the temporal relationship between image frames while retaining excellent image generation capabilities. Given a single input image, this model can use it as the first frame of a video, generate a sequence of latent vectors through a spatio-temporal joint generation network, and finally obtain a video of a specified number of frames and duration through denoising decoding. Auxiliary means such as frame interpolation and speed-up playback can be used to further obtain video segments of the required shot duration.
[0082] In step S104, an artificial intelligence model is used to generate the shots corresponding to each of the lyric segments. The shot is the content expressed in each scene, which is described and expressed by text rather than specific images or videos.
[0083] In one embodiment, the generating corresponding images / videos according to each sub-shot includes: generating a prompt for the corresponding image / video according to the sub-shot; inputting the prompt into an image or video generation intelligent model to generate the images / videos corresponding to each sub-shot. The prompt is the key information of the corresponding image / video of the sub-shot. The sub-shots are generated by an artificial intelligence model according to the story summary, each plot, and each lyric segment. Therefore, the generated sub-shots and prompts can closely match the emotions and artistic conceptions expressed by the lyrics. Among them, the prompts include any two or more of the following: camera movement, characters, character feature description, character action description, location, location scenery description, location item description, lighting atmosphere, and time. In some embodiments, the prompts further include the emotional tone. The emotional tone includes: relief, sadness, hope, loss, loneliness, etc. When generating each sub-shot, an appropriate shot scene can be selected in combination with the emotional tone. Schematically, in the above step S104, the prompts can correspond one by one to the sub-shots, that is, each sub-shot corresponds to at least one set of prompts, and the prompts include the basic environment of the sub-shot. The basic environment includes any one or more of the following: main color, location, time, season, and emotional tone.
[0084] In step S104, input each of the lyric segments and the corresponding plot into the artificial intelligence model to generate the sub-shots corresponding to each lyric segment. One lyric segment generates more than one sub-shot, and the multiple sub-shots corresponding to the same lyric segment maintain the same basic environment. For example, for the lyric segment at the beginning of the lyrics, one lyric segment can generate one sub-shot, while for the climax segments such as the chorus, one lyric segment can generate more than two sub-shots. When one lyric segment generates more than two sub-shots, these sub-shots belonging to the same lyric segment should be based on the same basic environment, that is, the time, location, characters, etc. in these sub-shots are the same or approximately the same. For example, the plot corresponding to a certain lyric segment is a chance encounter story that takes place in a coffee bar. This chance encounter story corresponds to multiple sub-shots, including: "entering the coffee bar", "making eye contact", "having a conversation", "exchanging contact information", etc., and these sub-shots all take place in the coffee bar. Therefore, these sub-shots have the same basic environment: "warm yellow tone, coffee bar, evening, spring, ambiguous".
[0085] In some embodiments, the basic environment can be extended, and the operation of extending the basic environment is also completed by an AI model performing an NLP text generation task. The extension means imagining based on the basic environment, the storyboard, and the plot to enrich the details of the picture as much as possible. For example, describing the items that may appear in the storyboard picture to be generated, the item attributes, the scenery, the lighting atmosphere, etc. In an output example, the basic environment is: main color: gold, location: beach, time: dusk, season: summer, emotional tone: relieved, cheerful, hopeful; the storyboard is: medium shot, a man by the sea, looking at the setting sun in the distance, full of hope; the extension is: a golden beach, a man wearing beach clothes, there are tourists around, in the distance is the setting sun at dusk, the clouds reflecting the sunset glow, the man looking up at the sky, the evening wind in summer blowing his hair and clothes, and his eyes full of hope.
[0086] In step S105, ffmpeg can be used to edit and process the images / videos corresponding to each storyboard to generate a song video. Among them, FFmpeg is an open-source computer program that can be used to record, convert digital audio and video, and convert them into streams. The duration of each storyboard corresponding image / video can be input into ffmpeg, and then the image / video of the storyboard and the audio of the song can be imported to generate the final song video.
[0087] In this embodiment, the production of the song video is divided into multiple steps: lyric information extraction, generating the story summary of the video, generating the storyboard for lyric segments, and generating images / videos from the storyboard, etc. Through the extracted information, the main ideas, images, etc. expressed by the lyrics can be obtained, and the information expressed by the lyrics (i.e., the main ideas and images of the lyrics) is combined in each step, so as to create a video exclusive to the song, making the video content better fit the artistic conception of the lyrics. And in the above technical solution, at least when generating the storyboard and generating the images / videos corresponding to each storyboard, an AI model is used, such as: large language model, image generation model, video generation model, etc., so the efficiency and quality of video production are greatly improved.
[0088] In one embodiment, when extracting the main ideas and images of the information expressed by the song in step S101, the lyrics are input into an AI model, and by performing text summarization and information extraction tasks in natural language processing technology, the main ideas and images to be expressed in the lyrics are obtained. Extracting the main ideas and images of the lyrics through an AI model not only has high efficiency but also fits the lyrics.
[0089] In some embodiments, when generating the storyline in step S103, the emotions expressed in the corresponding lyric segment can be used as an emotional reference and incorporated into the current storyline to avoid simply retelling the content of the lyric segment. And when generating the storyline, each storyline must be limited to the same location, time, season, and weather scene. Therefore, to meet this requirement, it is also possible to divide 1 sentence of the lyrics into 1 lyric segment and generate 1 storyline.
[0090] In one embodiment, when generating a song video, in order to better fit the mood of the lyrics and the song, character control is also introduced. Therefore, in this embodiment, the information extracted from the lyrics in step S101 includes, in addition to the above-mentioned main idea and imagery, a character description; wherein, the character description refers to the description and characterization of the characters and other objects involved in the lyrics, such as the character's gender, age, personality, appearance, etc. Through the character description, the generated storyline can better conform to the mood of the song.
[0091] Generating the story summary of the video according to the information includes: inputting the character description, the main idea, and the imagery into an artificial intelligence model, and the artificial intelligence model performs a text generation task using natural language processing technology to obtain the story summary.
[0092] As Figure 2 shown, after introducing character control, generating the story summary of the video according to the information includes the steps:
[0093] S201. Input the lyrics, and the LLM (i.e., the large language artificial intelligence model) performs the NLP (i.e., natural language processing) information extraction task to obtain the character description;
[0094] S202. Input the gender of the singer of the song, the "character description", the "main idea", and the "imagery", and the LLM performs the NLP text generation task to obtain the "story summary".
[0095] In one embodiment, a 3D video of the song can be generated. In this embodiment, steps S101 to S104 in the above-mentioned embodiment can be followed, but step S105 is improved. Specifically, as Figure 3 shown, the step of inputting the prompt into the image generation intelligent model to generate the images / videos corresponding to the respective sub-shots includes:
[0096] S301. Select 3D scenes and character models that meet the requirements of the prompt from the multimedia database according to the prompt;
[0097] S302. Load the character model into the 3D scene and generate an action sequence for the character model according to the story content in the storyboard or the prompt
[0098] S303. Control the character model to execute the action sequence in the 3D scene and render it into a video clip.
[0099] In step S302, the Story-to-Motion technology can be used to generate a character action sequence. Through the Story-to-Motion technology, according to the story content described in the plot of the storyboard, the character model can be controlled to perform corresponding actions, thus generating a series of action sequences.
[0100] In one embodiment, when the prompt is input into the image generation intelligent model to generate the images / videos corresponding to each storyboard, the scenery or actions in the generated images / videos echo the content of the corresponding lyric segments.
[0101] In this embodiment, the plot can be expressed through multiple storyboards. The storyboards focus on using professional storyboard description terms to describe the images to be formed. In some embodiments, it can also be to set a corresponding storyboard description for each lyric sentence.
[0102] In some embodiments, in order to better set off the emotions and atmosphere to be expressed in each plot, there is more than one atmosphere-setting shot in the storyboards generated corresponding to each lyric segment. The atmosphere-setting shot includes a special effect shot or a close-up shot of an object. That is, more than one atmosphere-setting shot is interspersed in each storyboard. For example, a close-up shot is interspersed in the storyboard to highlight the current scenery or a meaningful object. In this embodiment, the inner feelings of a character can also be expressed in an exaggerated way through a special effect shot. For example, the special effect shot is a very large room with a disproportionately sized little person sitting in it, thus expressing the loneliness in the heart of this little person.
[0103] As Figure 4 shown, in one embodiment, a video generation device 400 based on AIGC is provided. The video generation device 400 based on AIGC includes: an information extraction module 401, a segmentation module 402, a storyboard generation module 403, and a song video generation module 404.
[0104] The information extraction module 401 is used to extract the information expressed in the lyrics of a song and generate a story summary of the video according to the information; wherein, the information includes the main idea of the lyrics and the images that appear repeatedly in the lyrics, and the images include any one or more of scenery, items, concepts, and characters.
[0105] The segmentation module 402 is used to divide the lyrics into multiple lyric segments, input the story summary and each lyric segment into an artificial intelligence model, and generate the corresponding plot for each lyric segment.
[0106] The shot generation module 403 is used to input each lyric segment and the corresponding plot into the artificial intelligence model, generate the shots corresponding to each lyric segment, and then generate the corresponding images / videos according to each shot.
[0107] The song video generation module 404 is used to sort and process the images / videos generated by each shot in chronological order to obtain the song video corresponding to the song.
[0108] In this embodiment, the definitions and explanations of the nouns involved, as well as the production process of the song video, are the same as those in the above embodiment of the song video generation method. Therefore, the definitions and explanations in the above embodiment are also adopted, and the production process of the song video will not be repeated.
[0109] This embodiment divides the production of the song video into multiple steps, such as extracting the main ideas and images of the lyrics to generate the story summary of the video, generating the shots of the lyric segments, and generating images / videos from the shots. And in each step, the main ideas and images of the lyrics are combined to create a video exclusive to the song, so that the video content can better fit the artistic conception of the lyrics. And in the above technical solution, at least when generating the shots and the images / videos corresponding to each shot, an artificial intelligence model is used, so the efficiency and quality of video production are greatly improved.
[0110] As Figure 5 shown, in another embodiment, a computer-readable storage medium 500 is provided, which stores a computer program. When the computer program is run, it executes the AIGC-based video generation method described in any one of the above embodiments.
[0111] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of this application, the patent protection scope of this application cannot be limited thereby. Any technical solution obtained by equivalent structure or equivalent process substitution or modification using the content recorded in the text and drawings of the specification of this application based on the essential concept of this application, as well as any technical solution directly or indirectly implementing the above embodiments in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A video generation method based on AIGC, characterized in that, Including the following steps: Extract the information expressed in the lyrics of the song; Generate a story summary of the video based on the information; Divide the lyrics into multiple lyric segments, and input the story summary and each lyric segment into an artificial intelligence model to generate the plot corresponding to each lyric segment; Input each lyric segment and the corresponding plot into the artificial intelligence model to generate the sub-shots corresponding to each lyric segment, and then generate the corresponding images / videos based on each sub-shot; Sort and process the images / videos generated by each sub-shot in chronological order to obtain the song video corresponding to the song.
2. The video generation method based on AIGC according to claim 1, wherein The information expressed in the lyrics includes the main idea of the lyrics and the images that appear repeatedly in the lyrics. The images include any one or more of: scenery, items, concepts, and characters.
3. The AIGC-based video generation method according to claim 1 or 2, characterized in that, Generating the plot corresponding to each lyric segment further includes: generating the basic environment corresponding to each plot. The basic environment includes any one or more of: main color, location, time, season, and emotional tone.
4. The video generation method based on AIGC according to claim 3, wherein, When inputting each lyric segment and the corresponding plot into the artificial intelligence model to generate the sub-shots corresponding to each lyric segment, one lyric segment generates more than one sub-shot, and the multiple sub-shots corresponding to the same lyric segment maintain the same basic environment.
5. The AIGC-based video generation method according to claim 1 or 4, characterized in that The generating the corresponding images / videos based on each sub-shot includes: generating the prompt for generating the corresponding images / videos according to the sub-shot; inputting the prompt into the image or video generation intelligent model to generate the images / videos corresponding to each sub-shot.
6. The video generation method based on AIGC according to claim 5, wherein, The inputting the prompt into the image or video generation intelligent model to generate the images / videos corresponding to each sub-shot includes: Using the Stable Diffusion text-to-image technology to generate the corresponding images according to the prompt of each sub-shot; Or using the Imagen text-to-video technology to generate the corresponding videos according to the prompt of each sub-shot; Or further generating videos based on the images already generated by the Stable Diffusion technology using the Stable Video Diffusion video generation technology.
7. The video generation method based on AIGC according to claim 5, wherein The inputting the prompt into the image or video generation intelligent model to generate the images / videos corresponding to each sub-shot includes the steps of: Selecting the 3D scenes and character models that meet the requirements of the prompt from the multimedia database according to the prompt; Loading the character model into the 3D scene, and generating the action sequence of the character model according to the story content in the sub-shot or the prompt; Controlling the character model to execute the action sequence in the 3D scene and rendering it into a video clip.
8. The method for generating a video based on AIGC according to claim 2, wherein The information expressed in the lyrics further includes character descriptions; The generating the story summary of the video based on the information includes: inputting two or more of the character descriptions, the main idea, and the images into the artificial intelligence model, and the artificial intelligence model uses natural language processing technology to perform a text generation task to obtain the story summary.
9. The video generation method based on AIGC according to claim 5, wherein, The prompting words include any two or more of the following: camera movement, characters, character feature description, character action description, location, location scenery description, location item description, lighting atmosphere, and time.
10. The video generation method based on AIGC according to claim 1, wherein, In the generation of the sub-shots corresponding to each of the lyric segments, there is more than one atmosphere-enhancing shot, and the atmosphere-enhancing shot includes a special effect shot, a close-up shot of an object, or an empty shot.
11. The video generation method based on AIGC according to claim 1, wherein, The extraction of the information expressed in the lyrics of the song includes the steps of: Inputting the lyrics into an artificial intelligence model, and by performing the tasks of text summarization and information extraction in natural language processing technology, obtaining the main ideas and images to be expressed in the lyrics.
12. A video generation device based on AIGC, characterized in that, Including: An information extraction module for extracting the information expressed in the lyrics of the song and generating a story outline of the video according to the information; A segmentation module for dividing the lyrics into multiple lyric segments, inputting the story outline and each of the lyric segments into an artificial intelligence model, and generating the corresponding plot for each of the lyric segments; A sub-shot generation module for inputting each of the lyric segments and the corresponding plot into the artificial intelligence model, generating the sub-shots corresponding to each of the lyric segments, and then generating the corresponding images / videos according to the sub-shots; A song video generation module for sorting and processing the images / videos generated by each of the sub-shots in chronological order to obtain the song video corresponding to the song.
13. A computer-readable storage medium storing a computer program therein, characterized in that, When the computer program is run, it executes the AIGC-based video generation method according to any one of claims 1 to 11.
Citation Information
Cited By
Advertisement material automatic fission method and system based on artificial intelligence
CN121258610A
Artificial intelligence-based advertisement material automatic fission method and system
CN121258610B