Video generation method and device, equipment and storage medium

By combining dialogue big models and literary big models to generate literary big images, and inputting the video big models or literary big models, the problem that the existing technology cannot meet the diverse needs of users is solved, and the rapid generation of high-quality videos is achieved.

CN119996730APending Publication Date: 2025-05-13NANJING XUANJIA NETWORK TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131629.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art cannot meet the diverse needs of users when generating videos, and it is difficult to achieve high-quality and highly real-time multimodal video generation.

Method used

By obtaining the user-entered video generated key information and the number of video generated clips, combining dialogue big models and literary big models to generate literary big images, and input them into the video big models or literary big models to generate videos.

Benefits of technology

It realizes the rapid generation of high-quality videos while meeting the diverse needs of users, solving the problem that the existing technology cannot meet the diverse needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996730A_ABST
    Figure CN119996730A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and device, equipment and a storage medium. The method comprises the following steps: acquiring video generation key information and video generation fragment quantity input by a user, wherein the video generation key information comprises a video theme text, a sub-shot script text or an image; when the video generation key information comprises a video theme text or a sub-shot script text, generating a corresponding text image based on the video generation key information and the number of video generation clips in combination with a dialogue large model and a text image large model; and inputting the text image or the image into a large graph video model or a text video model to obtain a generated video. According to the method, the video can be quickly generated according to the information input by the user while diversified requirements of the user are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular to a video generation method, device, equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, artificial intelligence (AI) video generation technology has become an important research direction and is widely used in film and television production, advertising creativity, education and training and other fields. AI video generation technology can generate corresponding video content according to user needs or specific scenes, greatly improving the creation efficiency and being able to customize personalized video effects. With the improvement of computing power and the continuous optimization of deep learning algorithms, AI video generation is no longer limited to simple image synthesis, but is developing towards high-quality, real-time multi-modal generation.

[0003] However, most of the current technologies focus on applications in a single field, achieving single tasks in images, text, and videos, but lack comprehensive tasks that can integrate multiple capabilities. Although significant progress has been made in generation tasks in a single field (such as text to video), due to its closed-source nature, it is still challenging to replicate or even surpass the capabilities of advanced systems (such as OpenAI's Sora). Existing open source methods find it difficult to achieve comparable performance and are often hindered by insufficient training data quality, which results in an inability to improve the results of comprehensive tasks (such as video generation). Therefore, how to design a method that can both meet user needs and improve the results of comprehensive tasks (such as video generation, etc.) has become a technical challenge. Summary of the invention

[0004] The present invention provides a video generation method, device, equipment and storage medium to solve the problem that the prior art cannot meet the diverse needs of users when generating videos.

[0005] According to one aspect of the present invention, a video generation method is provided, the method comprising:

[0006] Acquire key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image;

[0007] When the video generation key information includes a video theme text or a storyboard script text, based on the video generation key information and the number of video generation segments, a corresponding Vincent image is generated in combination with a dialogue macro model and a Vincent image macro model;

[0008] The generated image or the image is input into a large image-generated video model or a large image-generated video model to obtain a generated video.

[0009] According to another aspect of the present invention, a video generation device is provided, the device comprising:

[0010] An acquisition module, used to acquire key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image;

[0011] A processing module, for generating a corresponding Vincent image based on the video generation key information and the number of video generation segments in combination with a dialogue macro model and a Vincent image macro model when the video generation key information includes a video theme text or a storyboard script text;

[0012] A generation module is used to input the Wensheng image or the image into a large image-generated video model or a Wensheng video model to obtain a generated video.

[0013] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising: at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video generating method described in any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video generation method described in any embodiment of the present invention when executed.

[0017] A video generation method, device, equipment and storage medium according to an embodiment of the present invention, the method comprises: obtaining key video generation information and the number of video generation segments input by a user, wherein the key video generation information comprises a video theme text, a storyboard script text or an image; when the key video generation information comprises a video theme text or a storyboard script text, based on the key video generation information and the number of video generation segments, a corresponding Vincent image is generated in combination with a dialogue large model and a Vincent graph large model; the Vincent image or the image is input into a graph video large model or a Vincent video model to obtain a generated video. The method can quickly generate a video according to the information input by the user while meeting the diverse needs of the user, thus solving the problem that the prior art cannot meet the diverse needs of the user when generating a video.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 A schematic diagram of a flow chart of a video generation method provided in Embodiment 1 of the present invention;

[0021] Figure 2 A schematic diagram of a flow chart of a video generation method provided by an embodiment of the present invention;

[0022] Figure 3 A schematic diagram of a flow chart of a video generation method provided in Embodiment 2 of the present invention;

[0023] Figure 4 A schematic diagram of the structure of a video generating device provided in Embodiment 3 of the present invention;

[0024] Figure 5 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only an embodiment of a part of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should belong to the scope of protection of the present invention. It should be understood that the various steps recorded in the method implementation of the present invention can be performed in different orders and / or in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0026] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, any variations of the terms "including" and "having" are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.

[0030] Embodiment 1

[0031] Figure 1 A flow chart of a video generation method provided in Embodiment 1 of the present invention is applicable to the case of generating a video based on text or images input by a user. The method can be executed by a video generation device, wherein the device can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes but is not limited to: computers and other devices.

[0032] like Figure 1 As shown, a video generation method provided by Embodiment 1 of the present invention includes the following steps:

[0033] S110, obtaining key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image.

[0034] The key information for video generation may include a video theme text, a storyboard script text, or an image. The video theme text may be a piece of text input by the user, the storyboard script text may be a manuscript input by the user that describes the video to be generated in units of shots, and the image may be a picture input by the user. The number of video generation segments may be the number of segments of the video that the user wishes to generate.

[0035] In this embodiment, the key information for video generation and the number of video generation segments input by the user can be obtained. The key information for video generation input by the user can include video theme text, storyboard script text or image. The user can input corresponding types of data according to their needs, such as generating videos by theme, generating videos by script editing, generating videos by images, and specifying the number of video generation segments as the number of generated videos for subsequent text / image generation videos.

[0036] S120. When the video generation key information includes a video theme text or a storyboard script text, a corresponding Vincent image is generated based on the video generation key information and the number of video generation segments in combination with a dialogue macro model and a Vincent image macro model.

[0037] Among them, the dialogue model can be an artificial intelligence model based on deep learning, which is mainly used to generate natural and fluent dialogue text. By training a large amount of dialogue text data, it learns language patterns, semantic understanding and dialogue strategies, and can generate reasonable responses in a given dialogue scenario. The text-generated image model can be an artificial intelligence model that can generate images based on input text descriptions. It combines natural language processing and computer vision technology to convert semantic information in text into visual image information. Text-generated images can be images generated based on text.

[0038] In this embodiment, when the key information of video generation includes video theme text or storyboard script text, it is necessary to obtain the corresponding text image based on the video theme text or storyboard script text and the number of video generation segments, combined with the dialogue model and the text image model. Figure 2 A schematic diagram of a video generation method provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, when a video is generated by text or script editing, the text input by the user can be expanded through a large dialogue model, and when a video is generated by an image, the image can be preprocessed.

[0039] The dialogue model of this embodiment can be selected at will (external model, open source model, local model), which is determined by the model's instruction-following ability and the richness of expansion. For localized deployment, if the model's instruction-following ability is weak, the instruction-following effect can be improved by fine-tuning the model, thereby improving the quality of text output. The specific method is as follows:

[0040] For scenarios where user text input data is used, a large model-assisted or manually constructed method is used to create a question-answer format dataset suitable for large model training.

[0041] Using a large model to assist in construction can take advantage of the natural language processing capabilities of the large model to automatically generate question-answer pairs in the required format. This method requires selecting a suitable pre-trained general domain large model (such as Qwen-7B-Chat, etc.) and setting specific prompts to automatically generate question-answer pairs that better adapt to scenario tasks. When building a question-answering dataset, the format and content of the question-answer pairs need to be designed according to the training objectives and expected effects. In this scenario, the training goal is to expand and split the user input topic, and output it on the large model side in a specific format (such as json format or standardized format). Therefore, the structure of the question-answer pair needs to meet this goal, so that when the model is trained, the input prompt is the user input, and the content that the model needs to generate is the split-shot content. Such question-answer pairs can help the model learn and improve the model response effect. Manual construction requires professionals to collect data for preprocessing and annotation, and batch build question-answer pairs according to scenario requirements.

[0042] In one embodiment, the method of generating corresponding Vincent images based on the video-generated key information and the number of video-generated segments in combination with a dialogue macro model and a Vincent diagram macro model includes: inputting the number of video-generated segments and a video theme text or a storyboard script text into the dialogue macro model to obtain segmentation prompt words corresponding to a plurality of video segments; and inputting the segmentation prompt words corresponding to a plurality of video segments into the Vincent diagram macro model to obtain corresponding Vincent images.

[0043] The segmentation prompt words may be a type of words or phrases used to divide paragraphs or guide the text to be generated according to the paragraph structure during the text processing process.

[0044] In this embodiment, when the key information of video generation includes video theme text or shot-by-shot script text, the number of video generation segments and the video theme text or shot-by-shot script text can be input into the dialogue macro model to obtain segmentation prompt words corresponding to multiple video segments, and the number of segmentation prompt words is the same as the number of video generation segments. After obtaining the segmentation prompt words, the segmentation prompt words can be input into the text-generated image macro model to obtain multiple corresponding text images through the text-generated image macro model.

[0045] Exemplarily, when the key information generated by the video input by the user is the video theme text, the video theme text input by the user can be expanded and processed by the dialogue big model according to the refined specific prompt words (prompt) and dialogue memory, and output on the big model side in a specific format (such as json format or standardized format), and the text content of each shot is segmented using a format processing script, and finally the segmented prompt words of the user-specified video length are output and temporarily stored for use in the next step.

[0046] When the key information generated by the video input by the user is a storyboard script text, the user-input storyboard script text can be expanded separately according to the fine-tuned specific prompt words (prompt) using the dialogue big model, and output on the big model side in a specific format (such as json format or standardized format), and finally the segmented prompt words of the user-specified video length are output and temporarily stored for use in the next step.

[0047] After obtaining the segmented prompt words, the segmented prompt words of the video length specified by the user can be used to generate images using the text-based image model according to the input window of the image-based video model: the text-based image model can be selected, and different operations can be performed according to the selected model category. For example, when the text-based image model supports Chinese and English bilingual models (most external models, such as Doubao, etc.), there is no need to process the segmented prompt words; when the text-based image model only supports English models (most open source models, such as Stable Diffusion, etc.), it is necessary to use the dialogue model to use a specific prompt (for example, it needs to be suitable for the Stable Diffusion style) to translate and modify the segmented prompt words. Finally, the processed segmented prompt words can be sent to the selected text-based image model separately or in parallel for image generation, and the generated images are stored in the order of the storyboards.

[0048] S130, inputting the text-generated image or the image into a large image-generated video model or a text-generated video model to obtain a generated video.

[0049] The image-generated video model can be an artificial intelligence model that can generate a video based on input image-related information. The text-generated video model can be an artificial intelligence model that generates a video based on input text descriptions.

[0050] In this embodiment, if the user inputs an image, the image can be directly input into the image-generated video model. If the user inputs a video theme text or a storyboard script text, the obtained text-generated image can be input into the image-generated video model or the text-generated video model to obtain the generated video. In this embodiment, when the key information for generating video input by the user is an image, the user input image can be pre-processed, cropped, scaled, etc., to conform to the input window of the image-generated video model.

[0051] A video generation method provided by the first embodiment of the present invention includes: obtaining key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image; when the key video generation information includes a video theme text or a storyboard script text, based on the key video generation information and the number of video generation segments, a corresponding Vincent image is generated in combination with a dialogue large model and a Vincent graph large model; the Vincent image or the image is input into a graph video large model or a Vincent video model to obtain a generated video. The method can quickly generate a video based on the information input by the user while meeting the diverse needs of the user, thus solving the problem that the prior art cannot meet the diverse needs of the user when generating a video.

[0052] Based on the above embodiment, a variant embodiment of the above embodiment is proposed. It should be noted that in order to make the description concise, only the differences from the above embodiment are described in the variant embodiment.

[0053] In one embodiment, when the user has a requirement for role consistency and provides a role reference picture, the segmented prompt words corresponding to multiple video clips are input into a large model of the Vincent picture to obtain corresponding Vincent images, including: inputting the role reference picture and the segmented prompt words corresponding to multiple video clips into the large model of the Vincent picture to obtain a Vincent image with role consistency.

[0054] The character consistency requirement may refer to the requirement of maintaining the consistency of the character features, so that the appearance, style, and features of the character in the generated new image are highly consistent with the original image. The character reference image may be an image related to the character provided by the user.

[0055] In this embodiment, when the user has a requirement for role consistency and the user provides a role reference map, the role reference map and the segmented prompt words corresponding to multiple video clips can be input into the Vincent map large model to obtain a Vincent image with role consistency, that is, using the role / character reference map additionally provided by the user, image generation based on the reference map is implemented in the Vincent map step, thereby maintaining role / character consistency.

[0056] In one embodiment, before inputting the Vincent image or the image into the Vincent video large model or the Vincent video model, the method further includes: when the user has a role consistency requirement, generating a Vincent image with role consistency based on the Vincent image through a pre-trained Vincent image large model; wherein the pre-trained Vincent image large model is trained using a role dataset and a low-rank adaptive method.

[0057] Among them, Low-Rank Adaptation (LoRA) is a training method based on a large model. By connecting the output layer of the large model with a small neural network (ie, the LoRA model), the capabilities of the large model can be used to perform training for specific tasks.

[0058] In this embodiment, after the image is generated, if the user has no special requirements for the role, the image can be enhanced with details and super-resolution to improve the image quality. If the user has consistency requirements for roles, characters, etc., and the user does not provide a reference image, the role data set can be connected to the Wenshengtu large model training in advance, and the LoRA method is used to select appropriate parameters for training according to the amount of training set data. After observing the loss and model convergence after training, the step is regenerated to obtain the required image containing the corresponding role, IP, etc. LoRA training needs to select appropriate training parameters according to the amount of training set data, such as learning rate, batch size, number of training rounds, etc. After the training is completed, it is necessary to observe the loss and model convergence after training. Loss refers to the error value of the model during the training process, which reflects the degree of fit of the model to the training set. If the loss value is small, it means that the model fits the training set well; if the loss value is large, it means that the model fits the training set poorly. At the same time, we also need to observe the convergence of the model. If the loss value of the model gradually decreases during the training process, it means that the model is gradually learning the characteristics of the training set; if the loss value of the model fluctuates or does not decrease during the training process, it means that the model may be overfitting or underfitting.

[0059] By selecting appropriate parameters for training and observing the loss and model convergence after training, we can determine whether the training effect of the model is optimal.

[0060] Embodiment 2

[0061] Figure 3 This is a flow chart of a video generation method provided by Embodiment 2 of the present invention. Embodiment 2 is optimized on the basis of the above embodiments. For details not yet fully described in this embodiment, please refer to Embodiment 1.

[0062] like Figure 3 As shown, a video generation method provided by Embodiment 2 of the present invention includes the following steps:

[0063] S210, obtaining key video generation information and the number of video generation segments input by the user, wherein the key video generation information includes a video theme text, a storyboard script text or an image.

[0064] S220. When the video generation key information includes a video theme text or a storyboard script text, a corresponding Vincent image is generated based on the video generation key information and the number of video generation segments in combination with a dialogue macro model and a Vincent image macro model.

[0065] S230: When the image to be processed is an image, or the image to be processed is a Vincent image and the user does not specify a Vincent video model, the image or the Vincent image is input into the image-generated video model to obtain a generated video.

[0066] S240: When the image to be processed is a Vincent image and the user has specified a Vincent video model, the Vincent image and the segmentation prompt words corresponding to the video theme text or the storyboard script text are input into the Vincent video model to obtain a generated video.

[0067] In this embodiment, the generated Vincent image or image can be used to perform the image-generated video pipeline to generate each video segment separately. The images of the storyboards can be sent separately or in parallel to the image-generated video model pipeline for video generation, and stored in the original order after processing. Alternatively, if the image to be processed is a Vincent image, and the user specifies that a Vincent video model needs to be used, the Vincent video pipeline with a text reference image is used for generation. At this time, the segmentation prompt words corresponding to the storyboard script text need to be used as the Vincent video pipeline input. If there is a requirement for the continuity of the video (video extension), the last frame of the first video segment can be used as the starting frame of the second video segment for video generation. At this time, the image generated by the original storyboard will be ignored.

[0068] A video generation method provided in the second embodiment of the present invention specifically inputs the Vincent image or the image into a large image-generated video model or a Vincent video model to obtain a generated video, including: when the image to be processed is an image, or the image to be processed is a Vincent image and the user has not specified a Vincent video model, the image or the Vincent image is input into a large image-generated video model to obtain a generated video; when the image to be processed is a Vincent image and the user has specified a Vincent video model, the Vincent image and the segmented prompt words corresponding to the video theme text or the storyboard script text are input into the Vincent video model to obtain a generated video. The method generates key information for different types of videos input by the user, can generate videos in different ways, and can improve the accuracy of video generation.

[0069] In one embodiment, after obtaining the generated video, the method further includes: interpolating frames of the video through a frame interpolation algorithm to obtain a video after interpolation; for each video segment in the video after interpolation, performing a smooth transition of intermediate frames in the order of each video segment through an audio and video codec tool to obtain a transitioned video; super-resolving the video after transition according to the scene type of the video to obtain a super-resolved video; and composing music for the super-resolved video through a music library or a large music generation model to obtain a processed video.

[0070] Among them, the interpolation algorithm can be a computer vision algorithm used to insert additional frames into the video to improve the smoothness and viewing experience of the video. The audio and video codec tool can be a type of software or hardware device used to process audio and video signals. The scene types of the video can include general scenes, animation scenes or other scenes. The music library can be a collection of music resources, which include various music works, such as songs, music, background music, etc., in the form of digital audio files. The music generation model can be an artificial intelligence model that can automatically generate new music works by learning a large amount of music data.

[0071] In this embodiment, for the generated video, the video can be interpolated by a frame insertion algorithm, and for each video segment of the interpolated video, the intermediate frames can be smoothly transitioned in the order of each video segment through the audio and video codec tool. For the video after the transition, the video can be super-resolved according to the scene type of the video, and then the music library or the music generation model can be used to compose music for the super-resolved video to obtain the final video. When processing the video itself, this embodiment can use video editing enhancement technology to enhance the features of each dimension of the video, including the number of frames and resolution, and use the interpolation algorithm / multi-script control / super-resolution to process the video to achieve high-quality generation. At the same time, since a pre-trained model is used, there is no need to train a new model additionally, and only the existing model needs to be used for processing. The system also provides scalability, and the latest technology can be used for post-processing of model updates.

[0072] For example, the general image-generated video model is limited by various factors (resources, hardware, computing power, etc.), and the output video frame rate is not enough (for example, stable video diffusion, each time generating 4s video with 6 frames per second, a total of 25 frames), then the interpolation algorithm is first used to fill the frame to 24 frames per second, so that the generated video picture is smoother. For each video segment processed by interpolation, it is merged according to the theme or script order. At the interval between each video segment, the intermediate transition frame is generated by a script based on Fast Forward Mpeg (ffmpeg for short), and after the smooth transition in the middle of the video, ffmpeg is used to merge the segmented video into a whole segment for storage. Among them, ffmpeg is an open source multimedia processing tool. The video resolution output by the image-generated video model is usually between 540p-720p (for example, stable video diffusion, 1024*576), and the resolution is relatively low. For the video after smooth transition, super-resolution can be performed according to demand to improve the video resolution and video effect. The super-resolution algorithm also involves various scenes, such as general super-resolution and animation super-resolution, and the appropriate algorithm can be intelligently selected for matching according to the specific video content.

[0073] In one embodiment, the method of composing music for the super-resolved video through a music library or a music generation model to obtain a processed video includes: matching the text content of the video with the names of music materials in the music library through a text model, and using the music material with the highest similarity to the text content among all music materials as background music; or, inputting the theme content of the video into the music generation model to obtain background music with the same length as the super-resolved video; segmenting the background music, and merging the segmented background music with each video clip of the super-resolved video through an audio and video codec tool to perform video and audio track merging.

[0074] The text big model may be an artificial intelligence model based on deep learning technology, mainly used to process tasks related to natural language text. The music material name may be the name of the music material.

[0075] In this embodiment, the video can be matched with music in two ways. For example, the text content of the video is matched with the name of the music material in the music library through the text big model, and the music material with the highest similarity to the text content among all the music materials is used as the background sound; or the theme content of the video is input into the music generation big model, and the background sound with the same length as the super-resolved video is generated through the music generation big model. Finally, the background sound is segmented, and the segmented background sound is merged with the video clips of the super-resolved video through the audio and video codec tool to perform video and audio track merging.

[0076] For example, since the video output by the image-generated video model has no sound, you can add music to the video as needed. The following two methods are provided:

[0077] Use the music material library prepared in advance, use the text big model to match the text content with the music material name, select the music material with high similarity, use the multimedia stream analysis tool (ffprobe) to obtain the video length, split the music material into the same length, and use ffmpeg to merge the video and the split music material into video and audio tracks to complete the video soundtrack processing.

[0078] Use the theme content as input and use the music generation model to generate it. After matching the corresponding length of the video, use ffmpeg to merge the video and the segmented music material into video and audio tracks to complete the video soundtrack processing.

[0079] The video generation method of this embodiment combines the arrangement and free combination of multiple tasks, covering text-to-text generation, text-to-image generation, text-to-video generation, text-conditional image-to-video generation, text-to-audio generation, expanding the generated video, video-to-video editing, connecting videos, etc., improving the generated video effect and realizing video customization.

[0080] This embodiment adopts a multi-agent (AI Agent) system framework and uses multiple modal models to collaborate on tasks including video generation. Not only can the multi-agent part collaborate and combine, but the model also supports customized training. Through model training within agent tasks such as text-to-text, text-to-image, and text-to-video, the performance and quality of video generation are further optimized.

[0081] Embodiment 3

[0082] Figure 4 This is a structural schematic diagram of a video generating device provided in Embodiment 3 of the present invention. The device may be applicable to situations where a video is generated based on text or images input by a user, wherein the device may be implemented by software and / or hardware and is generally integrated on an electronic device.

[0083] like Figure 4 As shown, the device comprises:

[0084] An acquisition module 310 is used to acquire key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image;

[0085] A processing module 320, configured to generate a corresponding Vincent image based on the video generation key information and the number of video generation segments in combination with a dialogue macro model and a Vincent image macro model when the video generation key information includes a video theme text or a storyboard script text;

[0086] The generation module 330 is used to input the Vincent image or the image into the image-generated video model or the Vincent video model to obtain a generated video.

[0087] This embodiment provides a video generation device, including: an acquisition module, used to acquire key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image; a processing module, used to generate a corresponding Vincent image based on the key video generation information and the number of video generation segments when the key video generation information includes a video theme text or a storyboard script text, in combination with a dialogue large model and a Vincent graph large model; a generation module, used to input the Vincent image or the image into a graph video large model or a Vincent video model to obtain a generated video. The device can quickly generate a video based on the information input by the user while meeting the diverse needs of the user, solving the problem that the prior art cannot meet the diverse needs of the user when generating a video.

[0088] Furthermore, the processing module 320 is specifically configured to:

[0089] Input the number of video generated segments and the video theme text or shot-by-shot script text into the dialogue model to obtain segmentation prompt words corresponding to the multiple video segments;

[0090] The segmentation prompt words corresponding to the multiple video clips are input into the Wensheng image large model to obtain the corresponding Wensheng images.

[0091] Further, when the user has a role consistency requirement and provides a role reference image, the segmented prompt words corresponding to the plurality of video clips are input into the Wensheng image large model to obtain the corresponding Wensheng image, including:

[0092] The character reference image and the segmented prompt words corresponding to the plurality of video clips are input into the Vincent image large model to obtain a Vincent image with character consistency.

[0093] Furthermore, before inputting the Vincent image or the image into the image-generated video model or the Vincent video model, the device is further used to:

[0094] When the user has a role consistency requirement, a pre-trained Vincent image large model is used to generate a Vincent image with role consistency based on the Vincent image;

[0095] Among them, the pre-trained Wensheng graph model is trained through the role dataset and low-rank adaptive method.

[0096] Furthermore, the generating module 330 is specifically used for:

[0097] When the image to be processed is an image, or the image to be processed is a Vincent image and the user does not specify a Vincent video model, the image or the Vincent image is input into the image-generated video model to obtain a generated video;

[0098] When the image to be processed is a Vincent image and the user has specified a Vincent video model, the Vincent image and the segmented prompt words corresponding to the video theme text or the storyboard script text are input into the Vincent video model to obtain a generated video.

[0099] Furthermore, after obtaining the generated video, the device is also used to:

[0100] Inserting frames into the video by using a frame insertion algorithm to obtain a frame-inserted video;

[0101] For each video segment in the video after frame interpolation, a smooth transition of intermediate frames is performed according to the order of each video segment by an audio and video codec tool to obtain a video after transition;

[0102] Performing video super-resolution on the transitioned video according to the scene type of the video to obtain a super-resolution video;

[0103] The super-resolved video is matched with music through a music library or a large music generation model to obtain a processed video.

[0104] Furthermore, the process of arranging music for the super-resolved video through a music library or a music generation model to obtain a processed video includes:

[0105] Match the text content of the video with the names of the music materials in the music library through the text big model, and use the music material with the highest similarity to the text content as the background sound; or input the theme content of the video into the music generation big model to obtain the background sound with the same length as the super-resolved video;

[0106] The background sound is segmented, and the segmented background sound is merged with the video and audio tracks of each video segment of the super-resolved video through an audio and video encoding and decoding tool.

[0107] The above-mentioned video generating device can execute the video generating method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the executing method.

[0108] Embodiment 4

[0109] Figure 5A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0110] like Figure 5 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0111] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0112] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as a video generation method.

[0113] In some embodiments, the video generation method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the video generation method in any other appropriate manner (e.g., by means of firmware).

[0114] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0115] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0116] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0118] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0119] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0120] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0121] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A video generation method, characterized in that: The method comprises: Acquire key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image; When the video generation key information includes a video theme text or a storyboard script text, based on the video generation key information and the number of video generation segments, a corresponding Vincent image is generated in combination with a dialogue macro model and a Vincent image macro model; The generated image or the image is input into a large image-generated video model or a large image-generated video model to obtain a generated video.

2. The method according to claim 1, characterized in that The generating a corresponding Vincent image based on the video generation key information and the number of video generation segments, combined with the dialogue large model and the Vincent image large model, includes: Input the number of video generated segments and the video theme text or shot-by-shot script text into the dialogue model to obtain segmentation prompt words corresponding to the multiple video segments; The segmentation prompt words corresponding to the multiple video clips are input into the Wensheng image large model to obtain the corresponding Wensheng images.

3. The method according to claim 2, characterized in that When the user has a role consistency requirement and provides a role reference image, the segmented prompt words corresponding to the plurality of video clips are input into the Wensheng image large model to obtain the corresponding Wensheng image, including: The character reference image and the segmented prompt words corresponding to the plurality of video clips are input into the Vincent image large model to obtain a Vincent image with character consistency.

4. The method according to claim 1, characterized in that: Before inputting the Vincent image or the image into the image-generated video model or the Vincent video model, the method further comprises: When the user has a role consistency requirement, a pre-trained Vincent image large model is used to generate a Vincent image with role consistency based on the Vincent image; Among them, the pre-trained Wensheng graph model is trained through the role dataset and low-rank adaptive method.

5. The method according to claim 1, characterized in that The step of inputting the Vincent image or the image into a large image-generated video model or a Vincent video model to obtain a generated video includes: When the image to be processed is an image, or the image to be processed is a Vincent image and the user does not specify a Vincent video model, the image or the Vincent image is input into the image-generated video model to obtain a generated video; When the image to be processed is a Vincent image and the user has specified a Vincent video model, the Vincent image and the segmented prompt words corresponding to the video theme text or the storyboard script text are input into the Vincent video model to obtain a generated video.

6. The method according to claim 1, characterized in that After obtaining the generated video, the method further includes: Inserting frames into the video by using a frame insertion algorithm to obtain a frame-inserted video; For each video segment in the video after frame interpolation, a smooth transition of intermediate frames is performed according to the order of each video segment by an audio and video codec tool to obtain a video after transition; Performing video super-resolution on the transitioned video according to the scene type of the video to obtain a super-resolution video; The super-resolved video is matched with music through a music library or a large music generation model to obtain a processed video.

7. The method according to claim 6, characterized in that The step of using a music library or a music generation model to compose music for the super-resolved video to obtain a processed video includes: Match the text content of the video with the names of the music materials in the music library through the text big model, and use the music material with the highest similarity to the text content as the background sound; or input the theme content of the video into the music generation big model to obtain the background sound with the same length as the super-resolved video; The background sound is segmented, and the segmented background sound is merged with the video and audio tracks of each video segment of the super-resolved video through an audio and video encoding and decoding tool.

8. A video generating device, characterized in that: The device comprises: An acquisition module, used to acquire key video generation information and the number of video generation segments input by a user, wherein the key video generation information includes a video theme text, a storyboard script text or an image; A processing module, for generating a corresponding Vincent image based on the video generation key information and the number of video generation segments in combination with a dialogue macro model and a Vincent image macro model when the video generation key information includes a video theme text or a storyboard script text; A generation module is used to input the Wensheng image or the image into a large image-generated video model or a Wensheng video model to obtain a generated video.

9. An electronic device, characterized in that: The device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the video generating method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video generation method according to any one of claims 1 to 7 when executed.