Long video generation method, device, equipment and storage medium

By generating target text through multiple rounds of interaction between the original text and a large language model, and combining the image and sub-text of the main object to generate audio and video, and then splicing them together, the problems of incoherence and high cost of long video generation in the existing technology are solved, and high-quality and low-cost long video generation is achieved.

CN117768746BActive Publication Date: 2025-10-03SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311863627.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-10-03
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

Existing long video generation methods require extensive training on large, well-annotated long video datasets. The generated long video content is incoherent and costly, and cannot meet users' requirements for video length.

Method used

Based on multiple rounds of interaction between the original text and the large language model, the target text is generated, including sub-texts and durations of multiple scenes. The image of the main object is generated and the main object corresponding to each scene is determined. Audio and video are generated based on the image and/or sub-text, and the audio and video of each scene are spliced ​​together.

Benefits of technology

The quality of generated long videos is improved and the cost of generating long videos is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117768746B_ABST
    Figure CN117768746B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention provide a method, apparatus, device and storage medium for generating long videos. The method comprises: performing multiple rounds of interactions with a large language model based on the original text to obtain a target text; wherein the target text includes subtexts of multiple scenes and the duration of each scene; generating an image of at least one main object based on the target text, and determining the main object corresponding to each scene; for each scene, generating the audio and video corresponding to the scene based on the image of the main object and / or the subtext; splicing the audio and video of each scene to obtain a target long video. The method for generating long videos provided by the embodiments of the present invention generates target long videos based on the main object, image of the main object and subtext of each scene, which can improve the quality of the generated long videos and reduce the cost of generating long videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of multimedia processing technology, and in particular to a method, apparatus, device and storage medium for generating a long video. Background Art

[0002] Current video generation methods primarily generate short videos of a few seconds by training image generation models. However, short videos cannot meet user requirements for length. Existing methods for generating long videos require extensive training on large, well-annotated long video datasets. The resulting long videos are incoherent and require manual design and processing of complex text descriptions, which is costly. Summary of the Invention

[0003] Embodiments of the present invention provide a method, apparatus, device, and storage medium for generating a long video, which can improve the quality of the generated long video and reduce the cost of generating the long video.

[0004] In a first aspect, an embodiment of the present invention provides a method for generating a long video, comprising:

[0005] Based on the original text and the large language model, multiple rounds of interaction are carried out to obtain the target text; wherein the target text includes sub-texts of multiple scenes and the duration of each scene;

[0006] generating an image of at least one subject object based on the target text, and determining the subject object corresponding to each scene;

[0007] For each scene, generating audio and video corresponding to the scene based on the image of the subject object and / or the subtext;

[0008] The audio and video of each scene are spliced ​​together to obtain the target long video.

[0009] In a second aspect, an embodiment of the present invention further provides a device for generating a long video, comprising:

[0010] A target text acquisition module is used to perform multiple rounds of interaction with the large language model based on the original text to obtain the target text; wherein the target text includes sub-texts of multiple scenes and the duration of each scene;

[0011] A subject object image generation module, configured to generate an image of at least one subject object based on the target text and determine the subject object corresponding to each scene;

[0012] An audio and video generation module, configured to generate, for each scene, audio and video corresponding to the scene based on the image of the subject object and / or the subtext;

[0013] The target long video acquisition module is used to splice the audio and video of each scene to obtain the target long video.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for generating a long video described in an embodiment of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for generating a long video described in an embodiment of the present invention when executed.

[0019] The embodiments of the present invention provide a method, apparatus, device and storage medium for generating long videos. Based on the original text, multiple rounds of interaction are performed with a large language model to obtain a target text; wherein the target text includes subtexts of multiple scenes and the duration of each scene; an image of at least one main object is generated based on the target text, and the main object corresponding to each scene is determined; for each scene, the audio and video corresponding to the scene is generated according to the image and / or subtext of the main object; the audio and video of each scene are spliced ​​to obtain a target long video. The long video generation method provided by the embodiment of the present invention generates a target long video based on the main object, image and subtext of each scene, which can improve the quality of the generated long video and reduce the cost of generating the long video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flowchart of a method for generating a long video in Example 1 of the present invention;

[0021] Figure 2 This is an example diagram of a target text in the first embodiment of the present invention;

[0022] Figure 3 This is an example diagram of generating a target text in the first embodiment of the present invention;

[0023] Figure 4 This is an example diagram of generating an image of a subject object in the first embodiment of the present invention;

[0024] Figure 5 This is an example diagram of determining the main object corresponding to each scene in the first embodiment of the present invention;

[0025] Figure 6 This is an example diagram of generating a video in the first embodiment of the present invention;

[0026] Figure 7 This is an example diagram of a target long video spliced ​​together in the first embodiment of the present invention;

[0027] Figure 8 This is a schematic structural diagram of a device for generating a long video in the second embodiment of the present invention;

[0028] Figure 9 It is a structural diagram of an electronic device in embodiment 3 of the present invention. DETAILED DESCRIPTION

[0029] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0030] Example 1

[0031] Figure 1 This is a flowchart of a method for generating a long video provided in Example 1 of the present invention. This embodiment is applicable to the case of generating a long video based on text. The method can be executed by a long video generation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC, or server. Specifically, it includes the following steps:

[0032] S110, performing multiple rounds of interactions with the large language model based on the original text to obtain the target text.

[0033] The target text includes subtext for multiple scenes and the duration of each scene. The original text can be a user's simple description of a specific situation, such as a story. For example, the original text can be as follows: "One day... A was sitting in the airport terminal... A arrived at a famous tourist attraction..." The large language model can be a deep learning model used to generate text. In this embodiment, any existing open source large language model can be used, without limitation.

[0034] The target text can be represented in the form of a script, including multiple sub-texts of consecutive scenes and the duration of each scene. Figure 2 is an example diagram of a target text in this embodiment, such as Figure 2 As shown, the target text includes S1, S2, S3, S4..., S nscenes, and each scene includes its corresponding subtext and duration.

[0035] Specifically, based on multiple rounds of interactions between the original text and the large language model, the target text can be obtained in the following manner: for the first round of interaction with the large language model, the original text and the prompt text corresponding to the first round of interaction are input into the large language model, and the intermediate text is output; for non-first round interactions with the large language model, the intermediate text output from the previous round of interaction and the prompt text corresponding to the current round of interaction are input into the large language model, and the intermediate text or target text corresponding to the current round of interaction is output.

[0036] Among them, non-first-round interaction includes: middle-round interaction or last-round interaction. Among them, the prompt text corresponding to the first-round interaction can be a sentence used to prompt the large language model to adjust or improve the original text. The prompt text corresponding to the non-first-round interaction can be a sentence used to prompt the large language model to adjust or improve the intermediate text. In this embodiment, the original text and the prompt text are first input into the large language model, and an intermediate text is output (which can be a text adjusted from the original text according to the prompt text); then the intermediate text and the prompt text are input into the large language model again, and an intermediate text is output (which can be a text adjusted from the input intermediate text according to the prompt text). The prompt text this time can have a logical progressive relationship with the prompt text input last time; and so on. After multiple rounds of interaction with the large language model, the final target text is output.

[0037] For example, Figure 3 is an example diagram of generating target text in this embodiment, such as Figure 3 As shown in the figure, taking the target text as a script, four rounds of interaction with the large language model are performed. In the first round of interaction, the original text (story description) and the prompt text (Please convert this story into a rough script) are input into the large language model, and the rough script is output, i.e., the intermediate text. In the second round of interaction, the intermediate text (rough script) and the prompt text (This is a rough script, please improve it) are input into the large language model, and the detailed script is output, i.e., the intermediate text. In the third round of interaction, the intermediate text (detailed script) and the prompt text (This is a detailed script, please check if there are any missing plot points) are input into the large language model, and the complete script is output, i.e., the intermediate text. In the final round of interaction, the intermediate text (complete script) and the prompt text (This is a complete script, please set the duration of each scene) are input into the large language model, and the final script is output, i.e., the target text.

[0038] S120: Generate an image of at least one main object based on the target text, and determine the main object corresponding to each scene.

[0039] The target text may include one or more subject objects. If the target text is a script, the subject object may be understood as the main actors in the script. The image of the subject object may be an image that can reflect the image of the subject object. For example, assuming that the subject object is a cat, the image may be an image of a cat.

[0040] Specifically, the method of generating an image of at least one subject object based on the target text can be: inputting the target text and the subject recognition prompt text into the large language model, and outputting the description text of at least one subject object; inputting the description text of at least one subject object into the image generation model respectively, and outputting the image of at least one subject object.

[0041] The subject recognition prompt text can be used to prompt the large language model to identify the subject object in the target text and output a descriptive text of the subject object. The descriptive text of the subject object is used to describe the characteristics of the subject object. The image generation model can generate a corresponding image based on the input descriptive text and is a pre-trained neural network model.

[0042] For example, Figure 4 is an example diagram of generating an image of a subject object in an embodiment of the present invention, such as Figure 4 As shown, the target text (script) and the subject identification prompt text (please identify the main actor from the script) are input into the large language model, and the description text of the subject object A (a brown teddy bear with soft hair) is output. Then the description text of A is input into the image generation model, and the image of A is output.

[0043] Specifically, the method for determining the subject object corresponding to each scene can be: input the target text, the description text of at least one subject object and the instruction text into the large language model, and output the subject object corresponding to each scene. The instruction text can be a sentence used to instruct the large language model to identify the subject object of each scene. For example, Figure 5 This is an example diagram of determining the subject objects corresponding to each scene in this embodiment. Figure 5 As shown, the target text (script), the descriptive text of the subject objects (main actors) included in the target text, and the instruction text (extracting the main actors of each scene) are input into the large language model, and the large language model outputs the subject objects of each scene, that is, the main actors of each scene.

[0044] S130 : For each scene, generate audio and video corresponding to the scene according to the image and / or subtext of the main object.

[0045] Here, audio and video can be understood as data including audio and video.

[0046] Specifically, the method of generating audio and video corresponding to a scene based on the image and / or sub-text of the main object can be: inputting the image and sub-text of the main object corresponding to the scene into a video generation model to obtain a video with a continuous duration; inputting the sub-text of the scene into an audio generation model to obtain the audio corresponding to the scene; encoding the video and audio to obtain the audio and video corresponding to the scene.

[0047] The video generation model may be a pre-trained neural network model that can generate videos of a set duration. In this embodiment, if the duration of a scene is less than the set duration, the image and subtext of the main object corresponding to the scene are input into the video generation model, a video of the set duration is output, and the video is cropped to a video of the set duration. If the duration of the scene is equal to the set duration, the image and subtext of the main object corresponding to the scene are directly input into the video generation model, and a video of the set duration is output. If the duration of the scene is greater than the set duration, multiple sub-videos of the set duration are generated by the video generation model based on the image and sub-text of the main object corresponding to the scene, and then these sub-videos are spliced.

[0048] Optionally, the image and sub-text of the main object corresponding to the scene are input into the video generation model, and the method for obtaining a video of continuous duration can be: input the image and sub-text of the main object of the scene into the video generation model, and output multiple sub-videos of set duration; splice multiple sub-videos of set duration to obtain a video of continuous duration.

[0049] Specifically, the image and subtext of the main object of the scene are input into the video generation model, and a method of outputting multiple sub-videos of set lengths can be: for the first sub-video, the image and sub-text of the main object of the scene are input into the video generation model, and the first sub-video is output; for non-first sub-videos, a set number of video frames and sub-text of the previous sub-video are input into the video generation model, and a non-first sub-video is output.

[0050] The set number of video frames of the previous sub-video may be the set number of video frames that are last in the previous sub-video.

[0051] In this embodiment, the image and subtext of the main object of the scene are first input into the video generation model to generate the first sub-video; then the last k video frames and subtext of the first sub-video are input into the video generation model again to generate the second sub-video; if the total duration of the generated sub-video does not reach the duration of the scene, the last k video frames and subtext of the second sub-video are input into the video generation model again to generate the third sub-video, and so on, until the cumulative duration of the generated sub-videos reaches the duration of the scene.

[0052] For example, Figure 6This is an example diagram of generating a video in this embodiment. Figure 6 As shown, scene S 10 The duration of the scene is 8s, and the duration of the video generated by the video generation model is 2s. First, the image and subtext of the main object corresponding to the scene are input into the video generation model to obtain a 2s sub-video; then it is judged whether the process is finished. If not, the last K frames and subtext of the newly generated sub-video are input into the video generation model again to generate a new sub-video; if it is finished, all the sub-videos are spliced.

[0053] Among them, the audio generation model can be a neural network model that can convert text into speech.

[0054] In this embodiment, the subtext of the scene is input into the audio generation model to obtain the audio corresponding to the scene. Finally, the video and audio of the scene are encoded to obtain audio and video data.

[0055] S140: Splice the audio and video of each scene to obtain a target long video.

[0056] In this embodiment, the audio and video of each scene are spliced ​​in chronological order to obtain the target long video. For example, Figure 7 is an example diagram of the target long video spliced ​​in this embodiment, such as Figure 7 As shown in FIG, the target long video is composed of audio and video splicing of each scene.

[0057] Optionally, the multiple sub-videos of set duration are spliced ​​together to obtain the video of the continuous duration in a manner that: if the duration of the spliced ​​video is greater than the continuous duration, the spliced ​​video is cropped to a video of the continuous duration.

[0058] Specifically, the video segment at the end may be cropped.

[0059] The technical solution of this embodiment is based on multiple rounds of interaction between the original text and a large language model to obtain a target text; the target text includes subtexts of multiple scenes and the duration of each scene; an image of at least one main object is generated based on the target text, and the main object corresponding to each scene is determined; for each scene, the corresponding audio and video are generated based on the image and / or subtext of the main object; and the audio and video of each scene are spliced ​​together to obtain a target long video. The long video generation method provided in this embodiment of the present invention generates a target long video based on the main object, image, and subtext of each scene, which can improve the quality of the generated long video and reduce the cost of generating the long video.

[0060] Example 2

[0061] Figure 8This is a schematic diagram of the structure of a long video generation device provided by the second embodiment of the present invention. Figure 8 As shown, the device includes:

[0062] The target text acquisition module 810 is configured to perform multiple rounds of interaction with the large language model based on the original text to obtain the target text; wherein the target text includes sub-texts of multiple scenes and the duration of each scene;

[0063] A subject object image generation module 820 is configured to generate an image of at least one subject object based on the target text and determine the subject object corresponding to each scene;

[0064] An audio and video generation module 830 is configured to generate audio and video corresponding to each scene based on the image and / or subtext of the subject object;

[0065] The target long video acquisition module 840 is used to splice the audio and video of each scene to obtain the target long video.

[0066] Optionally, the target text acquisition module 810 is further configured to:

[0067] For the first round of interaction with the large language model, the original text and the prompt text corresponding to the first round of interaction are input into the large language model, and the intermediate text is output;

[0068] For non-first-round interactions with the large language model, the intermediate text output from the previous round of interaction and the prompt text corresponding to the current round of interaction are input into the large language model, and the intermediate text or target text corresponding to the current round of interaction is output; among them, non-first-round interactions include: intermediate-round interactions or final-round interactions.

[0069] Optionally, the subject object image generation module 820 is further configured to:

[0070] Input the target text and subject recognition prompt text into the large language model, and output a description text of at least one subject object;

[0071] The description text of at least one subject object is input into the image generation model respectively, and the image of at least one subject object is output.

[0072] Optionally, the audio and video generation module 830 is further configured to:

[0073] Input the image and subtext of the main object corresponding to the scene into the video generation model to obtain a video with a continuous duration;

[0074] Input the subtext of the scene into the audio generation model to obtain the audio corresponding to the scene;

[0075] Encode the video and audio to obtain the audio and video corresponding to the scene.

[0076] Optionally, the audio and video generation module 830 is further configured to:

[0077] Input the image and subtext of the main object of the scene into the video generation model, and output multiple sub-videos of set length;

[0078] Multiple sub-videos of set duration are spliced ​​together to obtain a video of continuous duration.

[0079] Optionally, the audio and video generation module 830 is further configured to:

[0080] For the first sub-video, the image of the main object of the scene and the sub-text are input into the video generation model, and the first sub-video is output;

[0081] For a non-first sub-video, a set number of video frames and sub-text of the previous sub-video are input into the video generation model, and the non-first sub-video is output.

[0082] Optionally, the target long video acquisition module 840 is further configured to:

[0083] If the duration of the spliced ​​video is greater than the duration, the spliced ​​video will be cropped to the duration.

[0084] The above device can execute the methods provided by all the above embodiments of the present invention, and has the corresponding functional modules and beneficial effects of executing the above methods. For technical details not fully described in this embodiment, please refer to the methods provided by all the above embodiments of the present invention.

[0085] Example 3

[0086] Figure 9 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0087] like Figure 9As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0088] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0089] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for generating a long video.

[0090] In some embodiments, the method for generating a long video may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for generating a long video described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the method for generating a long video in any other appropriate manner (e.g., by means of firmware).

[0091] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0092] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0093] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0094] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0095] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0096] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0097] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0098] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for generating a long video, characterized in that: include: Based on the original text and the large language model, multiple rounds of interaction are carried out to obtain the target text; wherein the target text includes sub-texts of multiple scenes and the duration of each scene; generating an image of at least one subject object based on the target text, and determining the subject object corresponding to each scene; For each scene, generating audio and video corresponding to the scene based on the image of the subject object and / or the subtext; Splice the audio and video of each scene to obtain the target long video; The target text is obtained by performing multiple rounds of interactions between the original text and the large language model, including: For the first round of interaction with the large language model, the original text and the prompt text corresponding to the first round of interaction are input into the large language model, and an intermediate text is output; For non-first-round interactions with the large language model, the intermediate text output from the previous round of interaction and the prompt text corresponding to the current round of interaction are input into the large language model, and the intermediate text or target text corresponding to the current round of interaction is output; wherein the non-first-round interaction includes: an intermediate-round interaction or a final-round interaction, and the prompt text corresponding to the non-first-round interaction is a sentence used to prompt the large language model to adjust or improve the intermediate text.

2. The method according to claim 1, characterized in that Generating an image of at least one subject object based on the target text includes: Input the target text and subject recognition prompt text into the large language model, and output a description text of at least one subject object; The description text of the at least one subject object is respectively input into the image generation model, and the image of the at least one subject object is output.

3. The method according to claim 1, characterized in that Generating audio and video corresponding to the scene according to the image of the subject object and / or the subtext includes: Inputting the image and subtext of the main object corresponding to the scene into a video generation model to obtain a video of the duration; Inputting the subtext of the scene into an audio generation model to obtain audio corresponding to the scene; The video and the audio are encoded to obtain audio and video corresponding to the scene.

4. The method according to claim 3, characterized in that Inputting the image and subtext of the subject object corresponding to the scene into a video generation model to obtain a video of the duration, including: Inputting the image and subtext of the main object of the scene into the video generation model, and outputting multiple sub-videos of set length; The multiple sub-videos of the set duration are spliced ​​together to obtain a video of the set duration.

5. The method according to claim 4, characterized in that The image and subtext of the main object of the scene are input into the video generation model, and multiple sub-videos of set lengths are output, including: For the first sub-video, input the image and sub-text of the main object of the scene into the video generation model, and output the first sub-video; For a non-first sub-video, a set number of video frames of a previous sub-video and the sub-text are input into the video generation model, and the non-first sub-video is output.

6. The method according to claim 4, characterized in that The method of splicing the plurality of sub-videos of set duration to obtain a video of the set duration includes: If the duration of the spliced ​​video is greater than the duration, the spliced ​​video is cropped to a video of the duration.

7. A device for generating a long video, characterized in that: include: A target text acquisition module is used to perform multiple rounds of interaction with the large language model based on the original text to obtain the target text; wherein the target text includes sub-texts of multiple scenes and the duration of each scene; A subject object image generation module, configured to generate an image of at least one subject object based on the target text and determine the subject object corresponding to each scene; An audio and video generation module, configured to generate, for each scene, audio and video corresponding to the scene based on the image of the subject object and / or the subtext; The target long video acquisition module is used to splice the audio and video of each scene to obtain the target long video; The target text acquisition module is further used to: For the first round of interaction with the large language model, the original text and the prompt text corresponding to the first round of interaction are input into the large language model, and the intermediate text is output; For non-first-round interactions with the large language model, the intermediate text output from the previous round of interaction and the prompt text corresponding to the current round of interaction are input into the large language model, and the intermediate text or target text corresponding to the current round of interaction is output. Non-first-round interactions include intermediate-round interactions or final-round interactions, and the prompt text corresponding to the non-first-round interaction is a sentence used to prompt the large language model to adjust or improve the intermediate text.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for generating a long video according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for generating a long video according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • Generative artificial intelligence-based novel tweet video generation method and system

    CN117078782A