Video generation method, device and equipment, computer readable storage medium and product
By obtaining user material content and using large language models to generate plot text and storyboard pictures, the problem of cumbersome video generation operations is solved, and the effect of efficiently generating high-quality videos is achieved.
Patent Information
- Application Number
- CN202510864026.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-02
AI Technical Summary
In the prior art, video generation operations are cumbersome and have high professional requirements, resulting in low generation efficiency.
By obtaining the material content input by the user, using a large language model to generate plot text content, and generating a preset number of storyboard pictures based on plot text content and preset prompt content, and finally automatically generates the target video.
Without the need for users to shoot and edit video materials by themselves, they can quickly generate high-quality target videos that meet user needs, improving generation efficiency.
Smart Images

Figure CN120583294A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of data processing technology, and in particular to a video generation method, apparatus, device, computer-readable storage medium, and product. Background Art
[0002] With the improvement of terminal device hardware performance, users can generate video content on the terminal device. Users generally need to capture video content on the terminal device and generate the final target video through clipping and editing operations on the video content.
[0003] However, using the above method to generate videos is often complicated and requires a high level of professionalism from the user. Summary of the Invention
[0004] The embodiments of the present disclosure provide a video generation method, apparatus, device, computer-readable storage medium, and product, which are used to solve the technical problem that current video content generation operations are relatively cumbersome and difficult.
[0005] In a first aspect, an embodiment of the present disclosure provides a video generation method, comprising:
[0006] In response to a user-triggered video generation operation, obtaining material content determined by the user, the material content including one or more of character material, music material content, and video description text material;
[0007] In response to a preview operation triggered by the user, displaying plot text content used to generate a target video in a first display interface, where the plot text content is generated by a large language model based on the source content;
[0008] In response to the user-triggered storyboard generation operation, a preset number of storyboards generated based on the plot text content and a preset first prompt content are displayed in the first display interface, where the first prompt content is used to prompt the large language model to generate storyboards that meet a preset first condition based on the plot text content;
[0009] In response to the video generation operation triggered by the user, the target video is generated based on the preset number of storyboards.
[0010] In a second aspect, an embodiment of the present disclosure provides a video generation device, including:
[0011] an acquisition module, configured to acquire, in response to a user-triggered video generation operation, material content determined by the user, the material content including one or more of character material, music material content, and video description text material;
[0012] a display module, configured to display, in response to a preview operation triggered by the user, plot text content used to generate a target video in a first display interface, wherein the plot text content is generated by a large language model based on the source content;
[0013] a processing module configured to, in response to the storyboard generation operation triggered by the user, display, in the first display interface, a preset number of storyboards generated based on the plot text content and a preset first prompt content, wherein the first prompt content is configured to prompt the large language model to generate storyboards that meet a preset first condition based on the plot text content;
[0014] A generation module is used to generate the target video based on the preset number of storyboards in response to the video generation operation triggered by the user.
[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0016] The memory stores computer-executable instructions;
[0017] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video generation method as described in the first aspect and various possible designs of the first aspect.
[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0019] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the video generation method described in the first aspect and various possible designs of the first aspect.
[0020] The video generation method, apparatus, device, computer-readable storage medium, and product provided in this embodiment obtain the user-determined material content in response to a user-triggered video generation operation, thereby generating a complete plot text content based on the material content, and generating a preset number of storyboards based on the plot text content and a preset first prompt content. Therefore, the user can view the preset number of storyboards, and if the user's needs are met, the target video is automatically generated based on the preset number of storyboards. This eliminates the need for the user to shoot video material and edit the video material to generate the target video. The user only needs to select the material content to generate a high-quality target video that meets the user's actual needs, effectively improving the efficiency of generating the target video. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0022] Figure 1 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0023] Figure 2 A schematic diagram of the interface interaction provided by an embodiment of the present disclosure;
[0024] Figure 3 A flowchart of a video generation method provided by another embodiment of the present disclosure;
[0025] Figure 4 Another interface interaction diagram provided by an embodiment of the present disclosure;
[0026] Figure 5 A schematic diagram of a display interface provided in an embodiment of the present disclosure;
[0027] Figure 6 A flowchart of a video generation method provided by another embodiment of the present disclosure;
[0028] Figure 7 A schematic diagram of the structure of a video generating device provided in an embodiment of the present disclosure;
[0029] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0031] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0032] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0033] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0034] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0035] In related technologies, to generate a target video, users typically need to shoot multiple video clips based on their needs and then edit them to create the target video. However, these methods often require a high level of user expertise. Furthermore, these methods are often complex and cumbersome, resulting in low video generation efficiency.
[0036] In solving the above technical problems, the inventors discovered through research that it is possible to obtain user-entered content. This content can include textual content of the user-entered video inspiration, at least one protagonist associated with the video, and one or more musical materials. Therefore, based on this content, the generation capabilities of a large language model can be leveraged to automatically generate a target video that better meets the user's needs.
[0037] Optionally, in order to generate a higher quality target video, the generation process of the large language model may be constrained by preset prompt content.
[0038] In order to solve the technical problem that current video content generation operations are relatively cumbersome and difficult, the present disclosure provides a video generation method, apparatus, device, computer-readable storage medium and product.
[0039] It should be noted that the video generation method, apparatus, device, computer-readable storage medium and product provided by the present disclosure can be applied in any video generation scenario.
[0040] Figure 1 A flow chart of a video generation method provided by an embodiment of the present disclosure is shown in FIG. Figure 1As shown, the method includes:
[0041] Step 101: In response to a video generation operation triggered by a user, obtain material content determined by the user, where the material content includes one or more of character material, music material, and video description text material.
[0042] The execution subject of this embodiment is a video generation device. The video generation device can be coupled to a terminal device, so that it can determine the material content based on the user's trigger operation on the terminal device, and call the large language model to generate a target video based on the material content. Alternatively, the video generation device can be coupled to a server, which can be communicatively connected to the terminal device. Therefore, it can obtain the material content determined by the user in the terminal device, call the large language model to generate a target video based on the material content, and send the generated target video to the terminal device for display.
[0043] In this embodiment, the user can trigger a video generation operation on the terminal device. In response to the video generation operation, user-determined material content can be obtained, including one or more of character material, music material content, and video description text material.
[0044] The character material can be the main character in the target video. The music material can be the background music of the target video. The video description text material can be the video inspiration or video content summary text entered by the user according to actual needs.
[0045] As an implementable method, the character material and the music material are mandatory materials, while the video description text material can be optional materials. The user can determine whether to input the video description text material according to actual needs.
[0046] Optionally, users can generate target videos of different video types based on their actual needs. Video types include, but are not limited to, music videos, vlogs, and single-player videos. Users can specify different content for different video types. For example, if a user chooses to generate a music video, they can specify at least one character and music content. Furthermore, users can specify video description text based on their actual needs.
[0047] Step 102: In response to the preview operation triggered by the user, the plot text content used to generate the target video is displayed in the first display interface, where the plot text content is generated by the large language model based on the material content.
[0048] In this embodiment, the video description text material input by the user may be relatively brief, or the user may not have currently input the video description text material. Therefore, in order to achieve the generation of the target video, it is also necessary to determine the complete plot text content used to generate the target video.
[0049] Therefore, after obtaining the material content determined by the user, in response to the preview operation triggered by the user, the large language model can be called to generate the plot text content based on the material content. In order to enable the user to browse the plot text content more intuitively, the plot text content used to generate the target video can also be displayed in the first display interface.
[0050] When the user inputs a relatively short video description text material, the plot text content can be obtained by expanding the video description text material based on the video description text material and other material content. If the user does not input a video description text material, the plot text content can be automatically generated by the large language model based on the material content.
[0051] After displaying the plot text content, the user can edit the plot text content according to actual needs, wherein the editing operation includes but is not limited to modifying and deleting part of the plot text content, and adding plot text, etc., which is not limited in this disclosure.
[0052] Step 103: In response to the storyboard generation operation triggered by the user, a preset number of storyboards generated based on the plot text content and a preset first prompt content are displayed in the first display interface, where the first prompt content is used to prompt the large language model to generate a storyboard that meets a preset first condition based on the plot text content.
[0053] In this embodiment, after obtaining the plot text content, a preset number of storyboards can be generated based on the plot text content. The storyboards are also known as shot scripts. The shot scripts serve as an intermediate medium for converting text into three-dimensional audiovisual images. Their primary task is to design the corresponding screens based on the plot text content, configure the music and sound, and capture the rhythm and style of the target video.
[0054] Optionally, in response to a storyboard generation operation triggered by a user, the plot text content and a preset first prompt content can be input into the large language model, so that the large language model generates a storyboard that satisfies a preset first condition. The first prompt content is used to prompt the large language model to generate a storyboard that satisfies the preset first condition based on the plot text content. The first condition includes one or more of the number of storyboards, the screen type corresponding to each storyboard, the proportion information corresponding to different types of storyboards, the display content information corresponding to each storyboard, and the correspondence between each storyboard and the lyrics in the music material content. The screen type includes one or more of the following: an empty shot type, a lip sync type, and a character type.
[0055] It should be noted that since there are a preset number of storyboards, in order to generate a higher-quality storyboard image, the content displayed in each storyboard can be constrained in the first prompt content. For example, the first storyboard can be constrained to be an empty frame, and an atmospheric seaside image can be displayed in the frame. Therefore, each storyboard matches at least part of the content in the first prompt content.
[0056] Step 104: In response to the video generation operation triggered by the user, generate the target video based on the preset number of storyboards.
[0057] In this embodiment, after obtaining multiple storyboards, a higher-quality target video can be generated based on the storyboards. In response to a user-triggered video generation operation, the target video is generated based on a preset number of storyboards. The user can trigger the video generation operation by using a trigger generation control.
[0058] Figure 2 This is a schematic diagram of the interface interaction provided by the embodiment of the present disclosure, such as Figure 2 As shown, after obtaining the source content 21 determined by the user, in response to a user-triggered preview operation, plot text content 23 generated based on the source content can be displayed within a first display interface 22. In response to a user-triggered storyboard generation operation, a preset number of storyboards 24 generated based on the plot text content 23 can be switched and displayed within the first display interface 22. In response to a user-triggered video generation operation, a target video can be generated based on the preset number of storyboards.
[0059] The video generation method provided in this embodiment obtains user-determined source material content in response to a user-triggered video generation operation, thereby generating a complete plot text based on the source material content and generating a preset number of storyboards based on the plot text content and a preset first prompt content. Therefore, the user can view the preset number of storyboards, and if the user's needs are met, the target video is automatically generated based on the preset number of storyboards. This eliminates the need for the user to shoot and edit the video material to generate the target video. The user only needs to select the source material to generate a high-quality target video that meets the user's actual needs, effectively improving the efficiency of generating the target video.
[0060] Optionally, based on any of the above embodiments, after step 104, the method further includes:
[0061] The target video is generated in the background, and a preset content display interface is displayed, in which target videos generated historically by other users are displayed.
[0062] In this embodiment, since the target video generation process may take a long time, in order to avoid the user having to wait for a long time in the front-end interface, the target video can be generated in the background. The front-end displays a preset content display interface, and the target videos previously generated by other users are displayed in the content display interface, so that the user can continue to browse more target videos in the content display interface.
[0063] The video generation method provided in this embodiment generates the target video in the background and displays the content display interface on the front end, so there is no need to wait for the target video to be generated on the front end interface, and the user can continue to browse more target videos in the content display interface.
[0064] Figure 3 A flow chart of a video generation method provided by another embodiment of the present disclosure is provided. Based on any of the above embodiments, Figure 3 As shown, step 101 includes:
[0065] Step 301: Display selection controls associated with multiple video types in a second display interface.
[0066] Step 302: In response to the user triggering an operation on a selection control associated with any video type, display the first display interface.
[0067] Step 303: Display a material determination control corresponding to the video type associated with the selection control triggered by the user in the first display interface, wherein the material determination control includes one or more of a text input control, a role determination control, and a music determination control.
[0068] Step 304: Acquire the material content determined by the user based on the material determination control.
[0069] In this embodiment, users can generate target videos of different video types according to actual needs, where the video types include but are not limited to music MV types, Vlog types, single-person video types, etc. To facilitate the user's selection of video types, selection controls associated with multiple video types can be displayed in the second display interface.
[0070] For different video types, users can determine different material content.
[0071] Therefore, to facilitate users in quickly and easily determining the content of the material, a first display interface is displayed in response to a user triggering a selection control associated with any video type. Within the first display interface, a material determination control corresponding to the video type associated with the user-triggered selection control is displayed. The material determination control includes one or more of a text input control, a character determination control, and a music determination control. Thus, users can determine the content of the material based on the material determination control.
[0072] For example, when a user selects to generate a target video of the music video type, the user can specify at least one character material and music material content. The user can also specify the video description text material based on actual needs. Therefore, a text input control, a character determination control, and a music determination control can be displayed separately in the first display interface.
[0073] The video generation method provided in this embodiment displays selection controls associated with multiple video types within the second display interface, allowing users to generate different target videos based on their actual needs. Furthermore, by displaying a material determination control corresponding to the user-selected video type within the first display interface, users can more quickly determine material content that matches the currently selected video type, improving the efficiency and accuracy of material content determination.
[0074] Further, based on any of the above embodiments, step 304 includes:
[0075] In response to the user triggering the role determination control, a material determination interface is displayed.
[0076] A plurality of images to be selected in a preset first storage path and images to be selected associated with the user-associated virtual image are displayed in the material determination interface, where the virtual image is generated by the large language model based on the image material input by the user.
[0077] In response to a selection operation triggered by the user in the material interface, the person in the target image to be selected selected by the user is determined as the character material.
[0078] And / or, step 304 includes:
[0079] The video description text material input by the user based on the text input control is obtained.
[0080] And / or, step 304 includes:
[0081] In response to the user triggering the music determination control, a material determination interface is displayed.
[0082] All audios to be selected in the preset second storage path are displayed in the material determination interface.
[0083] In response to a selection operation triggered by the user in the material interface, the target audio selected by the user is determined as the music material content.
[0084] In this embodiment, the user can select the character material by triggering the character determination control.
[0085] Optionally, in response to the user triggering the role determination control, a material determination interface may be displayed. Multiple images to be selected in a preset first storage path and images to be selected associated with a user-associated virtual image may be displayed in the material determination interface. The virtual image is generated by a large language model based on the image material input by the user. The virtual image may be an AI image generated based on the user's photo or real-time image, and may have the user's personalized characteristics. The first storage path includes, but is not limited to, an album storage path within the user's terminal device.
[0086] Therefore, after the material determination interface is displayed, the user can trigger a selection operation in the material determination interface according to actual needs. In response to the selection operation triggered by the user in the material interface, the person in the target image selected by the user is determined as the character material.
[0087] Optionally, the user may determine the video description text material by triggering a text input control.
[0088] The text input control may be a text input box. In response to a user triggering operation on the text input control, an input keyboard may be pulled up to obtain the video description text material input by the user through the input keyboard.
[0089] Optionally, the user can determine the content of the music material by triggering the music determination control.
[0090] In response to a user triggering the music determination control, a material determination interface is displayed. The material determination interface displays all available audio files within a preset second storage path. For example, the material determination interface may display all available audio files stored within a folder within the terminal device. Therefore, after the material determination interface is displayed, the user can trigger a selection operation within the material determination interface based on actual needs. In response to the user triggering the selection operation within the material interface, the target audio file selected by the user is determined as the music material content.
[0091] Figure 4 Another interface interaction diagram provided by the embodiment of the present disclosure is as follows: Figure 4 As shown, selection controls 42 associated with multiple video types can be displayed within the second display interface 41. In response to a user triggering an operation on a selection control associated with any video type 42, a first display interface 43 is displayed. In response to the user triggering an operation on a selection control associated with a music MV type, a text input control 44, a character determination control 45, and a music determination control 46 can be displayed within the first display interface 43.
[0092] Taking the character selection process as an example, in response to a user triggering a character selection control 45, a material selection interface 47 is displayed. Within this interface, multiple candidate images 48 from a preset first storage path are displayed, along with a candidate image 49 associated with a user-associated avatar. The avatar is generated by the large language model based on the image material input by the user. In response to a user triggering a selection within the material interface, the person in the target candidate image selected by the user is determined as the character material.
[0093] The video generation method provided in this embodiment displays the material determination control corresponding to the video type selected by the user in the first display interface, so that the user can more quickly determine the material content that matches the currently selected video type, thereby improving the efficiency and accuracy of material content determination.
[0094] Optionally, based on any of the above embodiments, before step 102, the method further includes:
[0095] The material content and the preset second prompt content are input into the large language model to obtain feature information output by the large language model, wherein the feature information includes one or more of character features, character-associated clothing features, location features, and video style features, and the second prompt content is used to prompt the large language model to generate feature information that meets a preset second condition based on the material content.
[0096] The feature information and a preset third prompt content are input into the large language model to obtain the plot text content output by the large language model, wherein the third prompt content is used to prompt the large language model to expand the feature information according to a preset third condition to generate the plot text content.
[0097] In this embodiment, in order to accurately generate the plot text content, a feature extraction operation may be performed on the material content.
[0098] Optionally, the source material and a preset second prompt can be input into the large language model to obtain feature information output by the large language model. The feature information includes one or more of character features, character-associated clothing features, location features, and video style features. The second prompt is used to prompt the large language model to generate feature information that satisfies a preset second condition based on the source material. Furthermore, the music source material can be analyzed to determine the lyrics within the music source material, as well as features such as the scene, environment, and character relationships corresponding to the lyrics.
[0099] For example, the second prompt might ask you to summarize the main "characters, costumes, locations, and styles" based on the given material. The specific steps are as follows: 1. Carefully analyze the material. 2. If the material contains a costume description, output {character: xx}; if it does not, output {character: none}. 3. If the material contains a location or scene description, output {location: xx}; if it does not, output {location: none}. 4. If the material contains a style description, output {outfit: x}; if it does not, output {outfit: none}. 5. If the material contains a description of the protagonist, such as hair length or color, output {character: xx}; if it does not, output {character: none}. Please ensure that the output does not contain any XML tags.
[0100] Based on the image content, output the subject's gender, hair style, and skin color. Outputting other information is prohibited. Hair styles must be carefully distinguished and accurately described. Hair that reaches the chin and above is short, hair that reaches the chin to the shoulders is medium-long, and hair below the shoulders is long. Output examples: 1. Asian girl with long curly hair in xx color; 2. Asian boy with short straight hair in xx color.
[0101] Furthermore, after obtaining the characteristic information corresponding to the material content, the characteristic information and the preset third prompt content can be input into the large language model to obtain the plot text content output by the large language model, wherein the third prompt content is used to prompt the large language model to expand the characteristic information according to the preset third condition to generate the plot text content.
[0102] For example, the third prompt content can be 1. Based on the given music material content and the video description text material input by the user, output information such as [role, outfit, location, style]. 2. Role aspect: Integrate the character modeling description ① to determine the subject and user input. 3. Outfit aspect: If there is user input, give priority to the user input; if not, refer to the lyrics and output an outfit with atmosphere and good looks, ensuring that simple and basic colors are the main color. 4. Location aspect: If there is user input, give priority to the user input; if not, refer to the lyrics and output an atmosphere, good looks, super simple but aesthetic location. 5. Style aspect: If there is user input, give priority to the user input; if not, try to summarize the most suitable style from the lyrics. 6. Output format: Role: xx, Outfit: xx, Location: xx, Style: xx.
[0103] Therefore, based on the third prompt content and feature information mentioned above, the large language model can generate complete and high-quality plot text content.
[0104] The video generation method provided in this embodiment uses a large language model to extract feature information associated with the material content, and inputs the feature information and preset third prompt content into the large language model, so that the large language model can perform an expansion operation based on the feature information to obtain more complete and rich plot text content, and then can generate a higher quality target video based on the plot text content.
[0105] Furthermore, based on any of the above embodiments, step 103 includes:
[0106] In response to the user triggering the preset screen generation control in the first display interface, the plot text content and the preset third prompt content are input into the large language model to obtain a preset number of storyboards output by the large language model.
[0107] Among them, the first prompt content is used to prompt the large language model to generate a storyboard screen that meets a preset first condition based on the plot text content. The first condition includes one or more of the number of storyboard screens, the screen type corresponding to each storyboard screen, the proportion information corresponding to different types of storyboard screens, the display content information corresponding to each storyboard screen, and the correspondence between each storyboard screen and the lyrics in the music material content. The screen type includes one or more of the empty shot type, lip-sync type, and character type.
[0108] In this embodiment, after obtaining the plot text content, a preset number of storyboards can be generated based on the plot text content. The storyboards are also known as shot scripts. The shot scripts are the intermediate medium that converts text into three-dimensional audiovisual images. Their primary task is to design the corresponding screens based on the plot text content, configure the music and sound, and grasp the rhythm and style of the target video.
[0109] It should be noted that since there are a preset number of storyboards, in order to generate a higher-quality storyboard image, the content displayed in each storyboard can be constrained in the first prompt content. For example, the first storyboard can be constrained to be an empty frame, and an atmospheric seaside image can be displayed in the frame. Therefore, each storyboard matches at least part of the content in the first prompt content.
[0110] Therefore, in response to a user triggering a preset screen generation control in the first display interface, the plot text content and the preset third prompt content can be input into the large language model to obtain a preset number of storyboards output by the large language model. The first prompt content is used to prompt the large language model to generate storyboards that meet a preset first condition based on the plot text content. The first condition includes one or more of the number of storyboards, the screen type corresponding to each storyboard, the proportion information corresponding to different types of storyboards, the display content information corresponding to each storyboard, and the correspondence between each storyboard and the lyrics in the music material content. The screen type includes one or more of the following: empty shot type, lip sync type, and character type.
[0111] For example, the third prompt content can be based on the lyrics in the music material content, combined with the plot text content to help me split the lyrics by line into (preset number of storyboards, written in this format: storyboard n-scene: xx, role: xxxx (no need to describe when the protagonist is empty), picture description: xxxx, outfit: xxxx (no need to describe when the protagonist is empty), environment: xxxx, 1 type xX, color tone: xxXX, described in the form of prompt; the requirements are as follows: 1. {*} type PE / {x} output 2. Ensure that each picture is as simple as possible and the picture is beautiful, mainly to highlight the sense of atmosphere; 3. When characters appear in the picture, most of them are close-up beautiful photos of the front face, mainly presenting the beauty of the atmosphere, and the scene and action settings are super Simple; 4. When characters appear in the screen, describe them with their front faces, reduce side faces, and prohibit full side faces and descriptions of facial features; 5. Reduce the movements of the characters' hands, reduce the description of the hands next to the face, and avoid blocking the face with objects held by the hands; the movements should be simple, and avoid complex body expressions such as "crouching" and "crossing legs"; 6. The characters should not look up, down, or raise their heads, and the characters should not close their eyes; 7. The appearance of the characters is described according to the story setting, and each storyboard screen briefly describes the outfit set in the story setting; 8. The outfit needs to have a description of the clothing style to ensure that the clothing matching and clothing style of each storyboard screen are consistent; 9. Ensure the adaptability and coherence of the picture and the song. The environmental presentation of all storyboard screens must be consistent in style to avoid the situation where the background of the character screen is a European environment and the empty screen screen is a Chinese courtyard.
[0112] The video generation method provided in this embodiment, after acquiring the plot text content, uses a large language model to generate multiple storyboards based on the plot text content and the first prompt content, thereby generating a storyboard that better suits the user's current needs. Furthermore, by displaying these storyboards in the first display interface, the user can easily adjust and update the storyboards, thereby generating a target video that better suits the user's actual needs.
[0113] Furthermore, based on any of the above embodiments, after step 103, the following steps are further included:
[0114] Display preset regeneration controls in the display area associated with each storyboard.
[0115] In response to the user triggering an operation of regenerating a control associated with any storyboard, the first prompt content associated with the storyboard triggered by the user and the plot text content are re-input into the large language model to obtain an updated storyboard output by the large language model.
[0116] The updated storyboard is switched and displayed at the position of the storyboard triggered by the user.
[0117] In this embodiment, after a preset number of storyboards are displayed on the first display interface, if the user is not satisfied with the display effect of any storyboard, the user can re-generate the storyboard.
[0118] Optionally, a preset regeneration control may be displayed in a display area associated with each storyboard screen. For example, the regeneration control may be displayed on top of the storyboard screen.
[0119] Each storyboard is associated with at least a portion of the first prompt content. Therefore, in response to a user triggering a regeneration control associated with any storyboard, the first prompt content and plot text associated with the storyboard triggered by the user are re-input into the large language model, so that the large language model can regenerate the storyboard and obtain an updated storyboard output by the large language model.
[0120] In order to facilitate the user's viewing of the updated storyboard screen, the updated storyboard screen can be switched and displayed at the position of the storyboard screen triggered by the user.
[0121] Figure 5 A schematic diagram of a display interface provided by an embodiment of the present disclosure, such as Figure 5 As shown, a preset number of storyboards 52 can be displayed in the first display interface 51. A preset regeneration control 53 is displayed in the display area associated with each storyboard 52. Therefore, when the user is not satisfied with the display effect of any storyboard 52, the regeneration control 53 can be triggered to regenerate the storyboard.
[0122] The video generation method provided in this embodiment displays storyboards in a first display interface and displays preset regeneration controls in the display area associated with each storyboard, so that the user can quickly adjust and update the storyboards by triggering the regeneration controls, thereby generating a target video that better meets the user's actual needs.
[0123] Figure 6 A flow chart of a video generation method provided by another embodiment of the present disclosure is provided. Based on any of the above embodiments, Figure 6 As shown, step 104 includes:
[0124] Step 601: For each storyboard screen, determine the associated information of the storyboard screen, wherein the associated information includes the screen type associated with the storyboard screen, the display order associated with the storyboard screen, and the lyrics text paragraph corresponding to the storyboard screen, wherein the lyrics text paragraph is obtained by identifying the content of the music material.
[0125] Step 602: For each storyboard picture, determine a video processing method corresponding to the picture type associated with the storyboard picture, perform video processing on the storyboard picture according to the video processing method, and obtain a video segment corresponding to the storyboard picture.
[0126] Step 603: splicing the video segments corresponding to the preset number of storyboards according to the display order associated with the storyboards to obtain a video to be processed.
[0127] Step 604: add lyrics and subtitles to the video to be processed according to the lyrics text paragraphs corresponding to the storyboard images to obtain the target video.
[0128] In this embodiment, after a plurality of storyboards are generated based on the plot text content, a video segment may be generated based on each storyboard, and a splicing operation may be performed on the plurality of video segments to obtain a final target video.
[0129] Optionally, to more accurately generate video segments, associated information for each storyboard is determined. This information includes the associated screen type, the associated display order of the storyboards, and the corresponding lyric text segments. The lyric text segments are obtained by identifying the content of the music material. Different storyboards correspond to different lyric text segments, and the sum of the lyric text segments corresponding to all storyboards can constitute the complete lyrics of the music material.
[0130] Furthermore, the screen types include blank shot type, lip sync type, and portrait type. For different screen types, due to the different display content, the method for generating video segments also varies. Therefore, for each storyboard screen, a video processing method corresponding to the screen type associated with the storyboard screen is determined, and the storyboard screen is video processed according to the video processing method to obtain the video segment corresponding to the storyboard screen.
[0131] Therefore, after obtaining the video segments corresponding to each storyboard, a splicing operation can be performed on the video segments corresponding to a preset number of storyboards according to the presentation order of the storyboards to obtain the video to be processed. Transition effects, etc., can be added between the two video segments based on actual needs, and this disclosure does not impose any restrictions on this. Furthermore, lyrics and subtitles can be added to the video to be processed based on the lyrics text segments corresponding to the storyboards to obtain the target video.
[0132] The video generation method provided in this embodiment utilizes different video processing methods for storyboards of different screen types to generate corresponding video segments, thereby making the generated video segments more suitable for the current screen type. Furthermore, the target video is obtained by splicing the video segments corresponding to a preset number of storyboards according to their associated presentation order, and adding lyrics and subtitles to the processed video based on the corresponding lyrics text segments of the storyboards. This results in a higher-quality target video.
[0133] Furthermore, based on any of the above embodiments, the screen type includes an empty shot type, and the associated information of the storyboard screen also includes a portion of the plot text content associated with the storyboard screen. Step 602 includes:
[0134] The storyboard, part of the plot text content associated with the storyboard, and the fourth prompt content are input into the large language model to obtain the video segment corresponding to the storyboard output by the large language model, wherein the fourth prompt content is used to prompt the large language model to generate an empty shot video that meets the preset fourth condition based on the current input content.
[0135] In the present embodiment, the picture type can be an empty mirror type. Under the picture of the empty mirror type, no characters appear, and it can mostly be an environment or a landscape picture.
[0136] Therefore, the storyboard screen, part of the plot text content associated with the storyboard screen, and the fourth prompt content can be input into the large language model to obtain the video segment corresponding to the storyboard screen output by the large language model, wherein the fourth prompt content is used to prompt the large language model to generate an empty shot video that meets the preset fourth condition based on the current input content.
[0137] The fourth prompt can be used to prompt the large language model to generate a video segment with simple, aesthetically pleasing graphics and a sense of atmosphere. Alternatively, the fourth prompt can be used to prompt the large language model to generate a video segment that matches the music source. For example, if the music source is a song associated with the seaside, the large language model can be prompted to generate a video segment associated with the ocean or the seaside. In actual applications, the fourth prompt can be adjusted according to actual needs, and this disclosure does not impose any restrictions on this.
[0138] The video generation method provided in this embodiment obtains the video segments corresponding to the storyboards output by the large language model by inputting the storyboard images, part of the plot text content associated with the storyboard images, and the fourth prompt content into the large language model, thereby being able to quickly generate video segments of higher quality that are more in line with current needs.
[0139] Furthermore, based on any of the above embodiments, the picture type includes a lip sync type, and the associated information of the storyboard picture further includes a portion of the plot text content associated with the storyboard picture. Step 602 includes:
[0140] A face-changing operation is performed on the face area in the storyboard based on the face area in the character material determined by the user to obtain a processed face-changing image.
[0141] The face-swapped image, part of the plot text content associated with the storyboard, the lyrics text paragraph corresponding to the storyboard, and the preset fifth prompt content are input into the large language model to obtain a video paragraph corresponding to the storyboard output by the large language model, wherein the fifth prompt content is used to prompt the large language model to generate a video paragraph with a lip-syncing effect based on the current input content.
[0142] In this embodiment, the scene type can be lip sync. In a lip sync shot, a character area may appear, where the character area includes, but is not limited to, a face area, a body area including a face, etc. In the lip sync type, the character performs a lip syncing operation to the lyrics in the music material.
[0143] Optionally, to make the characters in the generated video clips more closely match the character material selected by the user, during the video generation process, a face-swap operation can be performed on the face area in the storyboard based on the face area in the character material determined by the user, thereby obtaining a processed face-swap image. Any face-swap method can be used to implement the face-swap operation on the storyboard image, and this disclosure does not impose any restrictions on this.
[0144] Furthermore, the face-changing image, part of the plot text content associated with the storyboard, the lyrics text paragraph corresponding to the storyboard, and the preset fifth prompt content can be input into the large language model to obtain the video paragraph corresponding to the storyboard output by the large language model, wherein the fifth prompt content is used to prompt the large language model to generate a video paragraph with a lip-syncing effect based on the current input content.
[0145] Among them, the fifth prompt content can constrain the large language model to generate a video singing to the face-swapped image, and the lip shape of the face-swapped image should match the lyrics text paragraph corresponding to the storyboard picture. In addition, the fifth prompt content can also constrain the large language model to generate a video paragraph with simple pictures, aesthetics, and a prominent atmosphere, or the fifth prompt content can also constrain the large language model to generate a video paragraph with uniform character clothing and hairstyle. Alternatively, the fifth prompt content can also constrain the large language model to generate a video paragraph with smooth and simple character movements, etc. In actual application, the sixth prompt content can be adjusted according to actual needs, and this disclosure does not impose any restrictions on this.
[0146] The video generation method provided in this embodiment, when the screen type is lip-sync, performs a face-swap operation on the face area of the storyboard based on the face area of the character material specified by the user, thereby generating a face-swap image that better matches the user-specified character material. Furthermore, by inputting the face-swap image, the portion of plot text associated with the storyboard, the lyrics text paragraph corresponding to the storyboard, and a preset fifth prompt into a large language model, a high-quality video segment that better matches the user-specified character material can be generated.
[0147] Furthermore, based on any of the above embodiments, the picture type includes a portrait type, and the associated information of the storyboard picture also includes a portion of the plot text content associated with the storyboard picture. Step 602 includes:
[0148] A face-changing operation is performed on the face area in the storyboard based on the face area in the character material determined by the user to obtain a processed face-changing image.
[0149] The face-swapped image, part of the plot text content associated with the storyboard, and the preset sixth prompt content are input into the large language model to obtain a video segment corresponding to the storyboard output by the large language model, wherein the sixth prompt content is used to prompt the large language model to generate a portrait video segment that meets the preset fifth condition based on the current input content.
[0150] In this embodiment, the screen type can be a portrait type. In a portrait type storyboard shot, a character area may appear, where the character area includes but is not limited to the body area, the head area, etc. In the portrait type, the character in the screen does not need to perform lip syncing.
[0151] Optionally, to make the characters in the generated video clips more closely match the character material selected by the user, during the video generation process, a face-swap operation can be performed on the face area in the storyboard based on the face area in the character material determined by the user, thereby obtaining a processed face-swap image. Any face-swap method can be used to implement the face-swap operation on the storyboard image, and this disclosure does not impose any restrictions on this.
[0152] Furthermore, after obtaining the face-swapped image, the face-swapped image, the portion of the plot text associated with the storyboard, and the preset sixth prompt content can be input into the large language model. This allows the large language model to output a video segment corresponding to the storyboard. The sixth prompt content prompts the large language model to generate a portrait video segment that meets the preset fifth condition based on the current input content.
[0153] For example, the sixth prompt could constrain the large language model to generate a video segment that matches the lyrics in the music clip. For example, the lyrics could be "Meeting by chance, relying on each other," and the video segment could depict two people meeting. If the user-selected character clip is a person, the video could instead depict a sense of dependence, for example, with the character standing at a dock, and the surroundings conveying the atmosphere of a meeting.
[0154] For example, the sixth prompt could constrain the large language model to generate video segments with consistent hairstyles, clothing, and attire. Alternatively, the sixth prompt could constrain the large language model to generate video segments with more aesthetic close-up shots of frontal faces and fewer profile or full-profile shots. Alternatively, the sixth prompt could constrain the large language model to generate video segments with fewer hand movements, etc. In actual applications, the sixth prompt can be adjusted based on actual needs, and this disclosure does not impose any restrictions on this.
[0155] The video generation method provided in this embodiment performs a face-swapping operation on the face area of a storyboard based on the face area of a user-specified character material when the screen type is a portrait type, thereby generating a face-swapping image that better matches the user-specified character material. Furthermore, by inputting the face-swapping image, a portion of the plot text associated with the storyboard, and a preset seventh prompt into a large language model, a high-quality video segment that better matches the user-specified character material can be generated.
[0156] Figure 7 A schematic diagram of the structure of a video generating device provided in an embodiment of the present disclosure is shown in FIG. Figure 7 As shown, the device includes: an acquisition module 71, a display module 72, a processing module 73, and a generation module 74. The acquisition module 71 is configured to, in response to a user-triggered video generation operation, acquire user-determined material content, including one or more of character material, music material, and video description text material. The display module 72 is configured to, in response to a user-triggered preview operation, display plot text content used to generate a target video within a first display interface. The plot text content is generated by a large language model based on the material content. The processing module 73 is configured to, in response to a user-triggered storyboard generation operation, display a preset number of storyboards generated based on the plot text content and a preset first prompt content within the first display interface. The first prompt content is configured to prompt the large language model to generate a storyboard that meets a preset first condition based on the plot text content. The generation module 74 is configured to, in response to the user-triggered video generation operation, generate the target video based on the preset number of storyboards.
[0157] Furthermore, based on any of the above embodiments, the acquisition module is configured to: display selection controls associated with multiple video types within the second display interface; display the first display interface in response to the user triggering an operation on a selection control associated with any video type; display a material determination control corresponding to the video type associated with the user-triggered selection control within the first display interface, wherein the material determination control includes one or more of a text input control, a character determination control, and a music determination control; and obtain the material content determined by the user based on the material determination control.
[0158] Further, based on any of the above embodiments, the acquisition module is configured to: in response to the user triggering the role determination control, display a material determination interface. The material determination interface displays multiple candidate images within a preset first storage path and candidate images associated with the user's associated avatar, where the avatar is generated by the large language model based on the image material input by the user. In response to the user triggering a selection operation within the material interface, the person in the target candidate image selected by the user is determined as the role material. And / or, the acquisition module is configured to: acquire the video description text material input by the user via the text input control. And / or, the acquisition module is configured to: in response to the user triggering the music determination control, display a material determination interface. The material determination interface displays all candidate audio files within a preset second storage path. In response to the user triggering a selection operation within the material interface, the target audio file selected by the user is determined as the music material content.
[0159] Furthermore, based on any of the above embodiments, the device further includes: an extraction module for inputting the material content and a preset second prompt content into the large language model to obtain feature information output by the large language model, wherein the feature information includes one or more of character features, character-associated clothing features, location features, and video style features, and the second prompt content is used to prompt the large language model to generate feature information that meets a preset second condition based on the material content. A generation module for inputting the feature information and a preset third prompt content into the large language model to obtain the plot text content output by the large language model, wherein the third prompt content is used to prompt the large language model to expand the feature information according to the preset third condition to generate the plot text content.
[0160] Furthermore, based on any of the above embodiments, the processing module is used to: in response to the user triggering the preset screen generation control in the first display interface, input the plot text content and the preset third prompt content into the large language model, and obtain a preset number of storyboards output by the large language model. The first prompt content is used to prompt the large language model to generate storyboards that meet a preset first condition based on the plot text content, wherein the first condition includes one or more of the number of storyboards, the screen type corresponding to each storyboard, the proportion information corresponding to different types of storyboards, the display content information corresponding to each storyboard, and the correspondence between each storyboard and the lyrics in the music material content. The screen type includes one or more of the blank shot type, the lip sync type, and the character type.
[0161] Furthermore, based on any of the above embodiments, the apparatus further comprises: a display module configured to display a preset regeneration control in a display area associated with each storyboard; a processing module configured to, in response to a user triggering the regeneration control associated with any storyboard, re-input the first prompt content associated with the storyboard triggered by the user and the plot text content into the large language model to obtain an updated storyboard output by the large language model; and a switching module configured to switch the display of the updated storyboard to the position of the storyboard triggered by the user.
[0162] Further, based on any of the above embodiments, the generation module is used to: determine the associated information of each storyboard screen, the associated information includes the screen type associated with the storyboard screen, the display order associated with the storyboard screen, and the lyrics text paragraph corresponding to the storyboard screen, and the lyrics text paragraph is obtained by identifying the content of the music material. For each storyboard screen, determine the video processing method corresponding to the screen type associated with the storyboard screen, perform video processing on the storyboard screen according to the video processing method, and obtain the video paragraph corresponding to the storyboard screen. Splice the video paragraphs corresponding to the preset number of storyboard screens according to the display order associated with the storyboard screen to obtain a video to be processed. Add lyrics subtitles to the video to be processed according to the lyrics text paragraph corresponding to the storyboard screen to obtain the target video.
[0163] Furthermore, based on any of the above embodiments, the screen type includes an empty shot type, and the associated information of the storyboard screen also includes a portion of the plot text content associated with the storyboard screen. The generation module is configured to input the storyboard screen, the portion of the plot text content associated with the storyboard screen, and a fifth prompt content into the large language model, and obtain a video segment corresponding to the storyboard screen output by the large language model, wherein the fifth prompt content is used to prompt the large language model to generate an empty shot video that meets a preset fifth condition based on the current input content.
[0164] Furthermore, based on any of the above embodiments, the picture type includes a lip-sync type, and the associated information of the storyboard picture also includes part of the plot text content associated with the storyboard picture. The generation module is used to: perform a face-changing operation on the face area in the storyboard picture based on the face area in the character material determined by the user, and obtain a processed face-changing image. The face-changing image, the part of the plot text content associated with the storyboard picture, the lyrics text paragraph corresponding to the storyboard picture, and a preset fifth prompt content are input into the large language model to obtain a video paragraph corresponding to the storyboard picture output by the large language model, wherein the fifth prompt content is used to prompt the large language model to generate a video paragraph with a lip-sync effect based on the current input content.
[0165] Furthermore, based on any of the above embodiments, the picture type includes a portrait type, and the associated information of the storyboard picture also includes a portion of the plot text content associated with the storyboard picture. The generation module is used to: perform a face-changing operation on the face area in the storyboard picture based on the face area in the character material determined by the user, and obtain a processed face-changing image. The face-changing image, the portion of the plot text content associated with the storyboard picture, and a preset sixth prompt content are input into the large language model to obtain a video segment corresponding to the storyboard picture output by the large language model, wherein the sixth prompt content is used to prompt the large language model to generate a portrait video segment that meets the preset fifth condition based on the current input content.
[0166] Furthermore, based on any of the above embodiments, the device further includes: a display module, configured to generate the target video in the background, and display a preset content display interface, and display target videos generated historically by other users in the content display interface.
[0167] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0168] In order to implement the above embodiments, the embodiments of the present disclosure further provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video generation method described in any of the above embodiments is implemented.
[0169] In order to implement the above embodiments, the embodiments of the present disclosure further provide a computer program product, including a computer program, which implements the video generation method as described in any of the above embodiments when executed by a processor.
[0170] In order to implement the above embodiment, the present disclosure further provides an electronic device, including: a processor and a memory;
[0171] The memory stores computer-executable instructions;
[0172] The processor executes the computer-executable instructions stored in the memory, so that the processor performs the video generation method as described in any of the above embodiments.
[0173] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. The electronic device 800 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0174] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0175] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0176] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0177] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0178] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0179] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0180] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0182] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0183] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0184] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0185] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0186] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0187] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A video generation method, characterized in that: include: In response to a user-triggered video generation operation, obtaining material content determined by the user, the material content including one or more of character material, music material content, and video description text material; In response to a preview operation triggered by the user, displaying plot text content used to generate a target video in a first display interface, where the plot text content is generated by a large language model based on the source content; In response to the user-triggered storyboard generation operation, a preset number of storyboards generated based on the plot text content and a preset first prompt content are displayed in the first display interface, where the first prompt content is used to prompt the large language model to generate storyboards that meet a preset first condition based on the plot text content; In response to the video generation operation triggered by the user, the target video is generated based on the preset number of storyboards.
2. The method according to claim 1, characterized in that The step of obtaining the material content determined by the user in response to the video generation operation triggered by the user includes: Displaying selection controls associated with multiple video types in the second display interface; In response to a triggering operation by the user on a selection control associated with any video type, displaying the first display interface; Displaying, in the first display interface, a material determination control corresponding to the video type associated with the selection control triggered by the user, wherein the material determination control includes one or more of a text input control, a role determination control, and a music determination control; The material content determined by the user based on the material determination control is acquired.
3. The method according to claim 2, characterized in that The acquiring the material content determined by the user based on the material determination control includes: In response to the user triggering the role determination control, displaying a material determination interface; Displaying, in the material determination interface, a plurality of images to be selected in a preset first storage path and images to be selected associated with the user's associated avatar, where the avatar is generated by the large language model based on the image material input by the user; In response to a selection operation triggered by the user in the material interface, determining the person in the target image to be selected selected by the user as the character material; And / or, obtaining the material content determined by the user based on the material determination control includes: Obtaining video description text material input by the user based on the text input control; And / or, obtaining the material content determined by the user based on the material determination control includes: In response to the user triggering the music determination control, displaying a material determination interface; Displaying all the audios to be selected in the preset second storage path in the material determination interface; In response to a selection operation triggered by the user in the material interface, the target audio selected by the user is determined as the music material content.
4. The method according to claim 1, wherein In response to the preview operation triggered by the user, before displaying the plot text content used to generate the target video in the first display interface, the method further includes: Inputting the source content and a preset second prompt content into the large language model to obtain feature information output by the large language model, the feature information including one or more of character features, character-associated clothing features, location features, and video style features, wherein the second prompt content is used to prompt the large language model to generate feature information that meets a preset second condition based on the source content; The feature information and a preset third prompt content are input into the large language model to obtain the plot text content output by the large language model, wherein the third prompt content is used to prompt the large language model to expand the feature information according to a preset third condition to generate the plot text content.
5. The method according to claim 1, characterized in that In response to the storyboard generation operation triggered by the user, displaying a preset number of storyboards generated based on the plot text content and preset first prompt content in the first display interface includes: In response to the user triggering an operation on a preset screen generation control in the first display interface, inputting plot text content and preset third prompt content into the large language model, and obtaining a preset number of storyboards output by the large language model; Among them, the first prompt content is used to prompt the large language model to generate a storyboard screen that meets a preset first condition based on the plot text content. The first condition includes one or more of the number of storyboard screens, the screen type corresponding to each storyboard screen, the proportion information corresponding to different types of storyboard screens, the display content information corresponding to each storyboard screen, and the correspondence between each storyboard screen and the lyrics in the music material content. The screen type includes one or more of the empty shot type, lip-sync type, and character type.
6. The method according to claim 1, characterized in that After the storyboard generation operation triggered by the user is displayed in the first display interface, a preset number of storyboards generated based on the plot text content and the preset first prompt content, the method further includes: Display preset regeneration controls in the display area associated with each storyboard; In response to a user triggering an operation of a regeneration control associated with any storyboard, re-inputting the first prompt content associated with the storyboard triggered by the user and the plot text content into the large language model to obtain an updated storyboard output by the large language model; The updated storyboard is displayed by switching at the position of the storyboard triggered by the user.
7. The method according to any one of claims 1 to 6, characterized in that The step of generating the target video based on the preset number of storyboards in response to the user-triggered video generation operation includes: For each storyboard, determining association information of the storyboard, the association information including a screen type associated with the storyboard, a display order associated with the storyboard, and a lyric text paragraph corresponding to the storyboard, the lyric text paragraph being obtained by identifying the content of the music material; For each storyboard picture, determining a video processing method corresponding to the picture type associated with the storyboard picture, performing video processing on the storyboard picture according to the video processing method, and obtaining a video segment corresponding to the storyboard picture; splicing the video segments corresponding to the preset number of storyboards according to the display order of the storyboards to obtain a video to be processed; Lyrics subtitles are added to the video to be processed according to the lyrics text paragraphs corresponding to the storyboard images to obtain the target video.
8. The method according to claim 7, characterized in that The picture type includes an empty shot type, and the associated information of the storyboard picture also includes part of the plot text content associated with the storyboard picture; The step of performing video processing on the storyboard according to the video processing method to obtain a video segment corresponding to the storyboard includes: The storyboard, part of the plot text content associated with the storyboard, and the fifth prompt content are input into the large language model to obtain a video segment corresponding to the storyboard output by the large language model, wherein the fifth prompt content is used to prompt the large language model to generate an empty shot video that meets the preset fifth condition based on the current input content.
9. The method according to claim 7, characterized in that The picture type includes a lip sync type, and the associated information of the storyboard picture also includes part of the plot text content associated with the storyboard picture; The step of performing video processing on the storyboard according to the video processing method to obtain a video segment corresponding to the storyboard includes: Performing a face-swapping operation on the face area in the storyboard based on the face area in the character material determined by the user to obtain a processed face-swapping image; The face-swapped image, part of the plot text content associated with the storyboard, the lyrics text paragraph corresponding to the storyboard, and the preset fifth prompt content are input into the large language model to obtain a video paragraph corresponding to the storyboard output by the large language model, wherein the fifth prompt content is used to prompt the large language model to generate a video paragraph with a lip-syncing effect based on the current input content.
10. The method according to claim 7, characterized in that The picture type includes a portrait type, and the associated information of the storyboard picture also includes part of the plot text content associated with the storyboard picture; The step of performing video processing on the storyboard according to the video processing method to obtain a video segment corresponding to the storyboard includes: Performing a face-swapping operation on the face area in the storyboard based on the face area in the character material determined by the user to obtain a processed face-swapping image; The face-swapped image, part of the plot text content associated with the storyboard, and the preset sixth prompt content are input into the large language model to obtain a video segment corresponding to the storyboard output by the large language model, wherein the sixth prompt content is used to prompt the large language model to generate a portrait video segment that meets the preset fifth condition based on the current input content.
11. The method according to any one of claims 1 to 6, characterized in that: After generating the target video based on the preset number of storyboards in response to the user-triggered video generation operation, the method further includes: The target video is generated in the background, and a preset content display interface is displayed, in which target videos generated historically by other users are displayed.
12. A video generating device, characterized in that: include: an acquisition module, configured to acquire, in response to a user-triggered video generation operation, material content determined by the user, the material content including one or more of character material, music material content, and video description text material; a display module, configured to display, in response to a preview operation triggered by the user, plot text content used to generate a target video in a first display interface, wherein the plot text content is generated by a large language model based on the source content; a processing module configured to, in response to the storyboard generation operation triggered by the user, display, in the first display interface, a preset number of storyboards generated based on the plot text content and a preset first prompt content, wherein the first prompt content is configured to prompt the large language model to generate storyboards that meet a preset first condition based on the plot text content; A generation module is used to generate the target video based on the preset number of storyboards in response to the video generation operation triggered by the user.
13. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the video generation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the video generation method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the video generating method according to any one of claims 1 to 11 is implemented.
Citation Information
Cited By
Video slicing method and system and readable storage medium
CN121462854A