Video generation method, device, equipment and storage medium

The video generation method with automatic display and user confirmation solves the problem of long video generation time in the prior art, achieves efficient generation and improves video quality.

CN120281994BActive Publication Date: 2025-09-30BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510704482.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-30
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing technologies require professional teams and equipment to generate videos, and the generation process is time-consuming, resulting in low efficiency.

Method used

A video generation method is provided, which automatically displays multiple storyboards and storyboard images in response to video description information input by a user, and allows the user to confirm and modify them until a final video is generated.

Benefits of technology

The efficiency and quality of video generation are improved, and users can participate in the generation process to ensure that the video meets their personal needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281994B_ABST
    Figure CN120281994B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, device, and storage medium, and relates to the field of artificial intelligence technology. The method can generate multiple storyboards based on video description information input by a user. After the user confirms the storyboards, multiple storyboard images are generated. After the user confirms the storyboard images, the final video is generated. The method only requires the user to input the video description information to automatically generate the video, thereby improving the efficiency of video generation. In addition, the generated storyboards and storyboard images are displayed midway through video generation, and the video is generated after the user confirms the video. This improves user participation and the controllability of the generated video. This not only improves the efficiency of video generation, but also ensures that the generated video better meets user needs, thereby improving the quality of the generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a video generation method, apparatus, device, and storage medium. Background Art

[0002] In related technologies, generating videos requires steps such as script writing, shooting, and editing, which not only requires a professional team and equipment, but also the generation process is time-consuming, resulting in low efficiency in video generation. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, device, and storage medium, which improves video generation efficiency and video generation quality. The technical solution of the present disclosure is as follows.

[0004] According to one aspect of an embodiment of the present disclosure, a video generation method is provided, the method comprising:

[0005] In response to a confirmation operation of the video description information input on the application interface, displaying a plurality of storyboards generated based on the video description information, the plurality of storyboards corresponding to at least one storyline;

[0006] In response to a confirmation operation on the plurality of storyboards, displaying a plurality of storyboard screens generated based on the plurality of storyboards, wherein each storyboard script corresponds to at least one storyboard screen;

[0007] In response to a confirmation operation on the plurality of storyboards, a video generated based on the plurality of storyboards is displayed.

[0008] In some embodiments, before displaying the plurality of storyboards generated based on the plurality of storyboards in response to the confirmation operation on the plurality of storyboards, the method further includes:

[0009] displaying at least one first character and a first screen style, wherein the at least one first character and the first screen style are both determined based on the plurality of storyboards;

[0010] The method of displaying a plurality of storyboards generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards comprises:

[0011] In response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, a plurality of storyboard screens generated based on the plurality of storyboards, the at least one first character, and the first picture style are displayed.

[0012] In some embodiments, before displaying the plurality of storyboards generated based on the plurality of storyboards in response to the confirmation operation on the plurality of storyboards, the method further includes:

[0013] In response to a modification operation on at least one storyboard among the plurality of storyboards, displaying at least one second character and a second screen style, wherein both the at least one second character and the second screen style are determined based on the modified storyboard;

[0014] The method of displaying a plurality of storyboards generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards comprises:

[0015] In response to a confirmation operation on the modified storyboard, the at least one second character, and the second picture style, a plurality of storyboards generated based on the modified storyboard, the at least one second character, and the second picture style are displayed.

[0016] In some embodiments, in response to a modification operation on at least one storyboard among the plurality of storyboards, displaying at least one second character and a second screen style comprises:

[0017] In response to a triggering operation on a replacement control on the application interface, displaying a plurality of storyboards regenerated by artificial intelligence based on the video description information;

[0018] In response to a confirmation operation on the regenerated plurality of storyboards, the at least one second character and the second screen style determined based on the regenerated plurality of storyboards are displayed.

[0019] In some embodiments, in response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, before displaying a plurality of storyboard images generated based on the plurality of storyboards, the at least one first character, and the first picture style, the method further includes:

[0020] In response to a modification operation on at least one of the at least one first character and the first picture style, displaying the modified character and picture style;

[0021] The method of displaying, in response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, a plurality of storyboard pictures generated based on the plurality of storyboards, the at least one first character, and the first picture style, comprises:

[0022] In response to a confirmation operation on the plurality of storyboards and the modified characters and picture styles, a plurality of storyboard screens generated based on the plurality of storyboards and the modified characters and picture styles are displayed.

[0023] In some embodiments, in response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style includes at least one of the following:

[0024] In response to a triggering operation on any of the at least one first character, an editing component of the triggered first character is displayed, wherein the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, the modified first character is displayed;

[0025] In response to a role adding operation on the interface where the at least one first role is located, the added role is displayed in the at least one first role.

[0026] In some embodiments, the first picture style is highlighted in a picture style list, the picture style list includes multiple picture styles, and the picture style highlighted in the picture style list is a picture style determined based on the multiple storyboards;

[0027] In response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style includes:

[0028] In response to a triggering operation on any picture style other than the first picture style in the picture style list, the triggered picture style is highlighted to obtain a modified picture style.

[0029] In some embodiments, the displaying of the plurality of storyboards generated based on the plurality of storyboard scripts includes:

[0030] Thumbnails of the plurality of storyboards are displayed on a screen preview interface. In response to a triggering operation on a first thumbnail among the plurality of thumbnails, a first storyboard corresponding to the first thumbnail is displayed in an enlarged manner on the screen preview interface.

[0031] In some embodiments, the method further comprises:

[0032] Displaying first picture information of the first storyboard picture on the picture preview interface, the first picture information including at least one of picture description information, character information, and subtitle information;

[0033] In response to a modification operation on the first picture information of the first storyboard picture, a first storyboard picture regenerated based on the modified first picture information is displayed.

[0034] In some embodiments, the method further comprises:

[0035] In response to a triggering operation on any area on the first storyboard screen on the screen preview interface, displaying at least one modification option for the triggered area, wherein the at least one modification option is used to modify the triggered area in at least one dimension;

[0036] In response to a triggering operation on any one of the at least one modification option, the first storyboard screen modified based on the triggered modification option is displayed.

[0037] In some embodiments, the method further comprises:

[0038] In response to a replacement operation on a first storyboard on the picture preview interface, displaying a plurality of candidate storyboards;

[0039] In response to a triggering operation on any candidate storyboard among the plurality of candidate storyboards, the first storyboard is replaced with the triggered candidate storyboard.

[0040] In some embodiments, displaying a video generated based on the plurality of storyboards includes:

[0041] Displaying the video on a video preview interface, wherein the video preview interface also displays thumbnails of the plurality of storyboards;

[0042] The method further comprises:

[0043] In response to a triggering operation on a second thumbnail among the plurality of thumbnails, a second storyboard corresponding to the second thumbnail is displayed in an enlarged manner on the video preview interface, and second picture information of the second storyboard is displayed on the video preview interface, where the second picture information includes at least one of picture description information, character information, subtitle information, music information, and camera movement mode;

[0044] In response to a modification operation on the second picture information of the second storyboard picture, a second storyboard picture regenerated based on the modified second picture information is displayed, and the regenerated second storyboard picture is used to regenerate the video.

[0045] In some embodiments, the video description information includes characteristic indication information of the video to be generated, and in response to a confirmation operation of the video description information input on the application interface, displaying multiple storyboards generated based on the video description information includes:

[0046] In response to a confirmation operation of the video description information input on the application interface, displaying a first story generated by artificial intelligence based on the video description information;

[0047] In response to a confirmation operation on the first story, multiple storyboards generated based on the first story are displayed; or, in response to a modification operation on the first story, multiple storyboards generated based on the modified first story are displayed.

[0048] In some embodiments, the video description information includes a second story, and the video description information is used to indicate that the video is generated based on the second story.

[0049] In some embodiments, in response to a confirmation operation on the plurality of storyboards, displaying a video generated based on the plurality of storyboards includes:

[0050] In response to a confirmation operation on the multiple storyboards, multiple videos generated based on the multiple storyboards are displayed, wherein the third screen information of at least one storyboard is different between the multiple videos, and the third screen information includes at least one of subtitle information, music information, and camera movement method.

[0051] In some embodiments, before displaying the video generated based on the multiple storyboards in response to the confirmation operation on the multiple storyboards, the method further includes:

[0052] In response to an operation of adjusting the order of the plurality of storyboard pictures, displaying the plurality of storyboard pictures after the order is adjusted;

[0053] The step of displaying a video generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards comprises:

[0054] In response to a confirmation operation on the plurality of storyboards after the order is adjusted, a video generated based on the plurality of storyboards after the order is adjusted is displayed.

[0055] According to another aspect of the present disclosure, a video generation device is provided, the device comprising:

[0056] A first display unit is configured to, in response to a confirmation operation on the video description information input on the application interface, display a plurality of storyboards generated based on the video description information, the plurality of storyboards corresponding to at least one storyline;

[0057] a second display unit configured to, in response to a confirmation operation on the plurality of storyboards, display a plurality of storyboard screens generated based on the plurality of storyboards, each storyboard corresponding to at least one storyboard screen;

[0058] The third display unit is configured to display a video generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards.

[0059] In some embodiments, the first display unit is further configured to perform:

[0060] displaying at least one first character and a first screen style, wherein the at least one first character and the first screen style are both determined based on the plurality of storyboards;

[0061] The second display unit is configured to perform:

[0062] In response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, a plurality of storyboard screens generated based on the plurality of storyboards, the at least one first character, and the first picture style are displayed.

[0063] In some embodiments, the apparatus further includes a first modifying unit configured to perform:

[0064] In response to a modification operation on at least one storyboard among the plurality of storyboards, displaying at least one second character and a second screen style, wherein both the at least one second character and the second screen style are determined based on the modified storyboard;

[0065] The second display unit is configured to perform:

[0066] In response to a confirmation operation on the modified storyboard, the at least one second character, and the second picture style, a plurality of storyboards generated based on the modified storyboard, the at least one second character, and the second picture style are displayed.

[0067] In some embodiments, the first modifying unit is configured to perform:

[0068] In response to a triggering operation on a replacement control on the application interface, displaying a plurality of storyboards regenerated by artificial intelligence based on the video description information;

[0069] In response to a confirmation operation on the regenerated plurality of storyboards, the at least one second character and the second screen style determined based on the regenerated plurality of storyboards are displayed.

[0070] In some embodiments, the apparatus further includes a second modifying unit configured to perform:

[0071] In response to a modification operation on at least one of the at least one first character and the first picture style, displaying the modified character and picture style;

[0072] The second display unit is configured to perform:

[0073] In response to a confirmation operation on the plurality of storyboards and the modified characters and picture styles, a plurality of storyboard screens generated based on the plurality of storyboards and the modified characters and picture styles are displayed.

[0074] In some embodiments, the second modifying unit is configured to perform at least one of the following:

[0075] In response to a triggering operation on any of the at least one first character, an editing component of the triggered first character is displayed, wherein the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, the modified first character is displayed;

[0076] In response to a role adding operation on the interface where the at least one first role is located, the added role is displayed in the at least one first role.

[0077] In some embodiments, the first picture style is highlighted in a picture style list, the picture style list includes multiple picture styles, and the picture style highlighted in the picture style list is a picture style determined based on the multiple storyboards;

[0078] The second modifying unit is configured to execute:

[0079] In response to a triggering operation on any picture style other than the first picture style in the picture style list, the triggered picture style is highlighted to obtain a modified picture style.

[0080] In some embodiments, the second display unit is configured to perform:

[0081] Thumbnails of the plurality of storyboards are displayed on a screen preview interface. In response to a triggering operation on a first thumbnail among the plurality of thumbnails, a first storyboard corresponding to the first thumbnail is displayed in an enlarged manner on the screen preview interface.

[0082] In some embodiments, the second display unit is further configured to perform:

[0083] Displaying first picture information of the first storyboard picture on the picture preview interface, the first picture information including at least one of picture description information, character information, and subtitle information;

[0084] In response to a modification operation on the first picture information of the first storyboard picture, a first storyboard picture regenerated based on the modified first picture information is displayed.

[0085] In some embodiments, the apparatus further includes a third modifying unit configured to execute:

[0086] In response to a triggering operation on any area on the first storyboard screen on the screen preview interface, displaying at least one modification option for the triggered area, wherein the at least one modification option is used to modify the triggered area in at least one dimension;

[0087] In response to a triggering operation on any one of the at least one modification option, the first storyboard screen modified based on the triggered modification option is displayed.

[0088] In some embodiments, the apparatus further comprises a replacing unit configured to perform:

[0089] In response to a replacement operation on a first storyboard on the picture preview interface, displaying a plurality of candidate storyboards;

[0090] In response to a triggering operation on any candidate storyboard among the plurality of candidate storyboards, the first storyboard is replaced with the triggered candidate storyboard.

[0091] In some embodiments, the third display unit is configured to perform:

[0092] Displaying the video on a video preview interface, wherein the video preview interface also displays thumbnails of the plurality of storyboards;

[0093] The apparatus further includes a fourth modifying unit configured to execute:

[0094] In response to a triggering operation on a second thumbnail among the plurality of thumbnails, a second storyboard corresponding to the second thumbnail is displayed in an enlarged manner on the video preview interface, and second picture information of the second storyboard is displayed on the video preview interface, where the second picture information includes at least one of picture description information, character information, subtitle information, music information, and camera movement mode;

[0095] In response to a modification operation on the second picture information of the second storyboard picture, a second storyboard picture regenerated based on the modified second picture information is displayed, and the regenerated second storyboard picture is used to regenerate the video.

[0096] In some embodiments, the video description information includes characteristic indication information of the video to be generated, and the first display unit is configured to execute:

[0097] In response to a confirmation operation of the video description information input on the application interface, displaying a first story generated by artificial intelligence based on the video description information;

[0098] In response to a confirmation operation on the first story, multiple storyboards generated based on the first story are displayed; or, in response to a modification operation on the first story, multiple storyboards generated based on the modified first story are displayed.

[0099] In some embodiments, the video description information includes a second story, and the video description information is used to indicate that the video is generated based on the second story.

[0100] In some embodiments, the third display unit is configured to perform:

[0101] In response to a confirmation operation on the multiple storyboards, multiple videos generated based on the multiple storyboards are displayed, wherein the third screen information of at least one storyboard is different between the multiple videos, and the third screen information includes at least one of subtitle information, music information, and camera movement method.

[0102] In some embodiments, the apparatus further comprises an adjusting unit configured to perform:

[0103] In response to an operation of adjusting the order of the plurality of storyboard pictures, displaying the plurality of storyboard pictures after the order is adjusted;

[0104] The third display unit is configured to perform:

[0105] In response to a confirmation operation on the plurality of storyboards after the order is adjusted, a video generated based on the plurality of storyboards after the order is adjusted is displayed.

[0106] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device including:

[0107] processor;

[0108] a memory for storing instructions executable by the processor;

[0109] The processor is configured to execute the instructions to implement the above-mentioned video generation method.

[0110] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above-mentioned video generation method.

[0111] According to another aspect of an embodiment of the present disclosure, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the video generation method is implemented.

[0112] The disclosed embodiments provide a video generation method, which can generate multiple storyboards based on video description information input by a user, generate multiple storyboards after the user confirms the storyboards, and generate a final video after the user confirms the storyboards. The method can automatically generate video only by the user inputting the video description information, thereby improving the efficiency of video generation; and, in the middle of video generation, the generated storyboards and storyboards are displayed, and the video is generated after the user confirms, thereby improving the user's participation, that is, improving the user's controllability over the generated video. This not only improves the efficiency of video generation, but also enables the generated video to better meet user needs, thereby improving the quality of the generated video.

[0113] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0115] Figure 1 It is a schematic diagram showing an implementation environment according to an exemplary embodiment.

[0116] Figure 2 The figure is a flowchart of a video generating method according to an exemplary embodiment.

[0117] Figure 3 The figure is a flowchart of another method for generating a video according to an exemplary embodiment.

[0118] Figure 4 It is a schematic diagram showing an application interface according to an exemplary embodiment.

[0119] Figure 5 is a schematic diagram of another application interface according to an exemplary embodiment.

[0120] Figure 6 The figure is a schematic diagram showing a screen preview result according to an exemplary embodiment.

[0121] Figure 7 The figure is a schematic diagram showing a video generation interface according to an exemplary embodiment.

[0122] Figure 8 The figure is a schematic diagram of a video preview interface according to an exemplary embodiment.

[0123] Figure 9 is a schematic diagram of another video preview interface according to an exemplary embodiment.

[0124] Figure 10 The figure is a block diagram of a video generating apparatus according to an exemplary embodiment.

[0125] Figure 11 It is a block diagram of a terminal according to an exemplary embodiment. DETAILED DESCRIPTION

[0126] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0127] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0128] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the video description information involved in this disclosure was obtained with full authorization.

[0129] The video generation method provided by the embodiment of the present disclosure can be executed by an electronic device, and the electronic device can be provided as at least one of a terminal and a server. Figure 1 This is a schematic diagram of an implementation environment provided by the embodiment of the present disclosure, see Figure 1 , the implementation environment includes: a terminal 101 and a server 102.

[0130] In the embodiment of the present disclosure, a target application is installed on the terminal 101, and the target application is used to automatically generate a video. The server 102 is a background server of the target application, and is used to provide background services for the target application.

[0131] In some embodiments, after inputting video description information on the application interface of terminal 101, the user confirms the video description information, and multiple storyboards corresponding to at least one storyline are automatically generated based on the video description information, and terminal 101 displays the multiple storyboards; and after the user confirms the multiple storyboards, multiple storyboard screens generated based on the multiple storyboards are displayed, and one storyboard script can correspond to one or more storyboard screens; finally, after the user confirms the multiple storyboard screens, terminal 101 displays the video generated based on the multiple storyboard screens.

[0132] Terminal 101 can be at least one of a smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and portable computer. Terminal 101 has communication capabilities and can access a wired or wireless network. Terminal 101 can generally refer to one of multiple terminals, and those skilled in the art will appreciate that the number of terminals can be greater or lesser. Server 102 can be a standalone physical server, a server cluster consisting of multiple physical servers, or a distributed file system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, server 102 is directly or indirectly connected to terminal 101 via wired or wireless communication, which is not limited in this embodiment. Optionally, the number of servers 102 can be greater or lesser, which is not limited in this embodiment. Of course, server 102 can also include other functional servers to provide more comprehensive and diverse services. Among them, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 or the terminal 101 can each undertake the computing work independently, and the embodiments of the present disclosure do not limit this.

[0133] Figure 2 is a flow chart of a video generation method according to an exemplary embodiment. Figure 2 As shown, the method is executed by the terminal, and the method includes the following steps.

[0134] In step S201, in response to a confirmation operation on video description information input on an application interface, the terminal displays a plurality of storyboards generated based on the video description information, wherein the plurality of storyboards correspond to at least one storyline.

[0135] In the embodiment of the present disclosure, the application interface is an interface within the target application, and the target application is used to generate a video based on the video description information input by the user. Furthermore, the target application generates the video through artificial intelligence, and the artificial intelligence generates multiple storyboards based on the video description information.

[0136] In some embodiments, an information input box is displayed on the application interface, which is used to input video description information, and a first confirmation control is displayed on the application interface. After the user confirms the video description information by triggering the first confirmation control, the terminal generates multiple storyboards based on the video description information and displays these multiple storyboards.

[0137] Alternatively, if the video is generated by artificial intelligence, the application interface is a dialogue interface between the user and the artificial intelligence character. The dialogue interface displays a dialogue input box for entering video description information, and a send control is displayed on the dialogue interface. Confirming the video description information triggers the send control. Accordingly, the multiple storyboards are the dialogue messages responded to by the artificial intelligence character on the dialogue interface.

[0138] In the disclosed embodiments, video description information is used to indicate the characteristics of the video to be generated, that is, to indicate the type of video to be generated. A storyboard is a visual script that presents a story or scene through a brief text description. A storyboard may include one or more of the following: a description of the shot content, the character's lines in the shot, the character's actions, the size and range of the subject in the shot, camera movement, sound effects, frame duration, visual effects, and notes.

[0139] In the embodiment of the present disclosure, multiple storyboards correspond to at least one storyline, and these storylines are obtained by artificial intelligence based on video description information, and then artificial intelligence generates multiple storyboards around these storylines.

[0140] The multiple storyboards corresponding to the same plot are arranged in chronological order, and the multiple storyboards sequentially present the development of the story, that is, according to the timeline, the multiple storyboards present the cause, process, and outcome of the story. Furthermore, the multiple storyboards corresponding to the same plot share at least one common element, such as the same characters, scenes, or background music.

[0141] In some embodiments, the multiple storyboards generated by AI can correspond to multiple plots, with each plot corresponding to at least one storyboard. These multiple plots can be connected through various means, such as theme, causality, characters, timelines, and locations, ensuring coherence between the multiple plots. For example, multiple plots may revolve around a single theme; or the outcome of one plot may be the cause of another; or multiple plots may revolve around the same character; or multiple plots may occur at different points in time, but form a complete timeline; or multiple plots may take place in the same location.

[0142] In step S202, in response to a confirmation operation on a plurality of storyboard scripts, the terminal displays a plurality of storyboard screens generated based on the plurality of storyboard scripts, where each storyboard script corresponds to at least one storyboard screen.

[0143] In some embodiments, a second confirmation control is displayed on the interface where the multiple storyboards are located, and the confirmation operation on the multiple storyboards is also a triggering operation on the second confirmation control.

[0144] Optionally, multiple storyboards are generated by artificial intelligence based on multiple storyboard scripts, and the artificial intelligence generates at least one storyboard for each storyboard script.

[0145] In step S203, in response to the confirmation operation on the plurality of storyboards, the terminal displays a video generated based on the plurality of storyboards.

[0146] In some embodiments, a third confirmation control is displayed on the interface where the multiple storyboard images are located, and the confirmation operation on the multiple storyboard images is also a triggering operation on the third confirmation control.

[0147] In the embodiments of the present disclosure, displaying a video on a terminal may refer to displaying a video cover, and triggering the cover may play the video. Displaying a video on a terminal may also refer to playing the video. The generated video may be an animation video, a comic video, or a cartoon video.

[0148] Optionally, the terminal generates a video based on multiple storyboards using artificial intelligence. In some embodiments, generating a video based on multiple storyboards can be performed by directly splicing the multiple storyboards to obtain a video. In other embodiments, storyboards are generated by interpolating between each two adjacent storyboards, so that the interpolated storyboards connect the two adjacent storyboards to further enrich the video content. Alternatively, for at least some of the multiple storyboards, a dynamic image is generated based on each storyboard in the at least some storyboards, each dynamic image is also a video clip, and then a video is obtained based on the remaining storyboards in the multiple storyboards and the video clips corresponding to the at least some storyboards, such as splicing the storyboards and video clips in chronological order to obtain a video.

[0149] The disclosed embodiments provide a video generation method, which can generate multiple storyboards based on video description information input by a user, generate multiple storyboards after the user confirms the storyboards, and generate a final video after the user confirms the storyboards. The method can automatically generate video only by the user inputting the video description information, thereby improving the efficiency of video generation; and, in the middle of video generation, the generated storyboards and storyboards are displayed, and the video is generated after the user confirms, thereby improving the user's participation, that is, improving the user's controllability over the generated video. This not only improves the efficiency of video generation, but also enables the generated video to better meet user needs, thereby improving the quality of the generated video.

[0150] above Figure 2 The following is only the basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 , Figure 3 The flowchart of another video generation method according to an exemplary embodiment is shown. The method is executed by a terminal and includes at least one of the following steps.

[0151] In step S301, in response to the confirmation operation of the video description information input on the application interface, the terminal displays multiple storyboards generated based on the video description information and displays at least one first character and a first picture style. The multiple storyboards correspond to at least one storyline, and the at least one first character and the first picture style are both determined based on the multiple storyboards.

[0152] In some embodiments, the video description information includes characteristic indication information of the video to be generated, that is, the video description information is used to indicate what kind of video to be generated. The characteristic indication information may indicate one or more of the story theme, plot, character characteristics, and screen characteristics of the video to be generated. When generating a video based on the video description information including the characteristic indication information, a story may be generated based on the video description information first. Accordingly, the above-mentioned process of displaying multiple storyboards generated based on the video description information by the terminal in response to a confirmation operation on the video description information input on the application interface includes the following steps: in response to a confirmation operation on the video description information input on the application interface, the terminal displays a first story generated by artificial intelligence based on the video description information; in response to a confirmation operation on the first story, the terminal displays multiple storyboards generated based on the first story; or, in response to a modification operation on the first story, the terminal displays multiple storyboards generated based on the modified first story.

[0153] In this embodiment, when the video description information indicates characteristic information for a video to be generated, a story is first generated based on the video description information and displayed. Once the user confirms the generated story, a corresponding storyboard is generated. This ensures that the generated storyboard meets the user's creative needs and is of higher quality. Furthermore, the generated storyboard can be modified, and a storyboard generated based on the modified storyboard. This allows the user to optimize the story content, providing greater flexibility and controllability, improving both creative efficiency and overall quality.

[0154] In other embodiments, the video description information includes a second story, and the video description information is used to instruct to generate a video based on the second story, that is, to generate a video based on an existing story.

[0155] Optionally, artificial intelligence generates a video based on the second story. If the video description information includes the second story, the artificial intelligence generates multiple storyboards based on the second story, each corresponding to a plurality of plot points in the second story. Furthermore, the video description information includes not only the second story but also story indication information, which indicates features of the video generated based on the second story. For example, the story indication information can be used to indicate visual features, sound effects, and other information of the video to be generated. The artificial intelligence then generates multiple storyboards based on the second story and the story indication information.

[0156] In the disclosed embodiment, the terminal can automatically generate storyboard scripts and storyboard images based on a given story, and then automatically generate a video, without the need for relevant workers to perform script writing, shooting, editing and other steps based on the story, thereby shortening the time spent on video generation and improving the efficiency of video generation.

[0157] In some embodiments, the application interface includes two generation options, the first generation option is used to generate a video based on the video description information including feature indication information, and the second generation option is used to generate a video based on the video description information including a story.

[0158] For example, see Figure 4 , Figure 4 is a schematic diagram illustrating an application interface according to an exemplary embodiment. The application interface may be a dialogue interface. The application interface includes a dialogue area, which includes a first generation option 401 and a second generation option 402. Optionally, after entering video description information in dialogue input box 403, a triggering operation on the first generation option 401 displays a first story generated based on the video description information, while a triggering operation on the second generation option 402 displays multiple storyboards generated based on the video description information.

[0159] In some embodiments, the application interface also includes a result display area, which is used to display information such as the generated storyboard, characters, and screen style. Figure 4 The left area on the application interface is the dialogue area, and the right area on the application interface is the result display area.

[0160] In some embodiments, the generated multiple storyboards can be displayed not only in the form of dialogue messages in the dialogue area, but also in the script editing box in the result display area. Figure 5 , Figure 5 1 is a schematic diagram of an application interface according to an exemplary embodiment. In the dialog area of ​​the application interface, multiple storyboards are displayed in the form of dialog messages. In the result display area of ​​the application interface, a script editing box 501 is displayed, and multiple storyboards are displayed in the script editing box 501.

[0161] In the disclosed embodiments, the character determined based on the storyboard is a virtual character, such as a virtual digital human. Optionally, artificial intelligence (AI) generates at least one character instantly based on multiple storyboards; or, the terminal selects at least one character from multiple preset characters based on multiple storyboards. The AI ​​determines the character based on the descriptions of the character, scene, etc. in the storyboard.

[0162] The terminal displays each character by displaying the name and corresponding photo of each character, including the character's appearance and clothing, to show the character's appearance; further, the character's voice and personality characteristics can also be displayed. Figure 5 , the result display area on the application interface displays multiple roles 502.

[0163] In the disclosed embodiment, the picture style is an indispensable part of the storyboard. It determines the visual presentation of the video through multiple elements such as color, composition, light and shadow, and dynamics, giving the story a unique visual language, enhancing emotional expression and the audience's sense of immersion. The picture style can be defined from multiple perspectives. For example, from the perspective of color, it can include styles such as cool colors, warm colors, contrasting colors, and complementary colors; from the perspective of composition, it can include styles such as realism, fantasy, high contrast, and low contrast; from the perspective of animation, it can include styles such as cartoon, realism, watercolor, and oil painting; from the perspective of details, it can include styles such as realism and abstraction; from the perspective of light and shadow, it can include styles such as natural light, artificial light, highlights, and shadows.

[0164] Displaying a picture style on the terminal refers to displaying the name and schematic diagram of the picture style. Furthermore, the terminal displays the picture styles determined based on multiple storyboards in a picture style list, with the first picture style highlighted in the picture style list. The picture style list includes multiple picture styles, and the highlighted picture style in the picture style list is the picture style determined based on the multiple storyboards. Accordingly, the terminal determines the first picture style from the multiple picture styles in the picture style list based on the multiple storyboards.

[0165] For example, see Figure 5 , a result display area on the application interface displays a picture style list 503 , and the first picture style 504 in the picture style list 503 is highlighted.

[0166] In the disclosed embodiments, users can generate stories through dialogue, which can then be used to create storyboards, or directly use existing stories to generate storyboards. The terminal uses artificial intelligence to automatically extract key information from video descriptions, construct a story framework, and then generate a storyboard. Furthermore, the artificial intelligence analyzes scene descriptions, character behaviors, and dialogue content in the storyboard to determine an appropriate visual style. The artificial intelligence then generates characters based on the story background and character characteristics, thereby improving the quality of the determined visual style and characters.

[0167] In step S302, in response to a confirmation operation on multiple storyboards, at least one first character, and a first picture style, the terminal displays multiple storyboard pictures generated based on the multiple storyboards, at least one first character, and the first picture style.

[0168] In some embodiments, multiple storyboards, a first character, and a first screen style are displayed on the application interface. A second confirmation control is displayed on the application interface. The second confirmation control is used to generate a storyboard screen. Therefore, the confirmation operation in step S302 is also a triggering operation of the second confirmation control. For example, continue to see Figure 5, a second confirmation control 505 is also displayed on the application interface.

[0169] In an embodiment of the present disclosure, the process of the terminal displaying multiple storyboard pictures generated based on the multiple storyboard scripts in response to the confirmation operation of the multiple storyboard scripts is implemented through the above-mentioned step S302. In other embodiments, the storyboard script can also be modified. Accordingly, in response to the modification operation of at least one storyboard script among the multiple storyboard scripts, the terminal displays at least one second character and a second picture style, and the at least one second character and the second picture style are both determined based on the modified storyboard script. The above-mentioned process of the terminal displaying multiple storyboard pictures generated based on the multiple storyboard scripts in response to the confirmation operation of the multiple storyboard scripts includes the following implementation method: in response to the confirmation operation of the modified storyboard script, the confirmation operation of the at least one second character and the second picture style, the terminal displays the multiple storyboard pictures generated based on the modified storyboard script, the at least one second character and the second picture style.

[0170] In this embodiment, the user can modify the generated storyboard script, so that the user can easily adjust various elements in the storyboard script, such as the lens angle, scene atmosphere, character movements, etc., to achieve personalized creation, and improve the user's controllability of the storyboard script. The modified storyboard script can also adaptively regenerate the characters and picture style, and the corresponding storyboard pictures can be generated instantly based on the modified content, thereby improving the efficiency of human-computer interaction.

[0171] It should be noted that if the elements in the storyboard modified by the user are associated with other storyboards, the terminal will automatically adaptively modify the other associated storyboards. For example, if the name of a character in a storyboard is modified, the terminal will automatically adaptively modify the name of the character in other storyboards.

[0172] In some embodiments, multiple storyboards are displayed in a script editing box, and the user can modify the multiple storyboards in the script editing box, and then the terminal generates and displays at least one second character and a second picture style based on the modified storyboard.

[0173] Optionally, the character, visual style, and storyboard are displayed on the same application interface, and the character and visual style are modified as the storyboard is modified, without requiring additional user operations. Alternatively, since the user may make multiple modifications to the storyboard at multiple locations, in order to conserve resources, in response to a confirmation operation on the modified storyboard, the terminal displays at least one second character and second visual style generated based on the modified storyboard.

[0174] In some embodiments, a storyboard replacement control is displayed on the application interface. The process in which the terminal displays at least one second character and a second picture style in response to a modification operation on at least one storyboard among a plurality of storyboards includes the following steps: in response to a triggering operation on the replacement control on the application interface, the terminal displays a plurality of storyboards regenerated by artificial intelligence based on the video description information; in response to a confirmation operation on the regenerated plurality of storyboards, the terminal displays at least one second character and a second picture style determined based on the regenerated plurality of storyboards.

[0175] For example, see Figure 5 , a storyboard replacement control 506 is displayed on the application interface.

[0176] In this embodiment, a storyboard replacement control is provided on the application interface, so that the storyboard can be regenerated by artificial intelligence with one click, and the characters and picture style can be regenerated based on the regenerated storyboard, thereby improving the efficiency of storyboard modification.

[0177] In the above embodiment, replacing multiple storyboards with one click is used as an example. In other embodiments, the application interface also provides a replacement control for each storyboard. In response to triggering the replacement control for any storyboard, a new storyboard generated based on the video description information is displayed. Alternatively, the replacement operation for any storyboard can be a long press or double-click operation on the storyboard, etc., which is not specifically limited here.

[0178] In some embodiments, the terminal can also modify the characters and screen styles determined by the storyboards. That is, in response to a confirmation operation on multiple storyboards, at least one first character, and a first screen style, before the terminal displays multiple storyboard images generated based on the multiple storyboards, at least one first character, and the first screen style, the terminal can display the modified characters and screen styles in response to a modification operation on at least one of the at least one first character and the first screen style. Accordingly, in response to a confirmation operation on the multiple storyboards, the modified characters, and the screen style, the terminal displays the multiple storyboard images generated based on the multiple storyboards, the modified characters, and the screen style.

[0179] In this embodiment, the user can also modify the characters and picture styles automatically determined by the storyboard script, and then generate storyboards based on the modified characters and picture styles. This can increase the user's participation in the generation of storyboards, making the storyboards more personalized, which not only improves the creation efficiency, but also increases the controllability of the storyboards, thereby improving the quality of the generated storyboards.

[0180] In some embodiments, in response to the modification operation on at least one of the at least one first character and the first screen style, the terminal displays the modified character and screen style, including at least one of the following implementations:

[0181] (1) In response to a triggering operation on any of at least one first character, the terminal displays an editing component of the triggered first character, where the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, the terminal displays the modified first character.

[0182] It should be noted that if any of the at least one first character is modified, the terminal can automatically modify the storyboard based on the modified character. For example, if the name of a character is modified, the terminal automatically modifies the name of the character in the storyboard.

[0183] In the embodiment of the present disclosure, the example in which the terminal displays the character refers to displaying the name and photo of the character is used for explanation. Then, the triggering operation on the first character may be a click operation on the name or photo of the first character.

[0184] This embodiment supports real-time preview and modification of automatically generated characters, allowing users to instantly view and optimize character effects during the creation process, improving interaction efficiency. Furthermore, a character editing component is provided, allowing users to adjust the character's name, appearance, clothing, voice, and personality traits as needed, thereby generating a character image that better meets creative needs. This high degree of customization allows each character to have a unique style and expressiveness, thereby improving the quality of the generated characters and, in turn, the quality of the videos generated based on the characters.

[0185] (2) In response to a role adding operation on the interface where the at least one first role is located, the terminal displays the added role in the at least one first role.

[0186] Optionally, the interface where the first character is located is taken as the above application interface as an example for explanation. A character adding control is displayed on the application interface. In response to the triggering operation of the character adding control, the terminal displays the character adding interface. The character adding interface is used to set at least one of the name, appearance, clothing, voice and personality characteristics of the added character. It should be noted that after adding a character to at least one first character, the terminal automatically modifies the storyboard script based on the added character, and then generates the storyboard screen based on the modified storyboard script. For example, continue to refer to Figure 5 , the role display area on the application interface also displays a role adding control 507.

[0187] In this embodiment, not only can the detailed features such as the name, appearance, clothing, voice and personality traits of any character be modified, but new characters can also be added, further increasing the user's controllability over the characters, and facilitating the generation of storyboards that meet creative needs based on the modified characters, thereby generating videos that meet creative needs and improving the quality of the generated videos.

[0188] In some embodiments, the terminal displays multiple candidate information for the narrator's voice on the screen displaying at least one character, including information such as the timbre and personality of the narrator. The user can set any of the candidate information for the narrator's voice on this screen. The terminal can also set other video attributes on this screen, such as the aspect ratio and resolution, which will not be detailed here.

[0189] In some embodiments, the first picture style is highlighted in a picture style list, the picture style list includes multiple picture styles, and the picture style highlighted in the picture style list is a picture style determined based on multiple storyboards; then the above-mentioned process of the terminal displaying the modified character and picture style in response to the modification operation of at least one first character and at least one of the first picture styles includes the following implementation method: in response to the triggering operation of any picture style other than the first picture style in the picture style list, the terminal highlights the triggered picture style to obtain the modified picture style.

[0190] In this embodiment, the picture style determined by the storyboard script is highlighted in the picture style list, and in response to a triggering operation on any other picture style in the picture style list, the triggered picture style is highlighted and used as the modified picture style. In this way, a variety of picture styles are provided through the picture style list, which increases the diversity of choices, and the picture style can also be modified by one-click triggering, which improves the human-computer interaction efficiency when modifying the picture style.

[0191] In an embodiment of the present disclosure, after the user inputs video description information on the application interface, artificial intelligence automatically generates multiple storyboards based on the video description information, generates characters and picture styles based on the storyboards, and displays the storyboards, characters, and picture styles. In this way, more dimensional information of the generated storyboard pictures is displayed, which improves the transparency of the information. After the user confirms this information, the storyboard pictures are regenerated, which improves the user's controllability over the generation of the storyboard pictures and increases the user's participation, thereby making the generated storyboard pictures more in line with the needs and improving the quality of the generated storyboard pictures.

[0192] In step S303, in response to the confirmation operation of multiple storyboard scripts, the terminal displays thumbnails of each of the multiple storyboard screens on the screen preview interface. In response to the trigger operation of the first thumbnail among the multiple thumbnails, the first storyboard screen corresponding to the first thumbnail is enlarged and displayed on the screen preview interface. Each storyboard script corresponds to at least one storyboard screen.

[0193] The first thumbnail may be any one of the multiple thumbnails. The triggering operation on any thumbnail may be a click operation on the thumbnail.

[0194] In some embodiments, the thumbnails of the multiple storyboards are displayed sequentially according to the order in which the multiple storyboards are displayed in the video. Furthermore, when the terminal zooms in on the first storyboard corresponding to the first thumbnail on the screen preview interface, it also displays the total length of the video and the time when the first storyboard was played in the video, thereby improving information transparency.

[0195] In some embodiments, the screen preview interface further displays a plurality of thumbnail size adjustment controls for adjusting the display area occupied by the triggered thumbnail in the area where the plurality of thumbnails are located, that is, for adjusting the size of the thumbnail.

[0196] In some embodiments, the terminal also displays first picture information of the first storyboard picture on the picture preview interface, and the first picture information includes at least one of picture description information, character information and subtitle information; in response to the modification operation of the first picture information of the first storyboard picture, the terminal displays the first storyboard picture regenerated based on the modified first picture information.

[0197] The screen description information includes information such as scene atmosphere, environmental elements, character actions, character dialogues, etc. Optionally, the terminal displays the screen description information in a screen editing box, and the screen editing box is used to modify the screen description information.

[0198] The character information includes the photo and name of the character in the first storyboard. Optionally, in response to a triggering operation on any character, an editing component of the triggered first character is displayed, and the editing component is used to modify at least one of the name, appearance, clothing, voice and personality characteristics of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, the modified first character is displayed; in response to a character addition operation in the editing interface where at least one first character is located, the added character is displayed in at least one first character, that is, the modification operation on the character on the screen preview interface is the same as the modification operation on the character on the application interface, which will not be repeated here.

[0199] The subtitle information includes the lines and narration information of each character in the first storyboard, etc. Optionally, the terminal displays the subtitle information in a subtitle editing box, and the subtitle editing box is used to modify the subtitle information.

[0200] In this embodiment, when a storyboard is magnified and displayed, the picture information of the storyboard is also displayed, and the storyboard is regenerated by modifying the picture information. In this way, the user can modify the storyboard through simple text prompts or operations without having to master complex image editing skills, which lowers the creation threshold and improves modification efficiency.

[0201] For example, see Figure 6 , Figure 6 6 is a schematic diagram illustrating a screen preview interface according to an exemplary embodiment. The screen preview interface displays multiple thumbnails 601. In response to a triggering operation on a thumbnail, the corresponding storyboard 602 is magnified and displayed on the screen preview interface. Furthermore, screen description information is displayed in a screen edit box 603. Character information 604 is displayed on the screen preview interface. Subtitle information is displayed in a subtitle edit box 605.

[0202] In some embodiments, a storyboard can be triggered to modify the storyboard. In response to a triggering operation on any area of ​​a first storyboard on the screen preview interface, the terminal displays at least one modification option for the triggered area, where the at least one modification option is used to modify the triggered area in at least one dimension. In response to a triggering operation on any of the at least one modification option, the terminal displays the first storyboard modified based on the triggered modification option.

[0203] The at least one modification option displayed varies depending on the triggered area. For example, if the triggered area is the character's area, at least one modification option is used to modify the character's appearance, clothing, etc. For another example, if the triggered area is the background area, at least one modification option is used to modify the background brightness, environment, etc. of the storyboard.

[0204] In some embodiments, the storyboard displayed enlarged on the screen preview interface is in an unmodifiable state by default. In response to a triggering operation on a modification control on the storyboard, the storyboard enters an editing state. In response to a triggering operation on any area on the storyboard in the editing state, corresponding modification options are displayed.

[0205] In this embodiment, any area on the storyboard can be modified by triggering the area, so that the user can make fine modifications to any area in the storyboard without affecting other parts. In this way, only part of the storyboard is redrawn or adjusted, and there is no need to regenerate the entire storyboard, which improves the modification efficiency.

[0206] In some embodiments, a storyboard can be replaced with a single click. In response to a replacement operation on the first storyboard on the preview interface, the terminal displays multiple candidate storyboards. In response to a trigger operation on any of the multiple candidate storyboards, the terminal replaces the first storyboard with the triggered candidate storyboard.

[0207] Among them, the multiple candidate storyboards can be storyboards generated by artificial intelligence based on a storyboard script, and the multiple candidate storyboards can also be images on the terminal locally, such as images in a local album.

[0208] The replacement operation for the first storyboard image can be a preset gesture operation such as a long press operation or a double-click operation on the first storyboard image. Alternatively, a replacement control for the first storyboard image is displayed on the screen preview interface, and the replacement operation for the first storyboard image is a triggering operation of the replacement control. For example, continue to refer to Figure 6 , a replacement control 606 for the first storyboard picture is displayed on the picture preview interface.

[0209] In this embodiment, the screen preview interface provides multiple candidate storyboards, allowing users to quickly compare different storyboards and select the one that best meets their needs, saving time and improving human-computer interaction efficiency. Furthermore, the one-click replacement function allows users to instantly see the modified effect, improving information transmission efficiency.

[0210] In some embodiments, after the terminal modifies any storyboard, it will also display a thumbnail of the storyboard before the modification on the screen preview interface, and triggering the thumbnail can also replace the storyboard with the storyboard before the modification corresponding to the thumbnail. It should be noted that if any storyboard is modified multiple times, the thumbnails of the multiple modified storyboards will be displayed on the screen preview interface, and can then be replaced with any modified version of the storyboard. For example, continue to see Figure 6 The picture preview interface also displays a thumbnail 607 of the storyboard picture before the first storyboard picture is modified.

[0211] In this disclosed embodiment, after the user confirms the storyboard, characters, and visual style, they can generate a video with one click. During the video generation process, a preview function for all storyboards is provided, allowing the user to modify the content at any time and then regenerate the video using artificial intelligence.

[0212] In the embodiment of the present disclosure, the process of displaying multiple storyboards generated based on multiple storyboard scripts is implemented through the above-mentioned step S303. In this embodiment, by displaying thumbnails, multiple storyboards can be displayed, thereby improving the transparency of information. By triggering any thumbnail, the corresponding storyboard can be enlarged and displayed. In this way, any storyboard can be viewed efficiently and conveniently, thereby improving the interaction efficiency.

[0213] In step S304, in response to the confirmation operation on the plurality of storyboards, the terminal displays a video generated based on the plurality of storyboards.

[0214] In some embodiments, the terminal displays a third confirmation control on the screen preview interface where the multiple storyboards are located. The confirmation operation on the multiple storyboards is also a triggering operation on the third confirmation control, and the third confirmation control is also used to generate a video. For example, continue to see Figure 6 , a third confirmation control 608 is also displayed on the screen preview interface.

[0215] In some embodiments, in response to the confirmation operation of the plurality of storyboards, the terminal displays a generation progress indicator in the video display area on the video generation interface, and the generation progress indicator is used to indicate the generation progress of the video. Figure 7 , Figure 7 FIG. 7 is a schematic diagram of a video generation interface according to an exemplary embodiment, wherein a generation progress indicator 702 is displayed on a video display area 701 on the video generation interface.

[0216] In some embodiments, the video generation interface also retains a display dialogue area. For example, continue to see Figure 7 ,The left area of ​​the video generation interface is the dialogue area.

[0217] In some embodiments, the video generation interface also displays the covers of videos generated by the user in history. These videos may include videos generated based on feature indication information and videos generated based on existing stories. Further, by triggering the cover of any video, the terminal can play the video. For example, continue to see Figure 7 The video generation interface also displays a cover 703 of the historically generated video.

[0218] In some embodiments, the video generation interface also displays a video publishing control, which is used to publish the video to the user's social account and can also be used to share the video with friends. Figure 7 , a video publishing control 704 is also displayed on the video generation interface.

[0219] In some embodiments, the terminal displays a video on a video preview interface, and the video preview interface also displays thumbnails of multiple storyboards; in response to a triggering operation on a second thumbnail among the multiple thumbnails, the terminal enlarges and displays a second storyboard corresponding to the second thumbnail on the video preview interface and displays second picture information of the second storyboard on the video preview interface, the second picture information including at least one of picture description information, character information, subtitle information, music information and camera movement mode; in response to a modification operation on the second picture information of the second storyboard, the terminal displays a second storyboard regenerated based on the modified second picture information, and the regenerated second storyboard is used to regenerate the video.

[0220] The second thumbnail may be any one of the plurality of thumbnails, and the triggering operation on any thumbnail may be a click operation on the thumbnail.

[0221] For example, see Figure 8 , Figure 8 : This is a schematic diagram of a video preview interface according to an exemplary embodiment. A certain storyboard 801 is displayed on the video preview interface, and a picture editing box 802 for the storyboard is also displayed on the video preview interface. The picture editing box 802 is used to modify the picture description information of the storyboard. The video preview interface also displays a regeneration control 803 for regenerating the storyboard with one click. The video preview interface also displays multiple attribute controls, which are used to modify the subtitles, soundtrack, and other attributes of the video, respectively, which are not specifically limited here. Furthermore, the video preview interface also displays a save control for saving the current modifications. After subsequent modifications to other attributes are completed, the storyboard is regenerated based on the saved modifications. Similarly, the video preview interface also displays a thumbnail of the storyboard before the second storyboard is modified, so that the second storyboard can be replaced with any previously modified version of the storyboard.

[0222] For example, see Figure 9 , Figure 9 1 is a schematic diagram illustrating a video preview interface according to an exemplary embodiment. Triggering the camera movement control on the video preview interface displays multiple preset camera movement modes 901. Selecting any of these modes allows you to modify the camera movement of the storyboard. Triggering the regenerate control regenerates the storyboard based on the selected camera movement mode with one click. Optionally, the selected camera movement mode is highlighted.

[0223] In this embodiment, after the video is generated, the user can also modify the video's storyboard and regenerate the video, which can significantly enhance the personalized features of the video, not only improving creation efficiency and reducing creation costs, but also increasing the controllability of video content, improving the flexibility and diversity of video creation. For example, by modifying the character's expression, action, language, etc., the character can be made more vivid, better conveying the emotions and the core of the story. By modifying the soundtrack, the theme and emotional direction of the video can be better matched. By modifying the camera movement, the user's attention can be attracted, and the tension and rhythm of the story can be enhanced. That is, by modifying the picture information, the quality of the video can be further improved.

[0224] In some embodiments, the terminal directly displays thumbnails on the video preview interface where the video is located. In other embodiments, the terminal displays thumbnails on the video preview interface in response to editing operations on the video, that is, thumbnails are displayed on the video preview interface only when the user has modification requirements for the video.

[0225] In some embodiments, the above-mentioned process in which the terminal displays a video generated based on multiple storyboards in response to a confirmation operation on multiple storyboards also includes the following implementation method: in response to a confirmation operation on multiple storyboards, the terminal displays multiple videos generated based on multiple storyboards, and the third picture information of at least one storyboard among the multiple videos is different, and the third picture information includes at least one of subtitle information, music information, and camera movement method.

[0226] Here, the third picture information includes at least one of subtitle information, music information and camera movement, etc., and is described as an example. The third picture information may also include characters, picture style, etc., which is not specifically limited here.

[0227] In the disclosed embodiment, the terminal can automatically generate a video based on the storyboard, and can generate a variety of videos based on the different picture information of the storyboard, which is beneficial for users to choose a satisfactory video from them, thereby improving user experience and video generation quality.

[0228] In some embodiments, before the terminal displays a video generated based on the multiple storyboards in response to a confirmation operation on the multiple storyboards, the terminal may further display the multiple storyboards after the order is adjusted in response to an operation to adjust the order of the multiple storyboards. Accordingly, in response to the confirmation operation on the multiple storyboards after the order is adjusted, the terminal displays the video generated based on the multiple storyboards after the order is adjusted.

[0229] Among them, multiple storyboards are displayed on the screen preview interface, and the multiple storyboards are arranged in sequence according to the display order in the video. Optionally, in response to the drag operation of at least one storyboard, the order of the multiple storyboards is adjusted, and the storyboards after the adjusted order are displayed.

[0230] In some embodiments, when a video is generated based on multiple storyboards, multiple storyboard scripts are also combined to generate the video. Optionally, the terminal automatically adjusts the multiple storyboard scripts based on the storyboards after the order is adjusted, and then generates a video based on the multiple storyboards after the order is adjusted and the multiple storyboard scripts.

[0231] In the disclosed embodiment, for a given plurality of storyboards, a user can adjust the order of the plurality of storyboards, and then generate a video based on the storyboards after the order has been adjusted. Since different storyboards can correspond to different storylines, the user can independently control the development order of each storyline in the video, so that the generated video better meets the user's needs and improves the quality of video generation.

[0232] It should be noted that the storyboard, characters, visual style, and storyboard images displayed by the terminal can be generated by the terminal itself. For example, the terminal may have a generative model embedded in it, such as an artificial intelligence model, and the terminal may use this generative model to generate the storyboard, characters, visual style, and storyboard images. Alternatively, the terminal may generate the storyboard, characters, visual style, and storyboard images through a server. Furthermore, the server may also generate the storyboards through a generative model, which is not specifically limited here.

[0233] In the disclosed embodiments, story content, storyboards, and storyboard images are generated in a conversational manner. Artificial intelligence leverages existing stories or extracts key information from video feature indicators to construct a story framework. This reduces reliance on innovative inspiration, which is difficult to obtain in traditional creative work, lowers the threshold for creation, provides users with more creative ideas and possibilities, and makes story creation easier. Furthermore, this method changes the traditional lengthy process from pre-production setup to post-production, and can quickly generate storyboards, storyboards, characters, and videos based on the input video description information, greatly improving creative efficiency and reducing creative time and cost. Furthermore, compared to the high requirements of traditional creation for large amounts of manpower and specialized equipment, this method utilizes artificial intelligence to automatically complete multiple tasks, reducing manpower and equipment investment, lowering the cost of creation, and making video creation more universal. Furthermore, it supports one-click video generation and allows for mid-process modification of video elements such as storyboards and storyboard images. After the video is generated, the camera movement and sound effects of the storyboard images can be edited again, allowing users to adjust video elements at any time according to their needs, ensuring that the final generated video is more in line with expectations, improving the accuracy and quality of the creation, giving users more independent control, and improving the quality of the generated video.

[0234] In the disclosed embodiments, a video can be generated using conversational interaction technology. Users use conversational instructions, descriptions, and other information, and AI accurately understands the semantics, extracts key information, and intelligently generates coherent and logically sound story content based on this information. A structured storyboard is then constructed, improving the quality of the generated content. Optionally, the terminal implements this process using advanced natural language processing algorithms and an intelligent story architecture model. Furthermore, based on the diverse information in the storyboard, such as scenes, character behaviors, and dialogue, AI uses image recognition and analysis, big data matching, and other technologies to accurately recommend appropriate visual styles. Furthermore, in conjunction with deep learning models, AI automatically generates unique, story-appropriate characters based on the story background and character characteristics. After AI completes the initial creative element preparation, the user triggers a one-click generation command, which rapidly integrates resources in the backend, calls the image rendering engine, and generates the video in real time, rendering it according to the predetermined visual style, characters, and storyboard, improving the efficiency of video generation.

[0235] The disclosed embodiments provide a video generation method, which can generate multiple storyboards by artificial intelligence for video description information input by a user, generate multiple storyboards after the user confirms the storyboards, and then generate a final video after the user confirms the storyboards. This method only requires the user to input video description information so that artificial intelligence can automatically generate the video, thereby improving the efficiency of video generation; and, in the middle of video generation, the generated storyboards, characters, picture styles, storyboards, etc. are displayed, and the video is generated after the user confirms, thereby improving the user's participation, that is, improving the user's controllability over the generated video. This not only improves the efficiency of video generation, but also enables the generated video to better meet user needs, thereby improving the quality of the generated video.

[0236] Figure 10 FIG. 1 is a block diagram of a video generating apparatus according to an exemplary embodiment. Figure 10 , the device comprises:

[0237] The first display unit 1001 is configured to, in response to a confirmation operation on the video description information input on the application interface, display a plurality of storyboards generated based on the video description information, wherein the plurality of storyboards correspond to at least one storyline;

[0238] The second display unit 1002 is configured to display a plurality of storyboards generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards, wherein each storyboard corresponds to at least one storyboard;

[0239] The third display unit 1003 is configured to display a video generated based on the multiple storyboards in response to a confirmation operation on the multiple storyboards.

[0240] In some embodiments, the first display unit 1001 is further configured to perform:

[0241] Displaying at least one first character and a first screen style, wherein the at least one first character and the first screen style are both determined based on a plurality of storyboards;

[0242] The second display unit 1002 is configured to execute:

[0243] In response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, a plurality of storyboard screens generated based on the plurality of storyboards, the at least one first character, and the first picture style are displayed.

[0244] In some embodiments, the apparatus further includes a first modifying unit configured to perform:

[0245] In response to a modification operation on at least one storyboard among the plurality of storyboards, displaying at least one second character and a second screen style, wherein the at least one second character and the second screen style are both determined based on the modified storyboard;

[0246] The second display unit 1002 is configured to execute:

[0247] In response to a confirmation operation on the modified storyboard, the at least one second character, and the second picture style, a plurality of storyboards generated based on the modified storyboard, the at least one second character, and the second picture style are displayed.

[0248] In some embodiments, the first modifying unit is configured to perform:

[0249] In response to a triggering operation on a replacement control on the application interface, a plurality of storyboards regenerated by artificial intelligence based on the video description information are displayed;

[0250] In response to a confirmation operation on the regenerated plurality of storyboards, at least one second character and a second screen style determined based on the regenerated plurality of storyboards are displayed.

[0251] In some embodiments, the apparatus further includes a second modifying unit configured to perform:

[0252] In response to a modification operation on at least one of the at least one first character and the first picture style, displaying the modified character and picture style;

[0253] The second display unit 1002 is configured to execute:

[0254] In response to a confirmation operation on the plurality of storyboards and the modified characters and screen styles, a plurality of storyboard screens generated based on the plurality of storyboards and the modified characters and screen styles are displayed.

[0255] In some embodiments, the second modifying unit is configured to perform at least one of the following:

[0256] In response to a triggering operation on any of the at least one first character, an editing component of the triggered first character is displayed, the editing component being used to modify at least one of the name, appearance, clothing, voice, and personality traits of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, the modified first character is displayed;

[0257] In response to a role adding operation on the interface where at least one first role is located, the added role is displayed in the at least one first role.

[0258] In some embodiments, the first picture style is highlighted in a picture style list, the picture style list includes multiple picture styles, and the picture style highlighted in the picture style list is a picture style determined based on multiple storyboards;

[0259] The second modification unit is configured to perform:

[0260] In response to a triggering operation on any picture style other than the first picture style in the picture style list, the triggered picture style is highlighted to obtain a modified picture style.

[0261] In some embodiments, the second display unit 1002 is configured to perform:

[0262] Thumbnails of respective multiple storyboards are displayed on a picture preview interface. In response to a triggering operation on a first thumbnail among the multiple thumbnails, a first storyboard corresponding to the first thumbnail is enlarged and displayed on the picture preview interface.

[0263] In some embodiments, the second display unit 1002 is further configured to perform:

[0264] Displaying first picture information of the first storyboard picture on the picture preview interface, the first picture information including at least one of picture description information, character information, and subtitle information;

[0265] In response to a modification operation on the first picture information of the first storyboard picture, a first storyboard picture regenerated based on the modified first picture information is displayed.

[0266] In some embodiments, the apparatus further includes a third modifying unit configured to execute:

[0267] In response to a triggering operation on any area on the first storyboard on the screen preview interface, at least one modification option for the triggered area is displayed, where the at least one modification option is used to modify the triggered area in at least one dimension;

[0268] In response to a triggering operation on any one of the at least one modification options, a first storyboard screen modified based on the triggered modification option is displayed.

[0269] In some embodiments, the apparatus further comprises a replacement unit configured to perform:

[0270] In response to a replacement operation on a first storyboard on the picture preview interface, a plurality of candidate storyboards are displayed;

[0271] In response to a triggering operation on any candidate storyboard among a plurality of candidate storyboards, the first storyboard is replaced with the triggered candidate storyboard.

[0272] In some embodiments, the third display unit 1003 is configured to perform:

[0273] The video is displayed on the video preview interface, and the video preview interface also displays thumbnails of multiple storyboards;

[0274] The apparatus further includes a fourth modifying unit configured to execute:

[0275] In response to a triggering operation on a second thumbnail among the plurality of thumbnails, a second storyboard corresponding to the second thumbnail is displayed in an enlarged manner on the video preview interface and second picture information of the second storyboard is displayed on the video preview interface, the second picture information including at least one of picture description information, character information, subtitle information, music information, and camera movement mode;

[0276] In response to a modification operation on the second picture information of the second storyboard picture, a second storyboard picture regenerated based on the modified second picture information is displayed, and the regenerated second storyboard picture is used to regenerate the video.

[0277] In some embodiments, the video description information includes characteristic indication information of the video to be generated, and the first display unit 1001 is configured to execute:

[0278] In response to a confirmation operation of the video description information input on the application interface, displaying a first story generated by artificial intelligence based on the video description information;

[0279] In response to a confirmation operation on the first story, multiple storyboards generated based on the first story are displayed; or, in response to a modification operation on the first story, multiple storyboards generated based on the modified first story are displayed.

[0280] In some embodiments, the video description information includes a second story, and the video description information is used to indicate that the video is generated based on the second story.

[0281] In some embodiments, the third display unit 1003 is configured to perform:

[0282] In response to the confirmation operation of multiple storyboards, multiple videos generated based on the multiple storyboards are displayed, and the third picture information of at least one storyboard is different between the multiple videos, and the third picture information includes at least one of subtitle information, music information and camera movement method.

[0283] In some embodiments, the apparatus further comprises an adjusting unit configured to perform:

[0284] In response to an operation of adjusting the order of the plurality of storyboards, displaying the plurality of storyboards after the order is adjusted;

[0285] The third display unit 1003 is configured to execute:

[0286] In response to a confirmation operation on the plurality of storyboards after the order is adjusted, a video generated based on the plurality of storyboards after the order is adjusted is displayed.

[0287] An embodiment of the present disclosure provides a video generation device, which can automatically generate multiple storyboards based on video description information input by a user, generate multiple storyboard screens after the user confirms the storyboards, and generate a final video after the user confirms the storyboard screens. The device can automatically generate video only after the user inputs the video description information, thereby improving the efficiency of video generation; and, in the middle of video generation, the generated storyboard scripts and storyboard screens are displayed, and the video is generated after the user confirms, thereby improving the user's participation, that is, improving the user's controllability over the generated video. This not only improves the efficiency of video generation, but also enables the generated video to better meet user needs, thereby improving the quality of the generated video.

[0288] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0289] In some embodiments, the electronic device is provided as a terminal. Figure 11The following is a block diagram of a terminal 1100 according to an exemplary embodiment of the present disclosure. Terminal 1100 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 1100 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0290] Typically, the terminal 1100 includes a processor 1101 and a memory 1102 .

[0291] Processor 1101 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 1101 may be implemented in hardware using at least one of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content required for display. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0292] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one program code, which is executed by the processor 1101 to implement the video generation method provided in the method embodiment of the present disclosure.

[0293] In some embodiments, terminal 1100 may optionally include a peripheral device interface 1103 and at least one peripheral device. Processor 1101, memory 1102, and peripheral device interface 1103 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1103 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.

[0294] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0295] RF circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. RF circuit 1104 optionally includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. RF circuit 1104 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, RF circuit 1104 may also include circuitry related to Near Field Communication (NFC), although this disclosure does not limit this.

[0296] Display screen 1105 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. If display screen 1105 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 1105. These touch signals can be input as control signals to processor 1101 for processing. Display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 1105, located on the front panel of terminal 1100. In other embodiments, there can be at least two display screens 1105, located on different surfaces of terminal 1100 or in a foldable design. In yet other embodiments, display screen 1105 can be a flexible display, located on a curved or foldable surface of terminal 1100. Display screen 1105 can also be configured as a non-rectangular, irregular shape, known as a special-shaped screen. The display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0297] The camera component 1106 is used to capture images or videos. Optionally, the camera component 1106 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera component 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0298] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1101 for processing, or input into the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0299] Power supply 1108 is used to power various components in terminal 1100. Power supply 1108 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1108 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0300] Those skilled in the art will understand that Figure 11 The structure shown in the figure does not constitute a limitation on the terminal 1100, and the terminal 1100 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0301] In an exemplary embodiment, a computer-readable storage medium is also provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above-described video generation method. Alternatively, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.

[0302] In an exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program. When the computer program is executed by a processor, the video generation method is implemented.

[0303] In some embodiments, the computer program product involved in the embodiments of the present disclosure may be deployed and executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected through a communication network. Multiple electronic devices distributed at multiple locations and interconnected through a communication network may constitute a blockchain system.

[0304] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0305] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video generation method, characterized in that: The method comprises: In response to a confirmation operation of video description information input on the application interface, a first story generated by artificial intelligence based on the video description information is displayed; in response to the confirmation operation of the first story, multiple storyboards generated by artificial intelligence based on the first story and at least one first character and a first visual style generated based on the multiple storyboards are displayed, wherein the multiple storyboards correspond to at least one storyline, and the video description information is used to indicate the story theme of the video to be generated; In response to a confirmation operation of the plurality of storyboards, the at least one first character, and the first picture style, displaying a plurality of storyboard frames generated by artificial intelligence based on the plurality of storyboards, the at least one first character, and the first picture style, wherein each storyboard corresponds to at least one storyboard frame; In response to a confirmation operation on the plurality of storyboards, displaying a video generated based on the plurality of storyboards; The method further comprises: displaying thumbnails of the plurality of storyboards on a screen preview interface; and in response to a triggering operation on a first thumbnail among the plurality of thumbnails, zooming in and displaying a first storyboard corresponding to the first thumbnail on the screen preview interface, and displaying first screen information of the first storyboard, wherein the first screen information includes at least one of screen description information, character information, and subtitle information, and the screen description information includes scene atmosphere, environmental elements, character actions, and character dialogues of the first storyboard; The method further includes: in response to a triggering operation on any area on the first storyboard on the screen preview interface, displaying at least one modification option for the triggered area, wherein the at least one modification option is used to modify the triggered area in at least one dimension; in response to a triggering operation on any modification option among the at least one modification option, displaying the first storyboard after the triggered area is modified based on the triggered modification option; wherein, if the triggered area is an area where any character is located, the at least one modification option is used to modify at least one of the character's name, appearance, clothing, timbre, and personality traits; In response to a modification operation on the first picture information of the first storyboard picture, displaying a first storyboard picture regenerated based on the modified first picture information; The method also includes: in response to a triggering operation on a replacement control on the application interface, displaying multiple storyboards regenerated by artificial intelligence, where the regenerated storyboards are used to regenerate characters, picture styles, and storyboard pictures.

2. The video generation method according to claim 1, wherein: Before displaying a plurality of storyboards generated by artificial intelligence based on the plurality of storyboards, the at least one first character, and the first picture style in response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, the method further includes: In response to a modification operation on at least one storyboard among the plurality of storyboards, displaying at least one second character and a second screen style, wherein both the at least one second character and the second screen style are determined based on the modified storyboard; The method further comprises: In response to a confirmation operation on the modified storyboard script, the at least one second character, and the second picture style, a plurality of storyboard pictures generated by artificial intelligence based on the modified storyboard script, the at least one second character, and the second picture style are displayed.

3. The video generation method according to claim 1, wherein: Before displaying a plurality of storyboards generated by artificial intelligence based on the plurality of storyboards, the at least one first character, and the first picture style in response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, the method further includes: In response to a modification operation on at least one of the at least one first character and the first picture style, displaying the modified character and picture style; The method of displaying, in response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, a plurality of storyboard images generated by artificial intelligence based on the plurality of storyboards, the at least one first character, and the first picture style, comprises: In response to a confirmation operation on the plurality of storyboards and the modified characters and picture styles, a plurality of storyboard pictures generated by artificial intelligence based on the plurality of storyboards and the modified characters and picture styles are displayed.

4. The video generation method according to claim 3, wherein: In response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style includes at least one of the following: In response to a triggering operation on any of the at least one first character, displaying an editing component of the triggered first character, the editing component being used to modify at least one of the name, appearance, clothing, voice, and personality traits of the triggered first character; In response to a modification operation on the triggered first character based on the editing component, displaying the modified first character; In response to a role adding operation on the interface where the at least one first role is located, the added role is displayed in the at least one first role.

5. The video generation method according to claim 3, characterized in that: The first picture style is highlighted in a picture style list, the picture style list includes a plurality of picture styles, and the picture style highlighted in the picture style list is a picture style determined based on the plurality of storyboards; In response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style includes: In response to a triggering operation on any picture style other than the first picture style in the picture style list, the triggered picture style is highlighted to obtain a modified picture style.

6. The video generation method according to claim 1, characterized in that The displaying of a video generated based on the plurality of storyboards includes: Displaying the video on a video preview interface, wherein the video preview interface also displays thumbnails of the plurality of storyboards; The method further comprises: In response to a triggering operation on a second thumbnail among the plurality of thumbnails, a second storyboard corresponding to the second thumbnail is displayed in an enlarged manner on the video preview interface, and second picture information of the second storyboard is displayed on the video preview interface, where the second picture information includes at least one of picture description information, character information, subtitle information, music information, and camera movement mode; In response to a modification operation on the second picture information of the second storyboard picture, a second storyboard picture regenerated based on the modified second picture information is displayed, and the regenerated second storyboard picture is used to regenerate the video.

7. The video generation method according to claim 1, characterized in that: The method further comprises: In response to the modification operation on the first story, a plurality of storyboards generated based on the modified first story are displayed.

8. The video generation method according to claim 1, wherein: The video description information includes a second story, and the video description information is used to indicate that a video is generated based on the second story.

9. The video generation method according to claim 1, wherein: The step of displaying a video generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards comprises: In response to a confirmation operation on the multiple storyboards, multiple videos generated based on the multiple storyboards are displayed, wherein the third screen information of at least one storyboard is different between the multiple videos, and the third screen information includes at least one of subtitle information, music information, and camera movement method.

10. The video generation method according to claim 1, characterized in that: Before displaying the video generated based on the multiple storyboards in response to the confirmation operation on the multiple storyboards, the method further includes: In response to an operation of adjusting the order of the plurality of storyboard pictures, displaying the plurality of storyboard pictures after the order is adjusted; The step of displaying a video generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards comprises: In response to a confirmation operation on the plurality of storyboards after the order is adjusted, a video generated based on the plurality of storyboards after the order is adjusted is displayed.

11. A video generating device, characterized in that: The device comprises: The first display unit is configured to, in response to a confirmation operation of video description information input on an application interface, display a first story generated by artificial intelligence based on the video description information; in response to the confirmation operation of the first story, display a plurality of storyboards generated by artificial intelligence based on the first story and at least one first character and a first screen style generated based on the plurality of storyboards, wherein the plurality of storyboards correspond to at least one storyline, and the video description information is used to indicate a story theme of the video to be generated; a second display unit configured to, in response to a confirmation operation on the plurality of storyboards, the at least one first character, and the first picture style, display a plurality of storyboard images generated by artificial intelligence based on the plurality of storyboards, the at least one first character, and the first picture style, wherein each storyboard corresponds to at least one storyboard image; a third display unit configured to display a video generated based on the plurality of storyboards in response to a confirmation operation on the plurality of storyboards; The second display unit is further configured to: display thumbnails of each of the plurality of storyboards on a picture preview interface; and in response to a triggering operation on a first thumbnail among the plurality of thumbnails, enlarge and display a first storyboard corresponding to the first thumbnail on the picture preview interface, and display first picture information of the first storyboard, wherein the first picture information includes at least one of picture description information, character information, and subtitle information, and the picture description information includes scene atmosphere, environmental elements, character actions, and character dialogues of the first storyboard; In response to a triggering operation on any area on the first storyboard on the picture preview interface, at least one modification option for the triggered area is displayed, wherein the at least one modification option is used to modify the triggered area in at least one dimension; in response to a triggering operation on any modification option among the at least one modification option, the first storyboard after the triggered area is modified based on the triggered modification option is displayed; wherein, if the triggered area is the area where any character is located, the at least one modification option is used to modify at least one of the character's name, appearance, clothing, timbre, and personality traits; in response to a modification operation on first picture information of the first storyboard, a first storyboard regenerated based on the modified first picture information is displayed; The first display unit is also configured to execute a trigger operation in response to a replacement control on the application interface, and display multiple storyboards regenerated by artificial intelligence, where the regenerated storyboards are used to regenerate characters, picture styles and storyboard pictures.

12. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the video generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generating method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the video generation method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Video generation method and device, readable medium, electronic equipment and program product

    CN119364091A

  • Method and device for generating video based on text

    CN119893165A