Video generation method and device, equipment and storage medium

Through the video generation method that automatically generates and confirms users, the problem of low video generation efficiency in the prior art is solved, an efficient and controllable video generation process is realized, and the quality of generated videos is improved.

CN120281994AActive Publication Date: 2025-07-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510704482.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-08
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The prior art requires professional steps such as script writing, shooting and editing when generating videos, resulting in low generation efficiency and long-term consumption.

Method used

A video generation method is provided. By automatically generating a storyboard script and a storyboard screen in response to the video description information input by the user, and generate a video after the user confirms. Using artificial intelligence to automatically generate multiple storyboard scripts and storyboard screens, the user can modify and confirm during the generation process.

Benefits of technology

It improves the efficiency and quality of video generation, enhances users' controllability of generated videos, and makes the generated video more in line with user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281994A_ABST
    Figure CN120281994A_ABST
Patent Text Reader

Abstract

The invention relates to a video generation method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. According to the method, for video description information input by a user, a plurality of split scripts can be generated, a plurality of split pictures are generated after the user confirms the split scripts, and a final video is generated after the user confirms the split pictures, and the video can be automatically generated only by inputting the video description information by the user, so that the video generation efficiency is improved; moreover, in the process of generating the video, the generated split script and the split picture are displayed, and the video is generated after the user confirms, so that the participation degree of the user is improved, namely, the controllability of the user on the generated video is improved, the video generation efficiency is improved, the generated video can better meet the requirements of the user, and the user experience is improved. And the quality of the generated video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a video generation method, apparatus, device, and storage medium. Background Art

[0002] In the related art, steps such as script writing, shooting, and editing are required when generating a video. This not only requires a professional team and equipment, but also takes a long time in the generation process, resulting in low efficiency of video generation. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, device, and storage medium, and this method improves the efficiency and quality of video generation. The technical solution of the present disclosure is as follows.

[0004] According to one aspect of the embodiments of the present disclosure, a video generation method is provided. The method includes: In response to a confirmation operation on video description information input on an application interface, displaying a plurality of storyboard scripts generated based on the video description information, where the plurality of storyboard scripts correspond to at least one plot; In response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, where each storyboard script corresponds to at least one storyboard frame; In response to a confirmation operation on the plurality of storyboard frames, displaying a video generated based on the plurality of storyboard frames.

[0005] In some embodiments, before the step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the method further includes: Displaying at least one first character and a first picture style, where both the at least one first character and the first picture style are determined based on the plurality of storyboard scripts; The step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts includes: In response to a confirmation operation on the plurality of storyboard scripts, the at least one first character, and the first picture style, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the at least one first character, and the first picture style.

[0006] In some embodiments, before the step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the method further includes: In response to a modification operation on at least one of the multiple storyboard scripts, display at least one second character and a second screen style, where the at least one second character and the second screen style are both determined based on the modified storyboard script; In response to a confirmation operation on the multiple storyboard scripts, display multiple storyboard frames generated based on the multiple storyboard scripts, including: In response to a confirmation operation on the modified storyboard script, the confirmation operation on the at least one second character and the second screen style, display multiple storyboard frames generated based on the modified storyboard script, the at least one second character, and the second screen style.

[0007] In some embodiments, the step of, in response to a modification operation on at least one of the multiple storyboard scripts, displaying at least one second character and a second screen style includes: In response to a trigger operation on a replacement control on the application interface, display multiple storyboard scripts regenerated by artificial intelligence based on the video description information; In response to a confirmation operation on the regenerated multiple storyboard scripts, display the at least one second character and the second screen style determined based on the regenerated multiple storyboard scripts.

[0008] In some embodiments, before the step of, in response to a confirmation operation on the multiple storyboard scripts, the at least one first character, and the first screen style, displaying multiple storyboard frames generated based on the multiple storyboard scripts, the at least one first character, and the first screen style, the method further includes: In response to a modification operation on at least one of the at least one first character and the first screen style, display the modified character and screen style; The step of, in response to a confirmation operation on the multiple storyboard scripts, the at least one first character, and the first screen style, displaying multiple storyboard frames generated based on the multiple storyboard scripts, the at least one first character, and the first screen style includes: In response to a confirmation operation on the multiple storyboard scripts and the modified character and screen style, display multiple storyboard frames generated based on the multiple storyboard scripts and the modified character and screen style.

[0009] In some embodiments, the step of, in response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style includes at least one of the following: In response to a triggering operation on any one of the at least one first character, display an editing component for the triggered first character, where the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality traits of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, display the modified first character. In response to a character addition operation on the interface where the at least one first character is located, display the added character among the at least one first character.

[0010] In some embodiments, the first screen style is prominently displayed in a screen style list, the screen style list includes multiple screen styles, and the prominently displayed screen style in the screen style list is the screen style determined based on the multiple storyboard scripts. The displaying of the modified character and screen style in response to a modification operation on at least one of the at least one first character and the first screen style includes: In response to a triggering operation on any screen style other than the first screen style in the screen style list, prominently display the triggered screen style to obtain the modified screen style.

[0011] In some embodiments, the displaying of the multiple storyboard images generated based on the multiple storyboard scripts includes: Display thumbnails of the multiple storyboard images on a screen preview interface, and in response to a triggering operation on a first thumbnail among the multiple thumbnails, magnify and display the first storyboard image corresponding to the first thumbnail on the screen preview interface.

[0012] In some embodiments, the method further includes: Display first screen information of the first storyboard image on the screen preview interface, where the first screen information includes at least one of screen description information, character information, and subtitle information. In response to a modification operation on the first screen information of the first storyboard image, display the first storyboard image regenerated based on the modified first screen information.

[0013] In some embodiments, the method further includes: In response to a triggering operation on any area of the first storyboard image on the screen preview interface, display at least one modification option for the triggered area, where the at least one modification option is used to modify the triggered area in at least one dimension. In response to a triggering operation on any one of the at least one modification option, display the first storyboard image modified based on the triggered modification option.

[0014] In some embodiments, the method further includes: In response to a replacement operation on the first storyboard frame in the screen preview interface, display a plurality of candidate storyboard frames; In response to a triggering operation on any one of the plurality of candidate storyboard frames, replace the first storyboard frame with the triggered candidate storyboard frame.

[0015] In some embodiments, the displaying the video generated based on the plurality of storyboard frames includes: Display the video on a video preview interface, and thumbnails of the plurality of storyboard frames are also displayed on the video preview interface; The method further includes: In response to a triggering operation on a second thumbnail among the plurality of thumbnails, magnify and display the second storyboard frame corresponding to the second thumbnail on the video preview interface and display second frame information of the second storyboard frame on the video preview interface, where the second frame information includes at least one of frame description information, character information, subtitle information, background music information, and camera movement method; In response to a modification operation on the second frame information of the second storyboard frame, display a second storyboard frame regenerated based on the modified second frame information, and the regenerated second storyboard frame is used to regenerate the video.

[0016] In some embodiments, the video description information includes feature indication information of the video to be generated. In response to a confirmation operation on the video description information input on the application interface, display a plurality of storyboard scripts generated based on the video description information, including: In response to a confirmation operation on the video description information input on the application interface, display a first story generated by artificial intelligence based on the video description information; In response to a confirmation operation on the first story, display a plurality of storyboard scripts generated based on the first story; or, in response to a modification operation on the first story, display a plurality of storyboard scripts generated based on the modified first story.

[0017] In some embodiments, the video description information includes a second story, and the video description information is used to indicate generating a video based on the second story.

[0018] In some embodiments, the displaying the video generated based on the plurality of storyboard frames in response to a confirmation operation on the plurality of storyboard frames includes: In response to a confirmation operation on the plurality of storyboard frames, display a plurality of videos generated based on the plurality of storyboard frames, where third frame information of at least one storyboard frame is different among the plurality of videos, and the third frame information includes at least one of subtitle information, background music information, and camera movement method.

[0019] In some embodiments, before displaying the video generated based on the multiple storyboard frames in response to a confirmation operation on the multiple storyboard frames, the method further includes: In response to an operation for adjusting the order of the multiple storyboard frames, displaying the multiple storyboard frames with the adjusted order; The displaying the video generated based on the multiple storyboard frames in response to a confirmation operation on the multiple storyboard frames includes: In response to a confirmation operation on the multiple storyboard frames with the adjusted order, displaying the video generated based on the multiple storyboard frames with the adjusted order.

[0020] According to another aspect of the embodiments of the present disclosure, there is provided a video generation device, the device includes: A first display unit, configured to execute and display multiple storyboard scripts generated based on the video description information in response to a confirmation operation on the video description information input on the application interface, the multiple storyboard scripts corresponding to at least one storyline; A second display unit, configured to execute and display multiple storyboard frames generated based on the multiple storyboard scripts in response to a confirmation operation on the multiple storyboard scripts, each storyboard script corresponding to at least one storyboard frame; A third display unit, configured to execute and display the video generated based on the multiple storyboard frames in response to a confirmation operation on the multiple storyboard frames.

[0021] In some embodiments, the first display unit is further configured to execute: Displaying at least one first character and a first picture style, both the at least one first character and the first picture style being determined based on the multiple storyboard scripts; The second display unit is configured to execute: In response to a confirmation operation on the multiple storyboard scripts, the at least one first character, and the first picture style, displaying multiple storyboard frames generated based on the multiple storyboard scripts, the at least one first character, and the first picture style.

[0022] In some embodiments, the device further includes a first modification unit, configured to execute: In response to a modification operation on at least one of the multiple storyboard scripts, displaying at least one second character and a second picture style, both the at least one second character and the second picture style being determined based on the modified storyboard script; The second display unit is configured to execute: In response to the confirmation operation for the modified storyboard script, the confirmation operation for the at least one second character and the second screen style, display a plurality of storyboard frames generated based on the modified storyboard script, the at least one second character, and the second screen style.

[0023] In some embodiments, the first modification unit is configured to perform: In response to a trigger operation on the replacement control on the application interface, display a plurality of storyboard scripts regenerated by artificial intelligence based on the video description information; In response to the confirmation operation for the regenerated plurality of storyboard scripts, display the at least one second character and the second screen style determined based on the regenerated plurality of storyboard scripts.

[0024] In some embodiments, the apparatus further includes a second modification unit configured to perform: In response to a modification operation on at least one of the at least one first character and the first screen style, display the modified character and screen style; The second display unit is configured to perform: In response to the confirmation operation for the plurality of storyboard scripts and the modified character and screen style, display a plurality of storyboard frames generated based on the plurality of storyboard scripts and the modified character and screen style.

[0025] In some embodiments, the second modification unit is configured to perform at least one of the following: In response to a trigger operation on any one of the at least one first character, display an editing component for the triggered first character, the editing component being used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to the modification operation on the triggered first character based on the editing component, display the modified first character; In response to a character addition operation on the interface where the at least one first character is located, display the added character among the at least one first character.

[0026] In some embodiments, the first screen style is prominently displayed in a screen style list, the screen style list includes a variety of screen styles, and the prominently displayed screen style in the screen style list is the screen style determined based on the plurality of storyboard scripts; The second modification unit is configured to perform: In response to a trigger operation on any screen style other than the first screen style in the screen style list, prominently display the triggered screen style to obtain the modified screen style.

[0027] In some embodiments, the second display unit is configured to perform: Display thumbnails of the respective multiple storyboard images on the screen preview interface, and in response to a trigger operation on a first thumbnail among the multiple thumbnails, magnify and display the first storyboard image corresponding to the first thumbnail on the screen preview interface.

[0028] In some embodiments, the second display unit is further configured to perform: Display first screen information of the first storyboard image on the screen preview interface, where the first screen information includes at least one of screen description information, character information, and subtitle information; In response to a modification operation on the first screen information of the first storyboard image, display the first storyboard image regenerated based on the modified first screen information.

[0029] In some embodiments, the apparatus further includes a third modification unit configured to perform: In response to a trigger operation on any region of the first storyboard image on the screen preview interface, display at least one modification option for the triggered region, where the at least one modification option is used to modify the triggered region in at least one dimension; In response to a trigger operation on any one of the at least one modification options, display the first storyboard image modified based on the triggered modification option.

[0030] In some embodiments, the apparatus further includes a replacement unit configured to perform: In response to a replacement operation on the first storyboard image on the screen preview interface, display multiple candidate storyboard images; In response to a trigger operation on any one of the multiple candidate storyboard images, replace the first storyboard image with the triggered candidate storyboard image.

[0031] In some embodiments, the third display unit is configured to perform: Display the video on the video preview interface, and thumbnails of the respective multiple storyboard images are also displayed on the video preview interface; The apparatus further includes a fourth modification unit configured to perform: In response to a trigger operation on a second thumbnail among the multiple thumbnails, magnify and display the second storyboard image corresponding to the second thumbnail on the video preview interface and display second screen information of the second storyboard image on the video preview interface, where the second screen information includes at least one of screen description information, character information, subtitle information, background music information, and camera movement mode; In response to a modification operation on the second screen information of the second storyboard screen, display the second storyboard screen regenerated based on the modified second screen information, and the regenerated second storyboard screen is used to regenerate a video.

[0032] In some embodiments, the video description information includes feature indication information of a video to be generated, and the first display unit is configured to perform: In response to a confirmation operation on the video description information input on the application interface, display a first story generated by artificial intelligence based on the video description information; In response to a confirmation operation on the first story, display a plurality of storyboard scripts generated based on the first story; or, in response to a modification operation on the first story, display a plurality of storyboard scripts generated based on the modified first story.

[0033] In some embodiments, the video description information includes a second story, and the video description information is used to indicate that a video is generated based on the second story.

[0034] In some embodiments, the third display unit is configured to perform: In response to a confirmation operation on the plurality of storyboard screens, display a plurality of videos generated based on the plurality of storyboard screens, where at least one of the third screen information of the plurality of storyboard screens is different, and the third screen information includes at least one of subtitle information, background music information, and camera movement mode.

[0035] In some embodiments, the device further includes an adjustment unit configured to perform: In response to an operation of adjusting the order of the plurality of storyboard screens, display the plurality of storyboard screens with the adjusted order; The third display unit is configured to perform: In response to a confirmation operation on the plurality of storyboard screens with the adjusted order, display a video generated based on the plurality of storyboard screens with the adjusted order.

[0036] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the above video generation method.

[0037] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above video generation method.

[0038] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes a computer program. When the computer program is executed by a processor, the above video generation method is implemented.

[0039] The embodiments of the present disclosure provide a video generation method. For the video description information input by a user, the method can generate multiple storyboard scripts. After the user confirms the storyboard scripts, multiple storyboard images are generated. After the user confirms the storyboard images, the final video is generated. The method can be automatically generated only by the user inputting the video description information, which improves the video generation efficiency. Moreover, in the middle of video generation, the generated storyboard scripts and storyboard images are displayed, and the video is generated after the user confirms, which improves the user's participation, that is, improves the controllability of the video generated by the user. In this way, not only the video generation efficiency is improved, but also the generated video can better meet the user's needs, thereby improving the quality of the generated video.

[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0041] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0042] Figure 1 It is a schematic diagram of an implementation environment shown according to an exemplary embodiment.

[0043] Figure 2 It is a flowchart of a video generation method shown according to an exemplary embodiment.

[0044] Figure 3 It is a flowchart of another video generation method shown according to an exemplary embodiment.

[0045] Figure 4 It is a schematic diagram of an application interface shown according to an exemplary embodiment.

[0046] Figure 5 It is a schematic diagram of another application interface shown according to an exemplary embodiment.

[0047] Figure 6 It is a schematic diagram of a screen preview result shown according to an exemplary embodiment.

[0048] Figure 7 It is a schematic diagram of a video generation interface shown according to an exemplary embodiment.

[0049] Figure 8It is a schematic diagram of a video preview interface shown according to an exemplary embodiment.

[0050] Figure 9 It is a schematic diagram of another video preview interface shown according to an exemplary embodiment.

[0051] Figure 10 It is a block diagram of a video generation device shown according to an exemplary embodiment.

[0052] Figure 11 It is a block diagram of a terminal shown according to an exemplary embodiment. Detailed implementation manners

[0053] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0054] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0055] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the video description information involved in the present disclosure is obtained under full authorization.

[0056] The video generation method provided by the embodiments of the present disclosure can be executed by an electronic device, and the electronic device can be provided as at least one of a terminal and a server. Figure 1 It is a schematic diagram of an implementation environment provided by the embodiments of the present disclosure. Refer to Figure 1 , this implementation environment includes: terminal 101 and server 102.

[0057] In the embodiments of the present disclosure, a target application is installed on the terminal 101, and the target application is used to automatically generate videos. The server 102 is the background server of the target application and is used to provide background services for the target application.

[0058] In some embodiments, after the video description information is input on the application interface of the terminal 101 and the user performs a confirmation operation on the video description information, a plurality of storyboard scripts corresponding to at least one storyline are automatically generated based on the video description information, and the terminal 101 displays the plurality of storyboard scripts; after the user performs a confirmation operation on the plurality of storyboard scripts, a plurality of storyboard images generated based on the plurality of storyboard scripts are displayed, and one storyboard script may correspond to one or more storyboard images; finally, after the user performs a confirmation operation on the plurality of storyboard images, the terminal 101 displays the video generated based on the plurality of storyboard images.

[0059] The terminal 101 may be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop portable computer. The terminal 101 has a communication function and can access a wired network or a wireless network. The terminal 101 may generally refer to one of a plurality of terminals, and those skilled in the art may know that the number of the above terminals may be more or less. The server 102 may be an independent physical server, or a server cluster or a distributed file system composed of a plurality of physical servers, or may be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 and the terminal 101 are directly or indirectly connected by a wired or wireless communication method, and the embodiments of the present disclosure do not limit this. Optionally, the number of the above servers 102 may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may also include other functional servers to provide more comprehensive and diversified services. Among them, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 or the terminal 101 can separately undertake the computing work, and the embodiments of the present disclosure do not limit this.

[0060] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 2 shown, this method is executed by the terminal, and this method includes the following steps.

[0061] In step S201, in response to a confirmation operation on the video description information input on the application interface, the terminal displays a plurality of storyboard scripts generated based on the video description information, and the plurality of storyboard scripts correspond to at least one storyline.

[0062] In an embodiment of the present disclosure, the application interface is an interface within a target application. The target application is used to generate a video based on video description information input by a user. Further, if the target application generates the video through artificial intelligence, then the artificial intelligence generates multiple storyboard scripts based on the video description information.

[0063] In some embodiments, an information input box is displayed on the application interface. The information input box is used to input video description information, and a first confirmation control is displayed on the application interface. After the user triggers the first confirmation control to confirm the video description information, the terminal generates multiple storyboard scripts based on the video description information and displays these multiple storyboard scripts.

[0064] Optionally, if the video is generated by artificial intelligence, the application interface is a dialogue interface between the user and an artificial intelligence character. A dialogue input box is displayed on the dialogue interface. The dialogue input box is used to input video description information, and a send control is displayed on the dialogue interface. The confirmation operation of the video description information is also the triggering operation of the send control. Correspondingly, the multiple storyboard scripts are the dialogue information replied by the artificial intelligence character on the dialogue interface.

[0065] In an embodiment of the present disclosure, the video description information is used to indicate the characteristics of the video to be generated, that is, it is used to indicate what kind of video to generate. A storyboard script is a visual script that presents a story or scene through short text descriptions. The storyboard script includes one or more of the description of the shot content, the lines of the characters in the shot, the actions of the characters, the size and range of the main body in the shot, the camera movement method, the sound effect, the duration of the picture, the visual effect, the remarks, etc.

[0066] In an embodiment of the present disclosure, the multiple storyboard scripts correspond to at least one storyline. These storylines are obtained by the artificial intelligence based on the video description information, and then the artificial intelligence generates multiple storyboard scripts around these storylines.

[0067] Among them, the multiple storyboard scripts corresponding to the same storyline are arranged in chronological order, and the multiple storyboard scripts show the development process of the story in sequence, that is, according to the timeline, the cause, process, and result of the story are shown through these multiple storyboard scripts. Further, there is at least one same element among the multiple storyboard scripts corresponding to the same storyline, such as the same characters, scenes, or background music in the multiple storyboard scripts.

[0068] In some embodiments, multiple storyboard scripts generated by artificial intelligence may correspond to multiple storylines, and one storyline corresponds to at least one storyboard script. These multiple storylines can be associated in various ways such as theme, causality, characters, timeline, location, etc., ensuring the coherence between multiple storylines. For example, multiple storylines unfold around one theme; or, the result of one storyline is the cause of another storyline; or, multiple storylines revolve around the same characters; or, multiple storylines can occur at different time points, but they form a complete timeline; or, multiple storylines take place at the same location.

[0069] In step S202, in response to the confirmation operation on the multiple storyboard scripts, the terminal displays multiple storyboard frames generated based on the multiple storyboard scripts, and each storyboard script corresponds to at least one storyboard frame.

[0070] In some embodiments, a second confirmation control is displayed on the interface where the multiple storyboard scripts are located, and the confirmation operation on the multiple storyboard scripts is also the triggering operation on this second confirmation control.

[0071] Optionally, the multiple storyboard frames are generated by artificial intelligence based on the multiple storyboard scripts, and the artificial intelligence generates at least one storyboard frame for each storyboard script.

[0072] In step S203, in response to the confirmation operation on the multiple storyboard frames, the terminal displays a video generated based on the multiple storyboard frames.

[0073] In some embodiments, a third confirmation control is displayed on the interface where the multiple storyboard frames are located, and the confirmation operation on the multiple storyboard frames is also the triggering operation on this third confirmation control.

[0074] In the embodiments of the present disclosure, the terminal displaying a video may refer to displaying the cover of the video, and triggering the cover can play the video. The terminal displaying a video may also refer to playing the video. The generated video can be an animated video, a comic video, or a cartoon video.

[0075] Optionally, the terminal generates a video based on multiple storyboard images by artificial intelligence. In some embodiments, generating a video based on multiple storyboard images may be directly splicing the multiple storyboard images to obtain a video. In other embodiments, intermediate storyboard images are generated by interpolating between every two adjacent storyboard images to connect the two adjacent storyboard images through the interpolated storyboard images, further enriching the video content. Alternatively, for at least some of the multiple storyboard images, a dynamic graphic is generated based on each of the at least some storyboard images, and each dynamic graphic is a video clip. Then, a video is obtained based on the remaining storyboard images among the multiple storyboard images and the video clips corresponding to the at least some storyboard images. For example, these storyboard images and video clips are spliced in chronological order to obtain a video.

[0076] Embodiments of the present disclosure provide a video generation method. For the video description information input by a user, the method can generate multiple storyboard scripts. After the user confirms the storyboard scripts, multiple storyboard images are generated. After the user confirms the storyboard images, the final video is generated. The method can be automatically generated only by the user inputting the video description information, improving the video generation efficiency. Moreover, in the middle of generating the video, the generated storyboard scripts and storyboard images are displayed, and the video is generated after the user confirms, improving the user participation, that is, improving the controllability of the user over the generated video. In this way, not only the video generation efficiency is improved, but also the generated video can better meet the user requirements, thereby improving the quality of the generated video.

[0077] The above Figure 2 only shows the basic process of the present disclosure. Based on a specific implementation manner below, the solution provided by the present disclosure will be further elaborated. Refer to Figure 3 , Figure 3 is a flowchart of another video generation method shown according to an exemplary embodiment. The method is executed by a terminal and includes at least one of the following steps.

[0078] In step S301, in response to a confirmation operation on the video description information input on the application interface, the terminal displays multiple storyboard scripts generated based on the video description information and displays at least one first character and a first picture style. The multiple storyboard scripts correspond to at least one storyline, and both the at least one first character and the first picture style are determined based on the multiple storyboard scripts.

[0079] In some embodiments, the video description information includes feature indication information of the video to be generated, that is, the video description information is used to indicate what kind of video to generate. The feature indication information may indicate one or more of the story theme, plot, character features, picture features, etc. of the video to be generated. When generating a video based on the video description information including the feature indication information, a story can be generated based on the video description information first. Correspondingly, the process of the terminal displaying multiple storyboard scripts generated based on the video description information in response to the confirmation operation of the video description information input on the application interface includes the following steps: in response to the confirmation operation of the video description information input on the application interface, the terminal displays a first story generated by artificial intelligence based on the video description information; in response to the confirmation operation of the first story, the terminal displays multiple storyboard scripts generated based on the first story; or, in response to the modification operation of the first story, the terminal displays multiple storyboard scripts generated based on the modified first story.

[0080] In this embodiment, when the video description information is the feature indication information indicating the video to be generated, a story is generated based on the video description information first and displayed. When the user confirms the generated story, the corresponding storyboard scripts are generated. This makes the generated storyboard scripts meet the user's creative needs and have higher quality. Moreover, the generated story can be modified, and the storyboard scripts can be generated based on the modified story. This allows the user to optimize the story content, providing the user with greater flexibility and controllability. While improving the creative efficiency, it also improves the overall quality of the creation.

[0081] In some other embodiments, the video description information includes a second story, and the video description information is used to indicate generating a video based on the second story, that is, generating a video based on an existing story.

[0082] Optionally, the artificial intelligence generates a video based on the second story. If the video description information includes the second story, the artificial intelligence generates multiple storyboard scripts based on this second story, and the multiple storyboard scripts correspond to multiple storylines in the second story. Further, the video description information not only includes the second story, but also includes story indication information, which is used to indicate the features of the video generated based on this second story. For example, the story indication information can be used to indicate information such as the picture features and sound effects of the video to be generated. Then the artificial intelligence generates multiple storyboard scripts based on the second story and the story indication information.

[0083] In the embodiments of the present disclosure, the terminal can automatically generate storyboard scripts and storyboard pictures based on a given story, and then automatically generate a video, without relevant workers having to perform steps such as script writing, shooting, and editing based on the story, shortening the time-consuming during video generation and improving the efficiency during video generation.

[0084] In some embodiments, the application interface includes two generation options. The first generation option is to generate a video based on video description information including feature indication information, and the second generation option is to generate a video based on video description information including a story.

[0085] For example, referring to Figure 4 , Figure 4 FIG. is a schematic diagram of an application interface shown according to an exemplary embodiment. Among them, the application interface can be a dialogue interface. The application interface includes a dialogue area, and the dialogue area includes a first generation option 401 and a second generation option 402. Optionally, after inputting video description information in the dialogue input box 403, in response to a trigger operation on the first generation option 401, a first story generated based on the video description information is displayed, and in response to a trigger operation on the second generation option 402, a plurality of storyboard scripts generated based on the video description information are displayed.

[0086] In some embodiments, the application interface further includes a result display area for displaying information such as the generated storyboard scripts, characters, and picture styles. For example, continuing to refer to Figure 4 , the left area on the application interface is the dialogue area, and the right area on the application interface is the result display area.

[0087] In some embodiments, the generated plurality of storyboard scripts can be displayed not only in the form of dialogue messages in the dialogue area but also in the script editing box in the result display area. For example, referring to Figure 5 , Figure 5 FIG. is a schematic diagram of an application interface shown according to an exemplary embodiment. Among them, a plurality of storyboard scripts are displayed in the form of dialogue messages in the dialogue area on the application interface, and a script editing box 501 is displayed in the result display area on the application interface, and a plurality of storyboard scripts are displayed in the script editing box 501.

[0088] In the embodiments of the present disclosure, the characters determined based on the storyboard scripts are virtual characters, such as virtual digital humans. Optionally, at least one character is instantaneously generated by artificial intelligence based on a plurality of storyboard scripts; or, the terminal filters out at least one character from a plurality of preset characters based on a plurality of storyboard scripts. Among them, artificial intelligence determines the characters based on the descriptions of characters, scenes, etc. in the storyboard scripts.

[0089] Among them, the terminal displaying each character means displaying the name of each character and the corresponding photo, and the photo includes information such as the appearance and clothing of the character, which is used to show the appearance image of the character; further, information such as the voice and personality characteristics of the character can also be displayed. For example, continuing to refer to Figure 5 , a plurality of characters 502 are displayed in the result display area on the application interface.

[0090] In the embodiments of the present disclosure, the screen style is an indispensable part of the storyboard screen. Through various elements such as color, composition, lighting, and dynamics, it determines the visual presentation of the video, endows the story with a unique visual language, and enhances emotional expression and the audience's sense of immersion. The screen style can be defined from multiple perspectives. For example, when defined from the color perspective, it can include styles such as cool tones, warm tones, complementary colors, and contrast colors; when defined from the composition perspective, it can include styles such as realistic, dreamy, high contrast, and low contrast; when defined from the animation perspective, it can include styles such as cartoon, realistic, watercolor, and oil painting; when defined from the detail perspective, it can include styles such as realistic and abstract; when defined from the lighting perspective, it can include styles such as natural light, artificial light, highlights, and shadows.

[0091] Among them, the terminal display screen style refers to the name of the display screen style and the schematic diagram of the screen style. Further, the terminal displays the screen style determined based on multiple storyboard scripts in the screen style list, and the first screen style is prominently displayed in the screen style list. The screen style list includes multiple screen styles, and the prominently displayed screen style in the screen style list is the screen style determined based on multiple storyboard scripts. Correspondingly, the terminal determines the first screen style from multiple screen styles in the screen style list based on multiple storyboard scripts.

[0092] For example, continuing to refer to Figure 5 , the result display area on the application interface displays the screen style list 503, and the first screen style 504 in the screen style list 503 is prominently displayed.

[0093] In the embodiments of the present disclosure, the user can generate a story in the form of a dialogue, and then generate a storyboard script, or directly use an existing story to generate a storyboard script. The terminal automatically extracts key information from the video description information by artificial intelligence, constructs a story framework, and then generates a storyboard script. And, the artificial intelligence analyzes the scene descriptions, character behaviors, and dialogue contents in the storyboard script to determine a suitable screen style, and the artificial intelligence generates characters according to the story background and character characteristics, thereby improving the quality of the determined screen style and characters.

[0094] In step S302, in response to the confirmation operation on multiple storyboard scripts, at least one first character, and the first screen style, the terminal displays multiple storyboard screens generated based on the multiple storyboard scripts, at least one first character, and the first screen style.

[0095] In some embodiments, multiple storyboard scripts, the first character, and the first screen style are displayed on the application interface, and a second confirmation control is displayed on the application interface. This second confirmation control is used to generate storyboard screens. Therefore, the confirmation operation in step S302 is also the triggering operation on this second confirmation control. For example, continuing to refer to Figure 5, a second confirmation control 505 is also displayed on the application interface.

[0096] In the embodiment of the present disclosure, through the above step S302, the process of the terminal displaying multiple storyboard frames generated based on multiple storyboard scripts in response to the confirmation operation of multiple storyboard scripts is realized. In some other embodiments, the storyboard script can also be modified. Correspondingly, in response to the modification operation of at least one storyboard script among multiple storyboard scripts, the terminal displays at least one second character and a second screen style, and both the at least one second character and the second screen style are determined based on the modified storyboard script. Then, the above process of the terminal displaying multiple storyboard frames generated based on multiple storyboard scripts in response to the confirmation operation of multiple storyboard scripts includes the following implementation manners: In response to the confirmation operation of the modified storyboard script and the confirmation operation of at least one second character and a second screen style, the terminal displays multiple storyboard frames generated based on the modified storyboard script, at least one second character, and a second screen style.

[0097] In this embodiment, the user can modify the generated storyboard script, so that the user can easily adjust various elements such as the camera angle, scene atmosphere, and character actions in the storyboard script, realize personalized creation, improve the controllability of the user over the storyboard script, and the modified storyboard script can also adaptively regenerate the character and screen style, and corresponding storyboard frames can be instantly generated based on the modified content, improving the human-computer interaction efficiency.

[0098] It should be noted that if the elements in the storyboard script modified by the user are associated with other storyboard scripts, the terminal automatically adaptively modifies other associated storyboard scripts. For example, if the name of a certain character in a storyboard script is modified, the terminal automatically adaptively modifies the name of this character in other storyboard scripts.

[0099] In some embodiments, multiple storyboard scripts are displayed in the script editing box, and the user can modify the multiple storyboard scripts in the script editing box, and then the terminal generates at least one second character and a second screen style based on the modified storyboard script and displays them.

[0100] Optionally, the character, screen style, and storyboard script are displayed on the same application interface, and the character and screen style are modified as the storyboard script is modified without the user having to perform additional operations. Or, since the user may modify the storyboard script at multiple locations multiple times, in order to save resources, in response to the confirmation operation of the modified storyboard script, the terminal then displays at least one second character and a second screen style generated based on the modified storyboard script.

[0101] In some embodiments, a replacement control for the storyboard script is displayed on the application interface. Then, the process in which the terminal displays at least one second character and a second screen style in response to a modification operation on at least one of the multiple storyboard scripts includes the following steps: In response to a trigger operation on the replacement control on the application interface, the terminal displays multiple storyboard scripts regenerated by artificial intelligence based on the video description information; in response to a confirmation operation on the regenerated multiple storyboard scripts, the terminal displays at least one second character and a second screen style determined based on the regenerated multiple storyboard scripts.

[0102] For example, continue to refer to Figure 5 , a replacement control 506 for the storyboard script is displayed on the application interface.

[0103] In this embodiment, a replacement control for the storyboard script is provided on the application interface, enabling the artificial intelligence to regenerate the storyboard script with one click and regenerating the character and screen style based on the regenerated storyboard script, which improves the modification efficiency of the storyboard script.

[0104] In the above embodiment, taking the one - click replacement of multiple storyboard scripts as an example, in some other embodiments, replacement controls for each storyboard script are also provided on the application interface. In response to a trigger operation on the replacement control of any storyboard script, a storyboard script regenerated based on the video description information is displayed. Alternatively, the replacement operation on any storyboard script can be a long - press operation or a double - click operation on the storyboard script, etc., which is not specifically limited herein.

[0105] In some embodiments, the terminal can also modify the character and screen style determined by the storyboard script. That is, before the terminal displays multiple storyboard scenes generated based on multiple storyboard scripts, at least one first character, and a first screen style in response to a confirmation operation on the multiple storyboard scripts, at least one first character, and a first screen style, the terminal can respond to a modification operation on at least one of at least one first character and a first screen style and display the modified character and screen style. Correspondingly, in response to a confirmation operation on the multiple storyboard scripts and the modified character and screen style, the terminal displays multiple storyboard scenes generated based on the multiple storyboard scripts and the modified character and screen style.

[0106] In this embodiment, for the character and screen style automatically determined by the storyboard script, the user can also make modifications, and then generate storyboard scenes based on the modified character and screen style. This can improve the user's participation in the generation of storyboard scenes, make the storyboard scenes more personalized, not only improve the creation efficiency, but also increase the controllability of the storyboard scenes, thereby improving the quality of the generated storyboard scenes.

[0107] In some embodiments, the process in which the terminal displays the modified character and screen style in response to a modification operation on at least one of the at least one first character and the first screen style includes at least one of the following implementation manners: (1) In response to a triggering operation on any one of the at least one first character, the terminal displays an editing component for the triggered first character, and the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, the terminal displays the modified first character.

[0108] It should be noted that if any one of the at least one first character is modified, the terminal can automatically modify the storyboard script based on the modified character. For example, if the name of a certain character is modified, the terminal automatically modifies the name of the character in the storyboard script.

[0109] In the embodiments of the present disclosure, taking the terminal displaying a character to mean displaying the name and photo of the character as an example, the triggering operation on the first character can be a click operation on the name or photo of the first character.

[0110] In this embodiment, for automatically generated characters, real-time preview and modification are supported, enabling users to instantaneously view the character effects and optimize them during the generation process, thereby improving the interaction efficiency. Moreover, by providing an editing component for the characters, users can adjust the name, appearance, clothing, voice, personality characteristics, etc. of the characters according to their needs, and then generate character images that better meet the creative requirements. This high degree of customization enables each character to have a unique style and expressiveness, that is, it improves the quality of the generated characters, and thus improves the quality of the video generated based on the characters.

[0111] (2) In response to a character addition operation on the interface where the at least one first character is located, the terminal displays the added character among the at least one first character.

[0112] Optionally, taking the application interface as an example of the interface where the first character is located, a character addition control is displayed on the application interface. In response to a triggering operation on the character addition control, the terminal displays a character addition interface, and the character addition interface is used to set at least one of the name, appearance, clothing, voice, and personality characteristics of the added character. It should be noted that after adding a character among the at least one first character, the terminal automatically modifies the storyboard script based on the added character, and then generates storyboard images based on the modified storyboard script. For example, continuing to refer to Figure 5 , a character addition control 507 is also displayed in the character display area on the application interface.

[0113] In this embodiment, not only can the detailed characteristics such as the name, appearance, clothing, voice color, and personality characteristics of any character be modified, but also new characters can be added, further increasing the user's controllability over the characters, facilitating the generation of storyboard frames that meet the creative requirements based on the modified characters, and then generating a video that meets the creative requirements, thereby improving the quality of the generated video.

[0114] In some embodiments, the terminal also displays multiple pieces of candidate information of the adjacent character on the interface where at least one character is located, including candidate information such as the voice color and personality characteristics of the adjacent character. The user can set any piece of candidate information of the narrator character on this interface. The terminal can also set some other attributes of the video on this interface, such as setting the aspect ratio and resolution of the video, which will not be elaborated in detail here.

[0115] In some embodiments, the first screen style is prominently displayed in the screen style list, and the screen style list includes multiple screen styles. The screen style prominently displayed in the screen style list is the screen style determined based on multiple storyboard scripts. Then, the process in which the terminal displays the modified character and screen style in response to the modification operation on at least one of the at least one first character and the first screen style includes the following implementation manners: in response to the triggering operation on any screen style other than the first screen style in the screen style list, the terminal prominently displays the triggered screen style to obtain the modified screen style.

[0116] In this embodiment, the screen style determined by the storyboard script is prominently displayed in the screen style list. In response to the triggering operation on any other screen style in the screen style list, the triggered screen style is prominently displayed and used as the modified screen style. In this way, a variety of screen styles are provided through the screen style list, improving the diversity of selection, and the screen style can be modified by one-key triggering, improving the human-computer interaction efficiency when modifying the screen style.

[0117] In the embodiments of the present disclosure, after the user inputs video description information on the application interface, the artificial intelligence automatically generates multiple storyboard scripts based on the video description information, generates characters and screen styles based on the storyboard scripts, and displays the storyboard scripts, characters, and screen styles. In this way, more-dimensional information for generating storyboard frames is displayed, improving the information transmission transparency. And after the user confirms this information, the storyboard frames are generated, improving the user's controllability over the generation of storyboard frames, increasing the user's participation, and thus making the generated storyboard frames more in line with the requirements and improving the quality of the generated storyboard frames.

[0118] In step S303, in response to the confirmation operation on multiple storyboard scripts, the terminal displays the thumbnails of multiple storyboard images respectively on the screen preview interface. In response to the trigger operation on the first thumbnail among the multiple thumbnails, the first storyboard image corresponding to the first thumbnail is enlarged and displayed on the screen preview interface, and each storyboard script corresponds to at least one storyboard image.

[0119] Among them, the first thumbnail can be any one of the multiple thumbnails. The trigger operation on any thumbnail can be a click operation on the thumbnail.

[0120] In some embodiments, the thumbnails of multiple storyboard images are arranged and displayed in sequence according to the display order of multiple storyboard images in the video. Further, while the terminal enlarges and displays the first storyboard image corresponding to the first thumbnail on the screen preview interface, it also displays the total duration of the video and the playback time point of the first storyboard image in the video, improving the information transmission rate.

[0121] In some embodiments, a size adjustment control for multiple thumbnails is also displayed on the screen preview interface, which is used to adjust the display area occupied by the triggered thumbnail in the area where multiple thumbnails are located, that is, to adjust the size of the thumbnail.

[0122] In some embodiments, the terminal also displays the first image information of the first storyboard image on the screen preview interface, and the first image information includes at least one of picture description information, character information, and subtitle information; in response to the modification operation on the first image information of the first storyboard image, the terminal displays the first storyboard image regenerated based on the modified first image information.

[0123] Among them, the picture description information includes information such as the scene atmosphere, environmental elements, character actions, and character dialogues of the picture. Optionally, the terminal displays the picture description information in a picture editing box, and the picture editing box is used to modify the picture description information.

[0124] Among them, the character information includes the photos and names of the characters in the first storyboard image. Optionally, in response to the trigger operation on any character, an editing component of the triggered first character is displayed, and the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to the modification operation on the triggered first character based on the editing component, the modified first character is displayed; in response to the character addition operation in the editing interface where at least one first character is located, the added character is displayed among at least one first character, that is, the modification operation on the character on the screen preview interface is the same as the modification operation on the character on the application interface, which will not be elaborated here.

[0125] Among them, the subtitle information includes the line information of each character and the voice-over information in the first storyboard frame, etc. Optionally, the terminal displays the subtitle information in a subtitle editing box, and the subtitle editing box is used to modify the subtitle information.

[0126] In this embodiment, when magnifying and displaying a certain storyboard frame, the frame information of the storyboard frame is also displayed, and by modifying the frame information, a new storyboard frame is regenerated. In this way, users can modify the storyboard frame through simple text prompts or operations without mastering complex image editing skills, reducing the creation threshold and improving the modification efficiency.

[0127] For example, referring to Figure 6 , Figure 6 is a schematic diagram of a frame preview interface shown according to an exemplary embodiment. Among them, a plurality of thumbnails 601 are displayed on the frame preview interface. In response to a trigger operation on a certain thumbnail, the corresponding storyboard frame 602 is magnified and displayed on the frame preview interface, and the frame description information is also displayed in the frame editing box 603, the character information 604 is displayed on the frame preview interface, and the subtitle information is displayed in the subtitle editing box 605.

[0128] In some embodiments, it is also possible to trigger the storyboard frame to modify the storyboard frame. Among them, in response to a trigger operation on any area of the first storyboard frame on the frame preview interface, the terminal displays at least one modification option for the triggered area, and at least one modification option is used to modify the triggered area in at least one dimension; in response to a trigger operation on any one of the at least one modification options, the terminal displays the first storyboard frame modified based on the triggered modification option.

[0129] Among them, based on the difference of the triggered area, the displayed at least one modification option is different. For example, if the triggered area is the area where the character is located, at least one modification option is used to modify the appearance, clothing, etc. of the character. Another example is that if the triggered area is the background area, at least one modification option is used to modify the background brightness, environment, etc. of the storyboard frame.

[0130] In some embodiments, the storyboard frame magnified and displayed on the frame preview interface is default in an unmodifiable state. In response to a trigger operation on the modification control on the storyboard frame, the storyboard frame enters the editing state. In response to a trigger operation on any area of the storyboard frame in the editing state, the corresponding modification option is displayed.

[0131] In this embodiment, by triggering any area on the storyboard frame, the area can be modified. In this way, users can make fine-grained modifications to any area in the storyboard frame without affecting other parts. In this way, only partial redrawing or adjustment of the storyboard frame is required, and there is no need to regenerate the entire storyboard frame, improving the modification efficiency.

[0132] In some embodiments, it is also possible to replace the storyboard screen with one key. Among them, in response to a replacement operation on the first storyboard screen on the screen preview interface, the terminal displays multiple candidate storyboard screens; in response to a trigger operation on any one of the multiple candidate storyboard screens, the terminal replaces the first storyboard screen with the triggered candidate storyboard screen.

[0133] Among them, the multiple candidate storyboard screens can be storyboard screens generated by artificial intelligence based on the storyboard script, and the multiple candidate storyboard screens can also be images on the local device of the terminal, such as images in the local photo album.

[0134] Among them, the replacement operation on the first storyboard screen can be a preset gesture operation such as a long press operation or a double click operation on the first storyboard screen. Or, a replacement control for the first storyboard screen is displayed on the screen preview interface, and the replacement operation on the first storyboard screen is a trigger operation on the replacement control. For example, continue to refer to Figure 6 , a replacement control 606 for the first storyboard screen is displayed on the screen preview interface.

[0135] In this embodiment, multiple candidate storyboard screens are also provided on the screen preview interface, so that users can quickly compare different storyboard screens, and then select the storyboard screen that best meets their needs, saving time and improving the human-computer interaction efficiency. And the one-key replacement function allows users to immediately see the modified effect, improving the information transmission efficiency.

[0136] In some embodiments, after the terminal modifies any storyboard screen, it also displays the thumbnail of the storyboard screen before the modification on the screen preview interface, and triggering the thumbnail can also replace the storyboard screen with the storyboard screen before the modification corresponding to the thumbnail. It should be noted that if any storyboard screen is modified multiple times, the thumbnails of the storyboard screens modified multiple times are all displayed on the screen preview interface, and thus can be replaced with any modified version of the storyboard screen. For example, continue to refer to Figure 6 , a thumbnail 607 of the storyboard screen before the modification of the first storyboard screen is also displayed on the screen preview interface.

[0137] In the embodiments of the present disclosure, after the user confirms the storyboard script, characters, and screen style, a video can be generated with one key. And during the process of generating the video, a function of previewing all storyboard screens is also provided, so that users can modify the screen content at any time, and then the artificial intelligence regenerates the video.

[0138] In the embodiments of the present disclosure, the process of displaying multiple storyboard images generated based on multiple storyboard scripts is implemented through the above step S303. In this embodiment, by displaying thumbnails, multiple storyboard images can be all displayed, improving the information transmission rate. And by triggering any thumbnail, the corresponding storyboard image can be enlarged and displayed, so that any storyboard image can be viewed efficiently and conveniently, improving the interaction efficiency.

[0139] In step S304, in response to the confirmation operation on the multiple storyboard images, the terminal displays a video generated based on the multiple storyboard images.

[0140] In some embodiments, the terminal displays a third confirmation control on the screen preview interface where the multiple storyboard images are located. The confirmation operation on the multiple storyboard images is also the triggering operation on this third confirmation control, and the third confirmation control is also used to generate a video. For example, continue to refer to Figure 6 , and a third confirmation control 608 is also displayed on the screen preview interface.

[0141] In some embodiments, in response to the confirmation operation on the multiple storyboard images, the terminal displays a generation progress indicator in the video display area on the video generation interface, and the generation progress indicator is used to indicate the generation progress of the video. For example, refer to Figure 7 , Figure 7 FIG. is a schematic diagram of a video generation interface shown according to an exemplary embodiment. Among them, a generation progress indicator 702 is displayed in the video display area 701 on the video generation interface.

[0142] In some embodiments, a dialogue area is still reserved and displayed on the video generation interface. For example, continue to refer to Figure 7 , and the left area of the video generation interface is the dialogue area.

[0143] In some embodiments, the covers of the videos previously generated by the user are also displayed on the video generation interface. These videos can include videos generated based on feature indication information and videos generated based on existing stories. Further, by triggering the cover of any video, the terminal can play the video. For example, continue to refer to Figure 7 , and the covers 703 of the previously generated videos are also displayed on the video generation interface.

[0144] In some embodiments, a video publishing control is also displayed on the video generation interface. The video publishing control is used to publish the video to the user's social account and can also be used to share the video with friends. For example, continue to refer to Figure 7 , and a video publishing control 704 is also displayed on the video generation interface.

[0145] In some embodiments, the terminal displays a video on the video preview interface, and thumbnails of multiple sub-scenes are also displayed on the video preview interface; in response to a triggering operation on the second thumbnail among the multiple thumbnails, the terminal magnifies and displays the second sub-scene corresponding to the second thumbnail on the video preview interface and displays the second scene information of the second sub-scene on the video preview interface, where the second scene information includes at least one of scene description information, character information, subtitle information, background music information, and camera movement mode; in response to a modification operation on the second scene information of the second sub-scene, the terminal displays the second sub-scene regenerated based on the modified second scene information, and the regenerated second sub-scene is used to regenerate the video.

[0146] Among them, the second thumbnail can be any one of the multiple thumbnails. The triggering operation on any thumbnail can be a click operation on the thumbnail.

[0147] For example, referring to Figure 8 , Figure 8 FIG. is a schematic diagram of a video preview interface shown according to an exemplary embodiment. Among them, a certain sub-scene 801 is displayed on the video preview interface, and a scene editing box 802 for the sub-scene is also displayed on the video preview interface. The scene editing box 802 is used to modify the scene description information of the sub-scene, and a regenerate control 803 is also displayed on the video preview interface for regenerating the sub-scene with one key. Among them, multiple property controls are also displayed on the video preview interface, which are respectively used to modify properties such as subtitles and background music of the video, and are not specifically limited here. Further, a save control is also displayed on the video preview interface to save the current modification. In response to subsequent completion of modification of other properties, the sub-scene is regenerated based on the saved modifications. Similarly, a thumbnail of the sub-scene before the modification of the second sub-scene is also displayed on the video preview interface to facilitate replacing the second sub-scene with any previous modified version of the sub-scene.

[0148] Again, for example, referring to Figure 9 , Figure 9 FIG. is a schematic diagram of a video preview interface shown according to an exemplary embodiment. Among them, when the camera movement control on the video preview interface is triggered, multiple preset camera movement modes 901 are displayed on the video preview interface. The camera movement mode of the sub-scene can be modified by selecting any one of them. When the regenerate control is triggered, the sub-scene is regenerated with one key based on the selected camera movement mode. Optionally, the selected camera movement mode is highlighted.

[0149] In this embodiment, after the video is generated, the user can also modify the storyboard frames of the video and regenerate the video, which can significantly enhance the personalized features of the video, not only improving the creation efficiency, reducing the creation cost, but also increasing the controllability of the video content, and enhancing the flexibility and diversity of video creation. For example, by modifying the expressions, actions, languages, etc. of the characters, the characters can be made more vivid to better convey the emotions and the core of the story. By modifying the background music, it can better match the theme and emotional trend of the video. By modifying the camera movement method, it can attract the user's attention and enhance the tension and rhythm of the story, that is, by modifying the frame information, the quality of the video can be further improved.

[0150] In some embodiments, the terminal directly displays the thumbnail on the video preview interface where the video is located. In other embodiments, in response to an editing operation on the video, the terminal displays the thumbnail on the video preview interface, that is, the thumbnail is displayed on the video preview interface only when the user has a modification requirement for the video.

[0151] In some embodiments, the process in which the terminal displays the video generated based on multiple storyboard frames in response to a confirmation operation on multiple storyboard frames further includes the following implementation manners: in response to a confirmation operation on multiple storyboard frames, the terminal displays multiple videos generated based on multiple storyboard frames, and the third frame information of at least one storyboard frame is different among the multiple videos, and the third frame information includes at least one of subtitle information, background music information, and camera movement method.

[0152] Among them, only taking the third frame information including at least one of subtitle information, background music information, and camera movement method as an example for illustration, the third frame information may also include characters, frame style, etc., which are not specifically limited herein.

[0153] In the embodiments of the present disclosure, the terminal can automatically generate a video based on the storyboard frames, and can generate diverse videos based on the differences in the frame information of the storyboard frames, which is beneficial for the user to select a satisfactory video from them, improving the user experience and the quality of video generation.

[0154] In some embodiments, before the terminal displays the video generated based on multiple storyboard frames in response to a confirmation operation on multiple storyboard frames, the terminal can also respond to an operation for adjusting the order of multiple storyboard frames and display the multiple storyboard frames after the order is adjusted. Correspondingly, the terminal responds to a confirmation operation on the multiple storyboard frames after the order is adjusted and displays the video generated based on the multiple storyboard frames after the order is adjusted.

[0155] Among them, when multiple storyboard frames are displayed on the frame preview interface and are arranged in sequence according to the display order in the video, optionally, in response to a dragging operation on at least one storyboard frame, the order of the multiple storyboard frames is adjusted, and the storyboard frames after the order is adjusted are displayed.

[0156] In some embodiments, when generating a video based on multiple storyboard frames, multiple storyboard scripts are also combined to generate the video. Optionally, the terminal automatically adjusts multiple storyboard scripts based on the storyboard frames after the adjustment order, and then generates a video based on the multiple storyboard frames and multiple storyboard scripts after the adjustment order.

[0157] In the embodiments of the present disclosure, for a given multiple storyboard frames, the user can adjust the order of the multiple storyboard frames, and then generate a video based on the storyboard frames after the adjustment order. Since different storyboard frames can correspond to different storylines, the user can independently control the development order of each storyline in the video, making the generated video more in line with the user's needs and improving the video generation quality.

[0158] It should be noted that the storyboard scripts, characters, picture styles, storyboard frames, etc. displayed on the terminal can be generated by the terminal itself. For example, a generative model is embedded in the terminal, and the generative model is an artificial intelligence model. Then, the terminal generates storyboard script characters, picture styles, storyboard frames, etc. through the generative model. Alternatively, the terminal generates storyboard scripts, characters, picture styles, storyboard frames, etc. through a server. Further, the server can also be generated through a generative model, which is not specifically limited herein.

[0159] In the embodiments of the present disclosure, by generating story content, storyboard scripts, and storyboard frames through dialogue, the artificial intelligence constructs a story framework by utilizing existing stories or obtaining key information from the feature indication information of the video, reducing the dependence on the difficult-to-obtain innovative inspiration in traditional creation, lowering the creation threshold, providing users with more creative ideas and possibilities, and making story creation easier to carry out. Moreover, it changes the long process from the early setting to the late production in the traditional way, can quickly generate a story framework, storyboard scripts, characters, videos, etc. according to the input video description information, greatly improving the creation efficiency and reducing the creation time cost. Moreover, compared with the high requirements for a large amount of labor and professional equipment in traditional creation, this method uses artificial intelligence to automatically complete multiple tasks, reducing the input of labor and equipment, lowering the creation cost, and making video creation more universal. Moreover, it supports one-key video generation and can modify video elements such as storyboard scripts and storyboard frames during the process. After generating the video, it can also edit the camera movement mode, sound effects, etc. of the storyboard frames, enabling users to adjust video elements at any time according to their needs, making the finally generated video more in line with expectations, improving the creation accuracy and quality, giving users more independent control, and improving the quality of the generated video.

[0160] In the embodiments of the present disclosure, videos can be generated through conversational interaction technology. The user's instructions, descriptions, and other information in the conversation are accurately understood semantically by artificial intelligence, and key information is extracted. Based on this, coherent and logically reasonable story content is intelligently generated, and a complete storyboard script is further constructed, improving the generation quality. Optionally, the terminal implements this process through advanced natural language processing algorithms and intelligent story architecture building models. Further, based on the multiple information in the storyboard script, such as scenes, character behaviors, dialogues, etc., artificial intelligence uses technologies such as image recognition and analysis, and big data matching to accurately recommend suitable picture styles. At the same time, combined with deep learning models, unique and story-fitting character roles are automatically generated according to the story background and character characteristics. After artificial intelligence completes the preparation of the preliminary creative elements, the user triggers a one-click generation instruction, and the background quickly integrates resources, calls the image rendering engine, and renders and generates videos in real time according to the established picture style, characters, and storyboard script, improving the efficiency of video generation.

[0161] The embodiments of the present disclosure provide a video generation method. For the video description information input by the user, multiple storyboard scripts can be generated by artificial intelligence. After the user confirms the storyboard scripts, multiple storyboard images are generated. After the user confirms the storyboard images, the final video is generated. This method can be automatically generated by artificial intelligence only by the user inputting video description information, improving the video generation efficiency. And, during the process of generating the video, the generated storyboard scripts, characters, picture styles, and storyboard images are displayed, and the video is generated after the user confirms, improving the user's participation, that is, improving the user's controllability of the generated video. In this way, not only the video generation efficiency is improved, but also the generated video can better meet the user's needs, thereby improving the quality of the generated video.

[0162] Figure 10 is a block diagram of a video generation device shown according to an exemplary embodiment. Refer to Figure 10 , the device includes: The first display unit 1001 is configured to execute, in response to a confirmation operation on video description information input on the application interface, display multiple storyboard scripts generated based on the video description information, and the multiple storyboard scripts correspond to at least one storyline; The second display unit 1002 is configured to execute, in response to a confirmation operation on the multiple storyboard scripts, display multiple storyboard images generated based on the multiple storyboard scripts, and each storyboard script corresponds to at least one storyboard image; The third display unit 1003 is configured to execute, in response to a confirmation operation on the multiple storyboard images, display a video generated based on the multiple storyboard images.

[0163] In some embodiments, the first display unit 1001 is further configured to execute: Display at least one first character and a first screen style, where both the at least one first character and the first screen style are determined based on a plurality of storyboard scripts; A second display unit 1002, configured to perform: In response to a confirmation operation on the plurality of storyboard scripts, the at least one first character, and the first screen style, display a plurality of storyboard images generated based on the plurality of storyboard scripts, the at least one first character, and the first screen style.

[0164] In some embodiments, the apparatus further includes a first modification unit, configured to perform: In response to a modification operation on at least one of the plurality of storyboard scripts, display at least one second character and a second screen style, where both the at least one second character and the second screen style are determined based on the modified storyboard script; The second display unit 1002, configured to perform: In response to a confirmation operation on the modified storyboard script, the at least one second character, and the second screen style, display a plurality of storyboard images generated based on the modified storyboard script, the at least one second character, and the second screen style.

[0165] In some embodiments, the first modification unit is configured to perform: In response to a triggering operation on a replacement control on the application interface, display a plurality of storyboard scripts regenerated by artificial intelligence based on video description information; In response to a confirmation operation on the regenerated plurality of storyboard scripts, display at least one second character and a second screen style determined based on the regenerated plurality of storyboard scripts.

[0166] In some embodiments, the apparatus further includes a second modification unit, configured to perform: In response to a modification operation on at least one of the at least one first character and the first screen style, display the modified character and screen style; The second display unit 1002, configured to perform: In response to a confirmation operation on the plurality of storyboard scripts and the modified character and screen style, display a plurality of storyboard images generated based on the plurality of storyboard scripts and the modified character and screen style.

[0167] In some embodiments, the second modification unit is configured to perform at least one of the following: In response to a triggering operation on any one of the at least one first character, display an editing component for the triggered first character, where the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality characteristics of the triggered first character; in response to a modification operation on the triggered first character based on the editing component, display the modified first character; In response to a character addition operation on the interface of at least one first character, the added character is displayed among the at least one first character.

[0168] In some embodiments, the first screen style is prominently displayed in a screen style list that includes multiple screen styles, and the prominently displayed screen style in the screen style list is a screen style determined based on multiple storyboard scripts. A second modification unit, configured to perform: In response to a trigger operation on any screen style other than the first screen style in the screen style list, the triggered screen style is prominently displayed to obtain a modified screen style.

[0169] In some embodiments, the second display unit 1002 is configured to perform: Thumbnails of multiple storyboard frames are displayed on a screen preview interface. In response to a trigger operation on the first thumbnail among the multiple thumbnails, the first storyboard frame corresponding to the first thumbnail is magnified and displayed on the screen preview interface.

[0170] In some embodiments, the second display unit 1002 is further configured to perform: First screen information of the first storyboard frame is displayed on the screen preview interface, and the first screen information includes at least one of screen description information, character information, and subtitle information. In response to a modification operation on the first screen information of the first storyboard frame, a first storyboard frame regenerated based on the modified first screen information is displayed.

[0171] In some embodiments, the device further includes a third modification unit, configured to perform: In response to a trigger operation on any area of the first storyboard frame on the screen preview interface, at least one modification option for the triggered area is displayed, and the at least one modification option is used to modify the triggered area in at least one dimension. In response to a trigger operation on any one of the at least one modification options, a first storyboard frame modified based on the triggered modification option is displayed.

[0172] In some embodiments, the device further includes a replacement unit, configured to perform: In response to a replacement operation on the first storyboard frame on the screen preview interface, multiple candidate storyboard frames are displayed. In response to a trigger operation on any one of the multiple candidate storyboard frames, the first storyboard frame is replaced with the triggered candidate storyboard frame.

[0173] In some embodiments, the third display unit 1003 is configured to perform: A video is displayed on a video preview interface, and thumbnails of multiple storyboard frames are also displayed on the video preview interface; The apparatus further includes a fourth modification unit configured to perform: In response to a triggering operation on a second thumbnail among the multiple thumbnails, the second storyboard frame corresponding to the second thumbnail is enlarged and displayed on the video preview interface, and second frame information of the second storyboard frame is displayed on the video preview interface, where the second frame information includes at least one of frame description information, character information, subtitle information, background music information, and camera movement method; In response to a modification operation on the second frame information of the second storyboard frame, a second storyboard frame regenerated based on the modified second frame information is displayed, and the regenerated second storyboard frame is used to regenerate the video.

[0174] In some embodiments, the video description information includes feature indication information of a video to be generated, and the first display unit 1001 is configured to perform: In response to a confirmation operation on the video description information input on the application interface, a first story generated by artificial intelligence based on the video description information is displayed; In response to a confirmation operation on the first story, multiple storyboard scripts generated based on the first story are displayed; or, in response to a modification operation on the first story, multiple storyboard scripts generated based on the modified first story are displayed.

[0175] In some embodiments, the video description information includes a second story, and the video description information is used to indicate generation of a video based on the second story.

[0176] In some embodiments, the third display unit 1003 is configured to perform: In response to a confirmation operation on multiple storyboard frames, multiple videos generated based on the multiple storyboard frames are displayed, and third frame information of at least one storyboard frame among the multiple videos is different, where the third frame information includes at least one of subtitle information, background music information, and camera movement method.

[0177] In some embodiments, the apparatus further includes an adjustment unit configured to perform: In response to an operation of adjusting the order of multiple storyboard frames, the multiple storyboard frames with the adjusted order are displayed; The third display unit 1003 is configured to perform: In response to a confirmation operation on the multiple storyboard frames with the adjusted order, a video generated based on the multiple storyboard frames with the adjusted order is displayed.

[0178] An embodiment of the present disclosure provides a video generation device. For the video description information input by a user, the device can automatically generate multiple storyboard scripts. After the user confirms the storyboard scripts, multiple storyboard images are generated. After the user confirms the storyboard images, the final video is generated. The device only needs the user to input the video description information to automatically generate, improving the video generation efficiency. Moreover, in the middle of video generation, the generated storyboard scripts and storyboard images are displayed, and the video is generated after the user's confirmation, improving the user's participation, that is, improving the controllability of the video generated by the user. This not only improves the video generation efficiency, but also enables the generated video to better meet the user's needs, thereby improving the quality of the generated video.

[0179] Regarding the device in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0180] In some embodiments, the electronic device is provided as a terminal. Figure 11 The structural block diagram of a terminal 1100 provided by an exemplary embodiment of the present disclosure is shown. The terminal 1100 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 1100 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.

[0181] Generally, the terminal 1100 includes a processor 1101 and a memory 1102.

[0182] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.

[0183] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one program code, and the at least one program code is used to be executed by the processor 1101 to implement the video generation method provided in the method embodiments of the present disclosure.

[0184] In some embodiments, the terminal 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.

[0185] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0186] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1104 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 may further include a circuit related to NFC (Near Field Communication), and this disclosure does not limit this.

[0187] The display screen 1105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input as control signals to the processor 1101 for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, which is disposed on the front panel of the terminal 1100; in other embodiments, there may be at least two display screens 1105, which are respectively disposed on different surfaces of the terminal 1100 or are in a folding design; in still other embodiments, the display screen 1105 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal 1100. Even further, the display screen 1105 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 1105 can be prepared using materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0188] The camera module 1106 is used to capture images or videos. Optionally, the camera module 1106 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting and VR (Virtual Reality) shooting functions or other fusion shooting functions by fusing the main camera and the wide-angle camera. In some embodiments, the camera module 1106 may further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0189] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0190] The power supply 1108 is used to supply power to each component in the terminal 1100. The power supply 1108 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1108 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0191] Those skilled in the art can understand that Figure 11 the structure shown in does not limit the terminal 1100, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0192] In an exemplary embodiment, there is also provided a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the above video generation method. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0193] In an exemplary embodiment, there is also provided a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above video generation method is implemented.

[0194] In some embodiments, the computer program product related to the embodiments of the present disclosure may be deployed to be executed on one electronic device, or on multiple electronic devices located at one place. Or, on multiple electronic devices distributed at multiple places and interconnected through a communication network. The multiple electronic devices distributed at multiple places and interconnected through a communication network may form a blockchain system.

[0195] Other embodiments of the present disclosure will be readily apparent to those skilled in the art in view of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims. All of the above alternative technical solutions may be combined arbitrarily to form alternative embodiments of the present application, which will not be elaborated herein one by one.

[0196] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video generation method, characterized in that, The method includes: In response to a confirmation operation on video description information input on an application interface, displaying a plurality of storyboard scripts generated based on the video description information, where the plurality of storyboard scripts correspond to at least one storyline; In response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, where each storyboard script corresponds to at least one storyboard frame; In response to a confirmation operation on the plurality of storyboard frames, displaying a video generated based on the plurality of storyboard frames.

2. The video generation method according to claim 1, wherein Before the step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the method further includes: Displaying at least one first character and a first picture style, where both the at least one first character and the first picture style are determined based on the plurality of storyboard scripts; The step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, includes: In response to a confirmation operation on the plurality of storyboard scripts, the at least one first character, and the first picture style, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the at least one first character, and the first picture style.

3. The video generation method according to claim 2, wherein Before the step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the method further includes: In response to a modification operation on at least one of the plurality of storyboard scripts, displaying at least one second character and a second picture style, where both the at least one second character and the second picture style are determined based on the modified storyboard script; The step of, in response to a confirmation operation on the plurality of storyboard scripts, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, includes: In response to a confirmation operation on the modified storyboard script, the at least one second character, and the second picture style, displaying a plurality of storyboard frames generated based on the modified storyboard script, the at least one second character, and the second picture style.

4. The video generation method according to claim 2, wherein Before the step of, in response to a confirmation operation on the plurality of storyboard scripts, the at least one first character, and the first picture style, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the at least one first character, and the first picture style, the method further includes: In response to a modification operation on at least one of the at least one first character and the first picture style, displaying the modified character and picture style; The step of, in response to a confirmation operation on the plurality of storyboard scripts, the at least one first character, and the first picture style, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts, the at least one first character, and the first picture style, includes: In response to a confirmation operation on the plurality of storyboard scripts and the modified character and picture style, displaying a plurality of storyboard frames generated based on the plurality of storyboard scripts and the modified character and picture style.

5. The video generation method according to claim 4, wherein In response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style, including at least one of the following: In response to a trigger operation on any one of the at least one first character, displaying an editing component for the triggered first character, where the editing component is used to modify at least one of the name, appearance, clothing, voice, and personality traits of the triggered first character; In response to a modification operation on the triggered first character based on the editing component, displaying the modified first character; In response to a character addition operation on the interface where the at least one first character is located, displaying the added character among the at least one first character.

6. The video generation method according to claim 4, wherein The first screen style is prominently displayed in a screen style list, where the screen style list includes multiple screen styles, and the prominently displayed screen style in the screen style list is the screen style determined based on the multiple storyboard scripts; In response to a modification operation on at least one of the at least one first character and the first screen style, displaying the modified character and screen style, including: In response to a trigger operation on any screen style other than the first screen style in the screen style list, prominently displaying the triggered screen style to obtain the modified screen style.

7. The video generation method according to claim 1, wherein The displaying of multiple storyboard images generated based on the multiple storyboard scripts includes: Displaying thumbnails of the multiple storyboard images on a screen preview interface, and in response to a trigger operation on a first thumbnail among the multiple thumbnails, magnifying and displaying the first storyboard image corresponding to the first thumbnail on the screen preview interface.

8. The video generation method according to claim 7, wherein The method further includes: Displaying first screen information of the first storyboard image on the screen preview interface, where the first screen information includes at least one of screen description information, character information, and caption information; In response to a modification operation on the first screen information of the first storyboard image, displaying the first storyboard image regenerated based on the modified first screen information.

9. The video generation method according to claim 7, wherein The method further includes: In response to a trigger operation on any area of the first storyboard image on the screen preview interface, displaying at least one modification option for the triggered area, where the at least one modification option is used to modify the triggered area in at least one dimension; In response to a trigger operation on any one of the at least one modification option, displaying the first storyboard image modified based on the triggered modification option.

10. The video generation method according to claim 1, wherein The displaying of a video generated based on the multiple storyboard images includes: Displaying the video on a video preview interface, and thumbnails of the multiple storyboard images are also displayed on the video preview interface; The method further includes: In response to a trigger operation on a second thumbnail among the multiple thumbnails, magnifying and displaying the second storyboard image corresponding to the second thumbnail on the video preview interface and displaying second screen information of the second storyboard image on the video preview interface, where the second screen information includes at least one of screen description information, character information, caption information, background music information, and camera movement method; In response to a modification operation on the second screen information of the second storyboard screen, display the second storyboard screen regenerated based on the modified second screen information, where the regenerated second storyboard screen is used to regenerate a video.

11. The video generation method according to claim 1, wherein The video description information includes feature indication information of the video to be generated. In response to a confirmation operation on the video description information input on the application interface, display a plurality of storyboard scripts generated based on the video description information, including: In response to a confirmation operation on the video description information input on the application interface, display a first story generated by an artificial intelligence based on the video description information. In response to a confirmation operation on the first story, display a plurality of storyboard scripts generated based on the first story; or, in response to a modification operation on the first story, display a plurality of storyboard scripts generated based on the modified first story.

12. The video generation method according to claim 1, wherein The video description information includes a second story, and the video description information is used to indicate generating a video based on the second story.

13. The video generation method according to claim 1, characterized in that, The response to a confirmation operation on the plurality of storyboard screens, displaying a video generated based on the plurality of storyboard screens, includes: In response to a confirmation operation on the plurality of storyboard screens, display a plurality of videos generated based on the plurality of storyboard screens, where the third screen information of at least one storyboard screen is different among the plurality of videos, and the third screen information includes at least one of subtitle information, background music information, and camera movement method.

14. The video generation method according to claim 1, wherein Before the response to a confirmation operation on the plurality of storyboard screens and displaying a video generated based on the plurality of storyboard screens, the method further includes: In response to an operation of adjusting the order of the plurality of storyboard screens, display the plurality of storyboard screens with the adjusted order. The response to a confirmation operation on the plurality of storyboard screens, displaying a video generated based on the plurality of storyboard screens, includes: In response to a confirmation operation on the plurality of storyboard screens with the adjusted order, display a video generated based on the plurality of storyboard screens with the adjusted order.

15. A video generation device, characterized in that, The apparatus includes: A first display unit configured to perform, in response to a confirmation operation on the video description information input on the application interface, display a plurality of storyboard scripts generated based on the video description information, where the plurality of storyboard scripts correspond to at least one plot. A second display unit configured to perform, in response to a confirmation operation on the plurality of storyboard scripts, display a plurality of storyboard screens generated based on the plurality of storyboard scripts, and each storyboard script corresponds to at least one storyboard screen. A third display unit configured to perform, in response to a confirmation operation on the plurality of storyboard screens, display a video generated based on the plurality of storyboard screens.

16. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the video generation method according to any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generation method according to any one of claims 1 to 14.

18. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the video generation method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Generative artificial intelligence-based novel tweet video generation method and system

    CN117078782A

  • Intelligent sub-shot script generation system and method

    CN118690739A

  • Video generation method, electronic device, storage medium and computer program product

    CN118748738A

  • Video generation method and device, readable medium, electronic equipment and program product

    CN119364091A

  • Video generation system and method based on intelligent agent

    CN119420994A

Cited By

  • Video generation method and device, equipment and medium

    CN121000949A

  • Video generation method and device, equipment and storage medium

    CN121126087A