Video generation method and device, electronic equipment, storage medium and program product

By responding to the video type selected by the user in the video generation interface, displaying the corresponding content input area and generating the target video, the problem of not combining the video type requirements in the prior art is solved, and efficient video generation and interaction are achieved.

CN120321457APending Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410058005.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art fails to effectively combine the user's demand for video types when generating videos, resulting in the need to edit after generation, which reduces the human-computer interaction efficiency and video generation efficiency.

Method used

Provide a video generation method, by displaying video generation controls in the video generation interface, displaying the corresponding content input area in response to the video type selected by the user, and generating the target video type based on the content input by the user, supporting parameter settings and storyboard disassembly, and optimizing the video generation process.

Benefits of technology

It improves video generation efficiency and human-computer interaction efficiency, reduces users' subsequent editing operations for generating videos, and meets users' dual needs for video content and types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321457A_ABST
    Figure CN120321457A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment, a computer readable storage medium and a computer program product, and the method comprises the steps: displaying a video generation interface which comprises at least one video generation control, and the at least one video generation control is used for generating videos of at least two video types; on the basis of the at least one video generation control, responding to a condition that a target video type in the at least two video types is in a selected state, and displaying a content input area corresponding to the target video type; in response to a content input operation in the content input area, displaying the input target content; and in response to the video generation instruction, displaying a video of a target video type generated based on the target content. According to the invention, the video generation efficiency and the man-machine interaction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technologies, and in particular, to a video generation method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] There are various types of videos, such as short videos, long videos, etc. The video parameters corresponding to different types of videos are different, and the video generation processes are also different. However, in the related technologies when generating videos, most are to directly generate videos based on the user's requirements for video content, that is, only considering the user's requirements for video content while ignoring the user's requirements for video types. In this way, after the video is generated, the user still needs to edit the video so that the edited video meets the user's requirements for video content and at the same time meets the user's requirements for video types. This not only makes the human-computer interaction efficiency too low, but also further reduces the video generation efficiency. Summary of the Invention

[0003] Embodiments of the present application provide a video generation method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the video generation efficiency and the human-computer interaction efficiency.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] Embodiments of the present application provide a video generation method, including:

[0006] Display a video generation interface, where the video generation interface includes at least one video generation control for generating videos of at least two video types;

[0007] Based on the at least one video generation control, in response to a target video type among the at least two video types being in a selected state, display a content input area corresponding to the target video type;

[0008] In response to a content input operation in the content input area, display the input target content;

[0009] In response to a video generation instruction, display the video of the target video type generated based on the target content.

[0010] Embodiments of the present application provide a video generation apparatus, including:

[0011] A first display module for displaying a video generation interface, where the video generation interface includes at least one video generation control for generating videos of at least two video types;

[0012] A second display module, configured to generate controls based on the at least one video, and display a content input area corresponding to the target video type in response to the target video type among the at least two video types being in a selected state;

[0013] A third display module, configured to display the input target content in response to a content input operation in the content input area;

[0014] A fourth display module, configured to display a video of the target video type generated based on the target content in response to a video generation instruction.

[0015] In the above solution, the number of the video generation controls is multiple, and different video generation controls correspond to different video types. The first display module is further configured to display a video generation interface including multiple video generation controls in response to an open instruction for the video generation interface. The apparatus further includes a determination module, and the determination module is configured to determine that the target video type is in a selected state in response to the video generation control corresponding to the target video type in the video generation interface being in a selected state.

[0016] In the above solution, the apparatus further includes a parameter setting module. The parameter setting module is further configured to display a parameter setting area in the video generation interface, where the parameter setting area is used to set parameters of the generated video. Based on the parameter setting area, in response to a parameter setting operation, display the video parameters set for the video of the target video type. The fourth display module is further configured to display a video of the target video type with the video parameters generated based on the target content in response to a video generation instruction.

[0017] In the above solution, the target video type is a long video type, and the long video includes multiple storyboards. The apparatus further includes a fifth display module, and the fifth display module is configured to display a storyboard disassembly interface and display multiple storyboard texts determined based on the target content in the storyboard disassembly interface in response to a determination instruction for the input target content. Each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard. The fourth display module is further configured to display a video of the target video type including storyboards corresponding to the respective storyboard texts generated in response to a video generation instruction triggered based on the multiple storyboard texts.

[0018] In the above solution, the device further includes a storyboard text quantity adjustment module. The storyboard text quantity adjustment module is used to display a storyboard text quantity adjustment control, and the storyboard text quantity adjustment control is used to add new storyboard texts; in response to a storyboard text addition operation triggered based on the storyboard text quantity adjustment control, in the storyboard disassembly interface, display the added new storyboard texts.

[0019] In the above solution, a first confirmation control is further displayed in the storyboard disassembly interface. The device further includes a sixth display module. The sixth display module is used to display a storyboard generation interface in response to a trigger operation on the first confirmation control; in the storyboard generation interface, display the storyboards corresponding to each of the storyboard texts, and display a second confirmation control; in response to a trigger operation on the second confirmation control, receive the video generation instruction.

[0020] In the above solution, each storyboard is associated with a parameter setting control, and the parameter setting control is used to set the storyboard parameters of the corresponding storyboard; the device further includes a storyboard parameter setting module. The storyboard parameter setting module is used to display the target storyboard parameters set for the target storyboard in response to a parameter setting operation triggered based on the parameter setting control associated with the target storyboard; in response to a confirmation instruction for the target storyboard parameters, update the storyboard parameters of the target storyboard to the target storyboard parameters.

[0021] In the above solution, the target video type is a long video type, and the long video type of video includes multiple storyboards; the device further includes a seventh display module. The seventh display module is used to display a storyboard generation interface in response to a confirmation instruction for the input target content, and display multiple storyboard texts and multiple storyboards in the storyboard generation interface, with each storyboard text corresponding to a storyboard, and the storyboard being generated based on the corresponding storyboard text; the fourth display module 4554 is further used to display a video of the target video type including the multiple storyboards in response to a video generation instruction based on the multiple storyboard texts and multiple storyboards.

[0022] In the above solution, the device further includes a selection module. The selection module is used to control the first target storyboard to be in a selected state in response to a selection operation on the first target storyboard among the multiple storyboards; the fourth display module is further used to display a video of the target video type including the first target storyboard in response to a video generation instruction based on the first target storyboard in the selected state and the storyboard text of the first target storyboard.

[0023] In the above solution, there is an arrangement order for the multiple storyboards, and videos formed by the multiple storyboards with different arrangement orders are different. A sorting adjustment control is also displayed on the storyboard generation interface; the device further includes a sorting adjustment module, and the sorting adjustment module is configured to, in response to a sorting adjustment operation for the multiple storyboards triggered based on the sorting adjustment control, change the sorting of the multiple storyboards from the current arrangement order to a target arrangement order; the fourth display module 4554 is further configured to, based on the multiple storyboard texts and the multiple storyboards, in response to a video generation instruction, display a video of a target video type including the multiple storyboards generated based on the target arrangement order.

[0024] In the above solution, a storyboard update control for updating the content of the storyboard is also displayed on the storyboard generation interface; the device further includes a storyboard update module, and the storyboard update module is configured to, based on the storyboard update control, in response to a content update operation for a second target storyboard among the multiple storyboards, update the second target storyboard to a first new storyboard, and the content of the first new storyboard is associated with the content of the second target storyboard.

[0025] In the above solution, there is an arrangement order for the multiple storyboards, and there is an upward association control for the third target storyboard among the multiple storyboards. The upward association control is configured to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard, and the third target storyboard is any one of the multiple storyboards except the first storyboard; the device further includes an association module, and the association module is configured to, in response to a trigger operation for the upward association control, update the third target storyboard to a second new storyboard.

[0026] In the above solution, the device further includes an introduction module, and the introduction module is configured to display function introduction information if the long video type of video is generated for the first time by the current account; wherein, the function introduction information is used to introduce the function of the upward association control.

[0027] In the above solution, the seventh display module is further configured to display multiple storyboard texts on the storyboard generation interface, and display multiple candidate storyboards corresponding to the respective storyboard texts; for each of the storyboard texts, in response to a selection operation for a target candidate storyboard among the multiple candidate storyboards, control the target candidate storyboard to be in a selected state, and determine the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

[0028] In the above solution, the device further includes an eighth display module. The eighth display module is configured to display at least one video template; in response to a selection operation for a target video template among the at least one video template, display a template content description of the target video template and a corresponding first editing control; in response to a triggering operation for the first editing control, determine the triggering operation as an input operation for the content input area; the third display module is further configured to, in response to the triggering operation for the first editing control, determine the triggering operation as the content input operation, display the template content description of the target video template in the content input area, and determine the displayed template content description as the input target content.

[0029] In the above solution, the device further includes a ninth display module. The ninth display module is configured to, in response to a triggering operation for a target video template among the at least one video template, display a details interface including details information of the target video template; wherein, the details interface further includes at least one of a second editing control and a first new video generation control. The second editing control is configured to regenerate a video of the target video type based on the template content description, and the first new video generation control is configured to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

[0030] In the above solution, the fourth display module is further configured to, in response to the video generation instruction, display a video of the target video type generated based on the target content on a video display interface; wherein, the video display interface further includes at least one of the following: details information of the generated video, a third editing control, and a first new video generation control; wherein, the third editing control is configured to regenerate a video of the target video type based on the target content, and the first new video generation control is configured to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

[0031] In the above solution, the video display interface includes the third editing control, the target video type is the long video type, and the long video type of video includes multiple storyboards; the device further includes a tenth display module, and the tenth display module is configured to, in response to a trigger operation on the third editing control, display a storyboard generation interface, and display a plurality of storyboard texts and a plurality of storyboards determined based on the target content on the storyboard generation interface; based on the storyboard generation interface, in response to an editing operation on the plurality of storyboard texts and the plurality of storyboards, display the edited plurality of storyboard texts and the plurality of storyboards; in response to a video generation instruction, display a long video type of video generated based on the edited plurality of storyboard texts and the plurality of storyboards.

[0032] In the above solution, the video display interface includes the third editing control, and the target video type is the short video type; the device further includes an eleventh display module, and the eleventh display module is configured to, in response to a trigger operation on the third editing control, display a video generation interface including a content input area, and the target content is displayed in the content input area; in response to an editing operation on the target content, display a new target content obtained by editing the target content; in response to a video generation instruction, display a short video type of video generated based on the new target content.

[0033] In the above solution, the fourth display module is further configured to, in response to a video generation instruction, display a first progress prompt message, and the first progress prompt message is used to prompt the generation progress of the video of the target video type; when the first progress prompt message indicates that the video of the target video type is generated, cancel the display of the progress prompt message, and display the video of the target video type generated based on the target content.

[0034] In the above solution, the device further includes a video generation prompt module, and the video generation prompt module is configured to display a video generation prompt message, and the video generation prompt message is used to prompt the number of videos that can be generated within a target time; the fourth display module is further configured to, in response to a video generation instruction, when the video generation prompt message indicates that the number of videos that can be generated within the target time is not zero, display the video of the target video type generated based on the target content.

[0035] In the above solution, the device further includes a twelfth display module. The twelfth display module is configured to display a sharing link of another video received; in response to a display instruction triggered based on the sharing link, if the current account has the playing permission for the other video, display the playing interface of the other video, and display the detailed information of the other video and a first new video generation control on the playing interface; the first display module is further configured to display the video generation interface in response to a triggering operation on the first new video generation control.

[0036] In the above solution, the video generation interface further includes a video display control. The video display control is used to display at least one asset video. The at least one asset video includes at least one of a successfully generated video that was successfully generated in the past, a failed video that failed to be generated in the past, a video being generated that is currently being generated, and a video waiting to be generated that is waiting to be generated; the device further includes an asset video display module. The asset video display module is configured to display at least one asset video including a video of the target video type in response to a triggering operation on the video display control.

[0037] In the above solution, the asset video includes a failed video. The failed video is associated with a regeneration control. The regeneration control is used to regenerate the failed video; the device further includes a regeneration module. The regeneration module is configured to regenerate the failed video in response to a triggering operation on the regeneration control associated with the target failed video; when the regeneration is successful, switch the failed video in the at least one asset video to the regenerated video; when the regeneration fails, display a regeneration failure prompt message.

[0038] In the above solution, the at least one asset video includes a video waiting to be generated. The device further includes a waiting prompt module. The waiting prompt module is configured to display a first waiting prompt message in the associated area of the video waiting to be generated. The first waiting prompt message is used to indicate that the corresponding video is waiting to be generated; in response to a triggering operation on the video waiting to be generated, display a waiting cancellation control. The waiting cancellation control is used to terminate the waiting process of the video waiting to be generated.

[0039] In the above solution, the video generation interface further includes a style setting control, and the style setting control is used to set the video style of the generated video; the device further includes a style setting module, and the style setting module is configured to display at least one video style in response to a trigger operation on the style setting control; in response to a selection operation on a target video style among the at least one video style, control the target video style to be in a selected state; the fourth display module is further configured to display a video of a target video type with the target video style generated based on the target content in response to a video generation instruction.

[0040] In the above solution, the device further includes a style identifier display module, and the style identifier display module is configured to display an identifier of the target video style on the style setting control in response to a determination instruction for the target video style in a selected state, and the identifier is associated with a style cancellation control; wherein, the identifier of the target video style is used to identify that the generated video of the target video type has the target video style; in response to a trigger operation on the style cancellation control, cancel the display of the identifier of the target video style on the style setting control; the fourth display module is further configured to display a video of a target video type with a default style generated based on the target content in response to a video generation instruction.

[0041] In the above solution, the fourth display module is further configured to display a second waiting prompt message in response to a video generation instruction, and the second waiting prompt message is used to indicate the duration to wait until the generation of the video of the target video type starts; when the waiting prompt message indicates that the duration to wait until the generation of the video of the target video type starts is zero, cancel the display of the second waiting prompt message and display the video of the target video type generated based on the target content.

[0042] In the above solution, the device further includes an interaction module, which is configured to display at least one interaction control for the generated video, and the interaction control is one of the following controls: a local editing control for locally editing the generated video, a canvas expansion control for editing the size of the generated video, a motion brush control for changing static objects in the generated video into dynamic objects, a shooting parameter setting control for setting the shooting parameters of the generated video, a background removal control for removing the background of the generated video, an object erasing control for erasing target objects in the generated video, a scene detection control for splitting the video into different video segments based on different scenes in the generated video, a depth of field setting control for setting the depth of field effect for the generated video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the generated video, an action sequence generation control for setting an action sequence for an object in the generated video, and a second new video generation control for generating a new video based on the audio data in the generated video; in response to a trigger operation on a target interaction control among at least one of the interaction controls, an interaction operation indicated by the target interaction control is performed on the video of the generated target video type.

[0043] An embodiment of the present application provides an electronic device, including:

[0044] A memory for storing computer-executable instructions;

[0045] A processor, when executing the computer-executable instructions stored in the memory, implements the video generation method provided by the embodiment of the present application.

[0046] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to implement the video generation method provided by the embodiment of the present application when executed.

[0047] An embodiment of the present application provides a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the video generation method provided by the embodiment of the present application.

[0048] The embodiment of the present application has the following beneficial effects:

[0049] First, at least one video generation control for generating videos of different video types is displayed. Then, based on the at least one video generation control, in response to the target video type among at least two video types being in a selected state, a content input area corresponding to the target video type is displayed, and based on the content input area, target content is input. Thus, based on the input target content, a video of the target video type is generated. In this way, not only is a video of the target video type corresponding to the target content generated based on the target content input by the user, so that the video content of the target video meets the user's needs, but also, based on the video generation control, a video of the target video type is generated, so that the type of the video also meets the user's needs. Compared with the solution where the user still needs to edit the video after it is generated, the present application reduces the editing operations that the user needs to perform on the generated video, not only increasing the human-computer interaction efficiency, but also improving the video generation efficiency. Description of the Drawings

[0050] Figure 1 is a schematic diagram of the architecture of the video generation system 100 provided by an embodiment of the present application;

[0051] Figure 2 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0052] Figure 3 is a schematic flowchart of the video generation method provided by an embodiment of the present application;

[0053] Figure 4 is a schematic diagram of the display of the content input area provided by an embodiment of the present application;

[0054] Figure 5 is a schematic diagram of the guiding information provided by an embodiment of the present application;

[0055] Figure 6 is a schematic diagram of the playback interface of other videos provided by an embodiment of the present application;

[0056] Figure 7 is a schematic diagram of the permission prompt information provided by an embodiment of the present application;

[0057] Figure 8 is a schematic diagram of at least one video template provided by an embodiment of the present application;

[0058] Figure 9 is a schematic diagram of the details interface including the details information of the target video template provided by an embodiment of the present application;

[0059] Figure 10 is a schematic diagram of the determination control provided by an embodiment of the present application;

[0060] Figure 11 is a schematic diagram of the display of the parameter setting area provided by an embodiment of the present application;

[0061] Figure 12 It is a schematic diagram of the storyboard disassembly interface provided by an embodiment of the present application;

[0062] Figure 13 It is a schematic diagram of the storyboard generation interface provided by an embodiment of the present application;

[0063] Figure 14 It is a schematic diagram of the function introduction information provided by an embodiment of the present application;

[0064] Figure 15 It is a schematic diagram of multiple candidate storyboards provided by an embodiment of the present application;

[0065] Figure 16 It is a schematic diagram of the video display interface of the long video provided by an embodiment of the present application;

[0066] Figure 17 It is a schematic diagram of the video display interface of the short video provided by an embodiment of the present application;

[0067] Figure 18 It is a schematic diagram of the first progress prompt information provided by an embodiment of the present application;

[0068] Figure 19 It is a schematic diagram of the second waiting prompt information provided by an embodiment of the present application;

[0069] Figure 20 It is a schematic diagram of the asset video display interface provided by an embodiment of the present application;

[0070] Figure 21 It is a schematic diagram of the deletion prompt information provided by an embodiment of the present application;

[0071] Figure 22 It is a schematic diagram of the generation failure prompt information, the re-generation control, and the backtracking control provided by an embodiment of the present application;

[0072] Figure 23 It is a schematic diagram of at least one video style provided by an embodiment of the present application;

[0073] Figure 24 It is a schematic diagram of the identifier of the target video style provided by an embodiment of the present application;

[0074] Figure 25 It is a schematic diagram of the operation control provided by an embodiment of the present application;

[0075] Figure 26 It is a schematic diagram of interacting with the generated video based on the local editing control provided by an embodiment of the present application;

[0076] Figure 27It is a schematic diagram of interacting with the generated video based on the canvas extension control provided by the embodiments of the present application;

[0077] Figure 28 It is a schematic structural diagram of the video generation model provided by the embodiments of the present application;

[0078] Figure 29 It is a schematic diagram of the sampling process provided by the embodiments of the present application;

[0079] Figure 30 It is a schematic diagram of the processing process of the skeleton-guided dance video model provided by the embodiments of the present application. Detailed implementation manners

[0080] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0081] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0082] In the following description, the terms "first\second\third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0083] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0084] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0085] 1) Responsive to, which is used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations to be executed can be real-time or can have a set delay; without special instructions, there is no limit on the execution order of the multiple operations to be executed.

[0086] 2) The client, also known as the user side, refers to a program that provides local services corresponding to the server. Except for some applications that can only run locally, it is generally installed on the terminal and needs to cooperate with the server to run. That is, there needs to be a corresponding server and service program in the network to provide corresponding services. In this way, a specific communication connection needs to be established between the client and the server side to ensure the normal operation of the application program. For example, a virtual scene client (such as a game client).

[0087] 3) Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0088] 4) Storyboarding refers to various video media such as movies, animations, TV dramas, advertisements, and music videos. Before actual shooting or drawing, it uses diagrams to illustrate the composition of the video, decomposes continuous pictures in units of one camera movement, and marks the camera movement method, time length, dialogue, special effects, etc.

[0089] 5) Camera movement, also known as a moving shot, is the shooting carried out by moving the camera position, or changing the camera optical axis, or varying the camera focal length within one shot. The picture taken through this shooting method is called a moving picture. Such as: push shots, pull shots, pan shots, tracking shots, crane shots, and comprehensive moving shots formed by push, pull, pan, track, follow, lift, and comprehensive camera movement.

[0090] See Figure 1 , Figure 1 is a schematic diagram of the architecture of the video generation system 100 provided by the embodiments of the present application. The terminal (exemplarily shows the terminal 400), and the terminal 400 is connected to the server 200 through the network 300. Among them, the network 300 can be a wide area network or a local area network, or a combination of the two, and uses wireless or wired links to achieve data transmission.

[0091] Among them, the server 200 is used to send interface data corresponding to the video generation interface including at least one video generation control to the terminal 400;

[0092] The terminal 400 is configured to receive interface data corresponding to a video generation interface including at least one video generation control; based on the interface data, display the video generation interface, the video generation interface including at least one video generation control, the at least one video generation control being configured to generate videos of at least two video types; based on the at least one video generation control, in response to a target video type among the at least two video types being in a selected state, display a content input area corresponding to the target video type; in response to a content input operation in the content input area, display the input target content; and in response to a video generation instruction, display a video of the target video type generated based on the target content.

[0093] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs, Content Deliver Network), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, a smart voice interaction device, a smart home appliance, a virtual reality device, a vehicle-mounted terminal, an aircraft, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device, a smart speaker, and a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0094] Next, the electronic device for implementing the video generation method provided in the embodiments of the present application will be described. Refer to Figure 2 , Figure 2 FIG. is a schematic structural diagram of the electronic device provided in the embodiments of the present application. The electronic device may be a server or a terminal. Taking the terminal shown in Figure 1 as an example, Figure 2 the electronic device shown in FIG. includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. The bus system 440 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.

[0095] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0096] The user interface 430 includes one or more output devices 431 that enable display of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0097] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0098] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0099] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0100] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0101] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;

[0102] A presentation module 453 for enabling information to be displayed via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.) (e.g., a user interface for operating a peripheral device and displaying content and information);

[0103] An input processing module 454 for detecting and translating one or more user inputs or interactions from one of one or more input devices 432.

[0104] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2 A video generation device 455 stored in the memory 450 is shown, which may be software in the form of a program, a plug-in, etc., and includes the following software modules: a first display module 4551, a second display module 4552, a third display module 4553, and a fourth display module 4554. These modules are logical, so they can be combined arbitrarily or further analyzed according to the functions implemented. The functions of each module will be described below.

[0105] In other embodiments, the device provided by the embodiments of the present application may be implemented in hardware. As an example, the video generation device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the video generation method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0106] In some embodiments, a terminal or a server may implement the video generation method provided by the embodiments of the present application by running a computer program. For example, the computer program may be a native program or a software module in an operating system; it may be a native application (APP), that is, a local client, that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP or a web browser APP; it may also be a small program, that is, a program that only needs to be downloaded to a browser environment to run; it may also be a small program that can be embedded in any APP. In short, the above computer program may be any form of client, module, or plug-in.

[0107] Based on the above description of the video generation system and the electronic device provided in the embodiments of the present application, the following describes the video generation method provided in the embodiments of the present application. In actual implementation, the video generation method provided in the embodiments of the present application can be implemented independently by a terminal or a server, or jointly implemented by a terminal and a server, taking the terminal 400 in Figure 1 as an example to separately execute the video generation method provided in the embodiments of the present application for illustration. Refer to Figure 3 , Figure 3 which is a schematic flowchart of the video generation method provided in the embodiments of the present application. Next, the steps shown in Figure 3 will be described.

[0108] Step 101, the terminal displays a video generation interface, and the video generation interface includes at least one video generation control, and the at least one video generation control is used to generate videos of at least two video types.

[0109] In actual implementation, the terminal is provided with a client supporting video generation. When the user opens the client on the terminal and the terminal runs the client, the terminal can, based on the client, display a video generation interface; wherein, the video generation interface includes at least one video generation control, and the at least one video generation control is used to generate videos of at least two video types; wherein, the video types include long videos and short videos.

[0110] In some embodiments, the number of video generation controls can be multiple, and different video generation controls correspond to different video types. Thus, the process of displaying the video generation interface can be, in response to an open instruction for the video generation interface, displaying a video generation interface including multiple video generation controls; thus, in response to the video generation control corresponding to the target video type in the video generation interface being in a selected state, determining that the target video type is in a selected state.

[0111] In other embodiments, the number of video generation controls is one. Thus, the process of displaying the video generation interface can be, in response to an open instruction for the video generation interface, displaying a video generation interface including one video generation control, and in response to a trigger operation on the video generation control, displaying at least two video type selection controls, where different video type selection controls are used to select different video types, and in response to a selection operation on the target video type selection control among the at least two video type selection controls, displaying the video generation interface of the target video type indicated by the video type selection control.

[0112] Step 102, based on the at least one video generation control, in response to the target video type among the at least two video types being in a selected state, display a content input area corresponding to the target video type.

[0113] It should be noted that when the video generation interface is displayed and there is one video generation control, after at least two video type selection controls are displayed in response to a trigger operation on the video generation control, which video type selection control the user triggers determines that the corresponding video type is in the selected state;

[0114] When the video generation interface is displayed and there are multiple video generation controls, one of the video generation controls can be automatically selected, or it can be based on the user's selection operation to control one of the video generation controls to be in the selected state, that is, the target video type is in the selected state; for example, first display a blank video generation interface and at least two video generation controls in the unselected state, and then in response to the selection operation on the target video generation control among the at least two video generation controls, control the target video generation control among the at least two video generation controls to be in the selected state, so as to display the content input area; or preset one of the at least two video generation controls to be in the selected state, so that when the video generation interface is displayed, first display the video generation interface including the target video generation control in the selected state, so as to display the content input area;

[0115] Based on this, in the process of determining that the target video type is in the selected state in response to the selected state of the video generation control corresponding to the target video type in the video generation interface, the target video generation control in the selected state can be determined based on the user's selection operation after the blank video generation interface is displayed, or can be preset, that is, directly determined when the video generation interface is displayed, or can also be determined based on the user's selection operation on another video generation control that is not in the selected state after the preset selected video generation control is displayed. In this regard, the embodiments of the present application do not make any limitations.

[0116] Exemplarily, refer to Figure 4 , Figure 4 is a schematic diagram of the display content input area provided by the embodiments of the present application. Based on Figure 4 a in, when the target video generation control is a short video generation control, in response to the selected state of the short video generation control indicated by 401, display the content input area indicated by 402; or, based on Figure 4 b in, when the target video generation control is a long video generation control, in response to the selected state of the long video generation control indicated by 403, display the content input area indicated by 404.

[0117] In some embodiments, based on at least one video generation control, in response to the target video type among at least two video types being in a selected state, while displaying the content input area corresponding to the target video type, guiding information will also be displayed in the form of a floating layer above the content input area. The guiding information is used to guide the input operation based on the content input area. Exemplarily, refer to Figure 5 , Figure 5 which is a schematic diagram of the guiding information provided by an embodiment of the present application. Based on Figure 5 , when the target video generation control is a short video generation control, in response to the short video generation control indicated by 501 being in a selected state, while displaying the content input area indicated by 502, the guiding information indicated by 503 will be displayed in the form of a floating layer.

[0118] In some embodiments, the user can directly perform a triggering operation on the function control for displaying the video generation interface when running the client to display the video generation interface; in other embodiments, it can also be based on a link shared by other users to display the video generation interface. Specifically, before displaying the video generation interface, it can also, in response to the display instruction triggered by the shared link, if the current account has the playback permission for other videos, display the playback interface of the other videos, and display the detailed information of the other videos and the first new video generation control on the playback interface; thus, the process of displaying the video generation interface can be, in response to the triggering operation on the first new video generation control, to display the video generation interface.

[0119] It should be noted that the first new video generation control is used to generate a new video, that is, input the target content in the content input area and generate a video of the target video type based on the input target content; the detailed information includes at least one of the duration, size, content description, prompt information on whether it includes audio, resolution, key image frame of the other video, where the key image frame is used to indicate the cover of the other video, or the picture input when generating the other video, etc.; the content description is used to indicate the text input when generating the other video, or the text content obtained by analyzing the video content of the other video after generating the other video; and whether it includes audio can indicate whether audio was input when generating the other video, or can indicate whether audio was generated when generating the other video. In this regard, the embodiments of the present application do not make any limitations.

[0120] Exemplarily, refer to Figure 6 , Figure 6 which is a schematic diagram of the playback interface of the other video provided by an embodiment of the present application. Based on Figure 6 a of Figure 6As shown in a of , where the detailed information of other videos is shown in the dashed box 601. In response to a trigger operation for generating a control for the first new video indicated by 602, a video generation interface as shown in Figure 4 b of is displayed; Based on Figure 6 b of , when the other video is a short video, the displayed playback interface is as shown in Figure 6 b of , where the detailed information of other videos is shown in the dashed box 603. In response to a trigger operation for generating a control for the first new video indicated by 604, a video generation interface as shown in Figure 4 a of is displayed.

[0121] In actual implementation, when the current account does not have the playback permission for other videos, a permission prompt message is displayed; Among them, the permission prompt message is used to prompt the target object that it does not have the playback permission for the target video.

[0122] It should be noted that whether the current account has the permission can be set by the account that generates other videos, or determined by information such as the level of the current account. In this regard, the embodiments of the present application do not make limitations.

[0123] Exemplarily, referring to Figure 7 , Figure 7 is a schematic diagram of the permission prompt message provided by the embodiments of the present application. Based on Figure 7 , when the current account does not have the playback permission for other videos, a permission prompt message as shown in Figure 7 is displayed.

[0124] In some embodiments, after displaying the video generation interface, at least one video template can also be displayed; In response to a selection operation for a target video template in the at least one video template, the template content description of the target video template and the corresponding first editing control are displayed; In response to a trigger operation for the first editing control, the trigger operation is determined as an input operation for the content input area; Thus, the process of displaying the input target content in response to the content input operation in the content input area can be that, in response to a trigger operation for the first editing control, the trigger operation is determined as a content input operation, in the content input area, the template content description of the target video template is displayed, and the displayed template content description is determined as the input target content.

[0125] It should be noted that the first editing control is used to generate a video associated with the target video template again based on the content of the target video template; when the content input area is displayed, at least one video template will also be displayed, and the video template here can be uploaded by other users or by the current user; the selection operation for the target video template among at least one video template is used to indicate that the target video template is in a selected state, rather than clicking on the target video template. For example, when the cursor controlled by the user moves above the target video template, it can be considered that the selection operation for the target video template is triggered.

[0126] Then, in response to the selection operation for the target video template among at least one video template, on the target video template, the template content description of the target video template and the corresponding first editing control are displayed; when a click operation is performed on the first editing control, the template content description of the target video template is filled back into the content input area, that is, in response to the triggering operation for the first editing control, a video generation interface including the content input area carrying the template content description of the target video template is displayed.

[0127] Exemplarily, refer to Figure 8 , Figure 8 which is a schematic diagram of at least one video template provided by an embodiment of the present application. Based on Figure 8 a of, in response to the selection operation for the target video template indicated by 801, the template content description of the target video template indicated by the dashed box 802 and the corresponding first editing control indicated by 803 are displayed; then in response to the click operation on the first editing control, when the target video template is a short video, the interface shown in Figure 8 b of is displayed, where Figure 8 in b of, what is indicated by 804 is the content input area carrying the template content description of the target video template;

[0128] When the target video template is a long video, the last interface when generating a long video is displayed, for example, the video generation interface shown in Figure 8 c of, where Figure 8 in c of, what is indicated by 805 is the content input area carrying the template content description of the target video template; or it is the storyboard disassembly interface mentioned later (such as directly generating the target video based on the storyboard disassembly interface, that is, a video of the target video type), or it can also be the storyboard generation interface mentioned later (such as directly generating the target video based on the storyboard generation interface, or generating the target video based on the storyboard generation interface after displaying the storyboard generation interface based on the storyboard disassembly interface). In this regard, the embodiments of the present application do not make any limitations.

[0129] In actual implementation, after displaying at least one video template, in response to a trigger operation on a target video template among the at least one video template, a details interface including details information of the target video template may further be displayed; wherein, the details interface further includes at least one of a second editing control and a first new video generation control. The second editing control is used to regenerate a video of the target video type based on the template content description. The first new video generation control is used to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

[0130] It should be noted that the second editing control is similar to the above-mentioned first editing control; meanwhile, the trigger operation on the target video template among the at least one video template may be a click operation on the target video template; the details information includes at least one of the duration, size, content description, prompt information on whether audio is included, key image frame, and resolution of the target video template; wherein, the key image frame is used to indicate the cover of the target video template, or is a picture input when generating the target video template, etc.; the content description is used to indicate the text input when generating the target video template, or is used to indicate the text content obtained by analyzing the video content of the target video template after generating the target video template; and whether audio is included may indicate whether audio was input when generating the target video template, or may indicate whether audio was generated when generating the target video template. In this regard, the embodiments of the present application do not make any limitations.

[0131] Exemplarily, continue to refer to Figure 8 and refer to Figure 9 , Figure 9 is a schematic diagram of the details interface including details information of the target video template provided by the embodiments of the present application. Based on Figure 8 a, when the target video template is a short video, in response to a trigger operation on the target video template indicated by 801, the details interface shown in Figure 9 a is displayed. Among them, the details information of the target video template is indicated by the dashed box 901, the second editing control is indicated by 902, and the first new video generation control is indicated by 903. Thus, in response to a trigger operation on the second editing control, the interface shown in Figure 8 b is displayed, or in response to a trigger operation on the new video control, the video generation interface shown in Figure 4 a is displayed;

[0132] Based on Figure 8 a, when the target video template is a long video, in response to a trigger operation on the target video template indicated by 801, the details interface shown in Figure 9The details interface shown in Figure b, where the details information of the target video template is indicated by the dashed box 904, the second editing control is indicated by 905, and the first new video generation control is indicated by 906, so as to display in response to a trigger operation on the new video control Figure 4 the video generation interface shown in Figure b; or, in response to a trigger operation on the second editing control, display the last interface when generating a long video, such as Figure 8 the video generation interface shown in Figure c, where Figure 8 the content input area indicated by 805 in Figure c carries the template content description of the target video template; or it is the storyboard disassembly interface mentioned later (such as directly generating the target video based on the storyboard disassembly interface), or it can also be the storyboard generation interface mentioned later (such as directly generating the target video based on the storyboard generation interface, or after displaying the storyboard generation interface based on the storyboard disassembly interface, generating the target video based on the storyboard generation interface). In this regard, the embodiments of the present application do not make any limitations.

[0133] Step 103, in response to a content input operation in the content input area, display the input target content.

[0134] It should be noted that the content input operation in the content input area can indicate the input of at least one of a picture, text, and audio, so that the target content can be at least one of the corresponding picture, text, and audio; among them, at least one input content editing control can also be displayed in the associated area of the content input area, where the at least one input content editing control is for the input pictures and audio, including a preview control, a re-upload control, and a delete control. Specifically, when a picture or audio input is displayed in response to an input operation on the content input area, in response to a trigger operation on the preview control, preview the input picture and audio, or in response to a re-upload operation triggered by the re-upload control, display the re-entered picture or audio, or in response to a trigger operation on the delete control, delete the input picture or audio. Exemplarily, continue to refer to Figure 4 , based on Figure 4 Figure a, the three input content editing controls indicated by the dashed box 405, where the first control is the preview control, the second control is the re-upload control, and the third control is the delete control. Correspondingly, based on Figure 4 Figure b, the three input content editing controls indicated by the dashed box 406, where the first control is the preview control, the second control is the re-upload control, and the third control is the delete control.

[0135] In some embodiments, in response to a target video type among at least two video types being in a selected state, a determination control in an inactive state may be displayed. The determination control is used to confirm target content input in a content input area. Thus, after a target content input operation in the content input area and the display of the input target content, the determination control may be switched from the inactive state to an active state. In response to a trigger operation on the determination control in the active state, a video generation instruction is received.

[0136] It should be noted that the inactive state is used to indicate that the determination control is in a dim state and cannot be triggered, and the active state is used to indicate that the determination control is in a normal state and can be triggered. Exemplarily, referring to Figure 10 , Figure 10 is a schematic diagram of the determination control provided in an embodiment of the present application. Based on Figure 10 a of, when the target video generation control is a short video generation control, in response to the short video type being in a selected state, a determination control in the inactive state indicated by 1001 in a of Figure 10 is displayed. After a target content input operation in the content input area and the display of the input target content, as shown by 1002 in b of Figure 10 , the determination control is switched from the inactive state to the active state. Thus, in response to a trigger operation on the determination control in the active state, a video generation instruction is received.

[0137] Correspondingly, based on Figure 10 c of, when the target video generation control is a long video generation control, in response to the long video type being in a selected state, a determination control in the inactive state indicated by 1003 in c of Figure 10 is displayed. After a target content input operation in the content input area and the display of the input target content, as shown by 1004 in d of Figure 10 , the determination control is switched from the inactive state to the active state. Thus, in response to a trigger operation on the determination control in the active state, a video generation instruction is received.

[0138] In some embodiments, while the display content input area is being shown, a parameter setting area is also shown in the video generation interface. The parameter setting area is used to set the parameters of the generated video. Based on the parameter setting area, in response to a parameter setting operation, the video parameters set for the video of the target video type are shown. Thus, in response to a video generation instruction, based on the target content, a video of the target video type with the video parameters is shown. For example, when the parameter setting of the target video is completed, in response to a trigger operation on the determination control in the active state, a video generation instruction is received. Thus, in the subsequent process, in response to the video generation instruction, a video of the target video type with the target parameters is generated. Among them, the parameters of the video include the duration, quality, ratio, resolution, etc. of the video.

[0139] Exemplarily, referring to Figure 11 , Figure 11 is a schematic diagram showing the parameter setting area provided by an embodiment of the present application. Based on Figure 11 a in, when the target video generation control is a short video generation control, the parameter setting area indicated by the dashed box 1101 is shown; or, based on Figure 11 b in, when the target video generation control is a long video generation control, the parameter setting area indicated by the dashed box 1102 is shown. Thus, when the parameter setting of the target video is completed, in response to a trigger operation on the determination control in the active state, a video generation instruction is received, and then a corresponding short video or long video is generated.

[0140] It should be noted that before parameter setting, default parameters are shown. For example, when the target video generation control is a short video generation control, the default parameters can be that the special effect is default none, the video quality is default 896*512, the ratio is default 7:4, the video is default 4s, and the audio is default none. When the target video generation control is a long video generation control, the default parameters can be that the video duration is default 6s, the video quality is default 896*512, the ratio is default 7:4, and the audio is none. When no parameter setting is performed and the trigger operation on the determination control in the active state is directly executed, the parameters of the video of the target video type generated in the subsequent process are the default video parameters. If parameter setting is performed, the parameters of the video of the target video type generated in the subsequent process are the target parameters.

[0141] In actual implementation, the video generation interface includes parameter setting controls, which are used to set the parameters of the generated target video. That is, the parameter setting area can be displayed based on the triggering of the parameter setting controls. Specifically, in response to the triggering operation on the parameter setting controls, the parameter setting area is displayed; based on the parameter setting area, in response to the parameter setting operation on the target video, the target parameters set for the target video are displayed; thus, in response to the video generation instruction, the process of displaying the target video generated based on the target content can be that when the setting of the parameters of the target video is completed, in response to the video generation instruction, the target video generated based on the target content is displayed; thus, in the subsequent process, in response to the video generation instruction, a target video with target parameters is generated. Exemplarily, continue to refer to Figure 4 , based on Figure 4 in a, the control indicated by 407 is the parameter setting control. Correspondingly, based on Figure 4 in b, the control indicated by 408 is the parameter setting control.

[0142] It should be noted that as described above, the video types include long videos and short videos, and the number of video generation controls can be multiple. For different video types, the process of receiving the video generation instruction is also different. Next, for the video generation controls of different video types, the process of receiving the video generation instruction will be described.

[0143] For the case where the target video generation control includes a short video generation control and the target video type includes the short video type.

[0144] In actual implementation, when the target video is a short video, directly based on the process described above, such as when the parameter setting of the target video is completed, in response to the triggering operation on the active confirmation control, or directly in response to the triggering operation on the active confirmation control, the video generation instruction can be received.

[0145] For the case where the target video generation control includes a long video generation control and the target video type includes the long video type.

[0146] In some embodiments, when the target video is a long video, the long video includes multiple storyboards. After displaying the input target content in response to the content input operation in the content input area, it is also possible to display the storyboard disassembly interface in response to the confirmation instruction for the input target content, and multiple storyboard texts determined based on the target content are displayed on the storyboard disassembly interface. Each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; thus, in the subsequent process, in response to the video generation instruction triggered by the multiple storyboard texts, a video of the target video type including the storyboards corresponding to the respective storyboard texts is displayed.

[0147] It should be noted that one shot here is used to indicate a video clip, and multiple video clips are combined to form the final target video; after the input target content is displayed, by parsing the target content, the corresponding text content of the target content is obtained, and then the text content is split to obtain multiple text paragraphs, so as to determine multiple shot texts, and then display them in the shot decomposition interface.

[0148] At the same time, the determination instruction for multiple shot texts can be triggered by the first determination control. Specifically, the shot decomposition interface further includes a first determination control, so that after the multiple shot texts determined based on the target content are displayed in the shot decomposition interface, a determination instruction for the multiple shot texts can also be received in response to the trigger operation for the first determination control.

[0149] Exemplarily, refer to Figure 12 , Figure 12 is a schematic diagram of the shot decomposition interface provided by an embodiment of the present application. Based on Figure 12 , what the dashed box 1201 indicates are four shot texts, and each shot text is used to describe the content of the corresponding shot. What 1202 indicates is the first determination control, so that in response to the trigger operation for the first determination control, a determination instruction for the multiple shot texts is received.

[0150] In actual implementation, a shot text quantity adjustment control can also be displayed. The shot text quantity adjustment control is used to adjust the quantity of shot texts, such as adding new shot texts or deleting original shot texts; in response to the shot text quantity adjustment operation triggered based on the shot text quantity adjustment control, such as the shot text addition operation or the shot text deletion operation, in the shot decomposition interface, the shot texts with adjusted quantities are displayed, such as the newly added shot texts or the shot texts after deletion.

[0151] As an example, taking adding new text as an example, specifically, in response to the trigger operation for the shot text quantity adjustment control, a text input box is displayed in the associated area of multiple shot texts; in response to the input operation for the text input box, the input text is displayed; in response to the determination instruction for the input text, the newly added shot text is displayed, and the new shot text is associated with the input text. Thus, in response to the determination operation triggered based on the multiple shot texts including the new shot text, a video generation instruction is received.

[0152] It should be noted that the associated area of multiple shot texts can be, for example, the area below the multiple shot texts, and the determination instruction for the input text can be triggered by a determination control. In this regard, the embodiments of the present application do not make any limitations.

[0153] In actual implementation, parameters of the shots corresponding to the respective storyboard texts can also be set. Specifically, each storyboard text is associated with a parameter setting control, and the parameter setting control is used to set the shot parameters of the shot corresponding to the respective storyboard text. Thus, after the storyboard disassembly interface is displayed, it is also possible to, in response to a parameter setting operation triggered by the parameter setting control associated with the target storyboard text, display the target shot parameters set for the target storyboard text; in response to a determination instruction for the target shot parameters, update the shot parameters of the shot corresponding to the target storyboard text to the target shot parameters. Thus, in the subsequent process, the generated target video, that is, the video of the target video type, includes the shot parameters of the target shot as the target shot parameters.

[0154] It should be noted that the target storyboard text here can be some of the multiple storyboard texts or all of the storyboard texts. In this regard, the embodiments of the present application do not make any limitations. At the same time, after the triggering operation for the parameter setting control is executed, a parameter setting area is displayed, and thus the parameter setting operation is implemented based on the parameter setting area. The parameter setting area is as described above, and in this regard, the embodiments of the present application will not elaborate.

[0155] Exemplarily, continue to refer to Figure 12 , the control indicated in the dashed box 1203 is the parameter setting control. Thus, in response to the triggering operation for the parameter setting control corresponding to one of the storyboard texts, a parameter setting area is displayed, and then in response to the parameter setting operation triggered based on the parameter setting area, the target parameters set for the target storyboard text are displayed. In this way, the parameters of the shots corresponding to each storyboard text are set in sequence, so that the finally generated target video conforms to the set parameters.

[0156] In actual implementation, the arrangement order of the storyboard texts can also be changed. Specifically, there is an arrangement order for the multiple storyboard texts, and the content of the video formed by the shots corresponding to the multiple storyboard texts with different arrangement orders is different. Each storyboard text is associated with a sorting adjustment control. Thus, after the storyboard disassembly interface is displayed, it is also possible to, in response to a sorting adjustment operation for the multiple storyboard texts triggered by the sorting adjustment control, change the multiple storyboard texts from the current arrangement order to the target arrangement order. Thus, in the subsequent process, based on the multiple storyboard texts, in response to a video generation instruction, a video of the target video type including multiple shots generated based on the target arrangement order is displayed.

[0157] It should be noted that as described above, the target video is composed of multiple shots, and there is a corresponding relationship between the shots and the storyboard texts. Therefore, the arrangement order of the multiple storyboard texts also indicates the combination order of the multiple shots. Thus, different arrangement orders of the multiple shots result in different combination orders of the multiple shots, thereby causing the content of the target video to be different.

[0158] It should be noted that the sorting of multiple storyboard texts can be achieved by dragging the corresponding storyboard texts by dragging the sorting adjustment control, or by performing a trigger operation on the sorting adjustment control to display an input box in the associated area of each storyboard text. Then, in response to the input operation on the input box, the serial numbers entered for each storyboard text are displayed, so as to change the arrangement order of the multiple storyboard texts from the current order to the target order.

[0159] Exemplarily, continue to refer to Figure 12 , the control indicated in the dashed box 1204 is a sorting adjustment control for sorting multiple storyboard texts by dragging the corresponding storyboard texts by dragging the sorting adjustment control. The order of the storyboards corresponding to the multiple storyboard texts from top to bottom forms a target video. Thus, in response to the sorting adjustment operation on the multiple storyboard texts triggered by the sorting adjustment control, the arrangement order of the multiple storyboard texts is changed, and then based on the changed order, a video generation instruction is received.

[0160] In actual implementation, it is also possible to select some storyboard texts to generate a target video. Specifically, in response to the selection operation on one or more target storyboard texts among the multiple storyboard texts, the target storyboard texts are controlled to be in a selected state; thus, in the subsequent process, based on the target storyboard texts in the selected state, in response to the video generation instruction, a video of the target video type including the target storyboard texts is displayed.

[0161] In actual implementation, it is also possible to modify the storyboard texts. Specifically, in response to the text modification operation on the target storyboard text among the multiple storyboard texts, the modified target storyboard text is displayed. Thus, in the subsequent process, based on the modified target storyboard text, in response to the video generation instruction, a video of the target video type including the modified target storyboard text is displayed.

[0162] In actual implementation, as described above, a first determination control is also displayed in the storyboard disassembly interface. When the first determination control is triggered, a video generation instruction can be received, that is, a target video is directly generated; or a storyboard generation interface can also be displayed. Specifically, in the process of receiving the video generation instruction in response to the determination instruction for the multiple storyboard texts, it can be that in response to the trigger operation on the first determination control, the storyboard generation interface is displayed; in the storyboard generation interface, the storyboards corresponding to each storyboard text are displayed, and a second determination control is also displayed; in response to the trigger operation on the second determination control, a video generation instruction is received.

[0163] Exemplarily, continue to refer to Figure 12 And refer to Figure 13 , Figure 13 is a schematic diagram of the storyboard generation interface provided by the embodiment of the present application. As Figure 12As shown, the one indicated by 1202 is the first determination control. Thus, in response to a determination instruction for multiple storyboard texts triggered based on the first determination control, a storyboard generation interface as indicated by Figure 13 is displayed. In the storyboard generation interface, the storyboards corresponding to each storyboard text generated as indicated by the dashed box 1301 are displayed, and a second determination control as indicated by 1302 is displayed. Thus, in response to a trigger operation for the second determination control, a video generation instruction is received.

[0164] It should be noted that the storyboard generation interface includes multiple storyboard texts and multiple storyboards, with each storyboard text corresponding to one storyboard, and the storyboard is generated based on the corresponding storyboard text. Thus, based on the multiple storyboard texts and multiple storyboards, in response to the video generation instruction, a video of the target video type including multiple storyboards is displayed.

[0165] In actual implementation, each storyboard is associated with a parameter setting control, and the parameter setting control is used to set the storyboard parameters of the corresponding storyboard. After the multiple storyboard texts and multiple storyboards are displayed in the storyboard generation interface, it is also possible to, in response to a parameter setting operation triggered based on the parameter setting control associated with the target storyboard, display the target storyboard parameters set for the target storyboard; in response to a determination instruction for the target storyboard parameters, update the storyboard parameters of the target storyboard to the target storyboard parameters. Thus, in the subsequent process, the storyboard parameters of the target storyboard included in the generated video of the target video type are the target storyboard parameters.

[0166] It should be noted that the target storyboard here can refer to some of the multiple storyboards or all of the storyboards. In this regard, the embodiments of the present application do not make any limitations. At the same time, when a trigger operation for the parameter setting control is executed, a parameter setting area is displayed, and thus the parameter setting operation is implemented based on the parameter setting area. The parameter setting area is as described above, and in this regard, the embodiments of the present application will not elaborate.

[0167] Exemplarily, continue to refer to Figure 13 , the one indicated by 1303 is the parameter setting control for one of the storyboards, and the parameter setting controls for other storyboards are similar. Thus, in response to a trigger operation for the parameter setting control corresponding to one of the storyboards, a parameter setting area is displayed, and then in response to a parameter setting operation triggered based on the parameter setting area, the target storyboard parameters set for the target storyboard are displayed. In this way, the storyboard parameters of each storyboard are set in sequence, so that the finally generated target video, that is, the video of the target video type, conforms to the set storyboard parameters.

[0168] In actual implementation, it is also possible to select some of the storyboards to generate the target video. Specifically, after multiple storyboard texts and multiple storyboards are displayed on the storyboard generation interface, it is also possible to control the first target storyboard to be in a selected state in response to a selection operation on the first target storyboard among the multiple storyboards. Thus, in the subsequent process, based on the first target storyboard in the selected state and the storyboard text of the first target storyboard, in response to a video generation instruction, a video of the target video type including the first target storyboard is displayed.

[0169] It should be noted that the number of the first target storyboards here can be one, or multiple, such as two or three. In this regard, the embodiments of the present application do not make any limitations.

[0170] Exemplarily, continue to refer to Figure 13 , based on the control indicated by the dashed box 1304, in response to the selection operations on the 1st, 2nd, and 4th storyboards, control the 1st, 2nd, and 4th storyboards to be in the selected state, where "√" indicates that the corresponding storyboard is in the selected state, so as to generate a target video including the 1st, 2nd, and 4th storyboards.

[0171] In actual implementation, it is also possible to change the arrangement order of the storyboards. Specifically, there is an arrangement order for the multiple storyboards, and the videos formed by the storyboards corresponding to the multiple storyboards in different arrangement orders are different. Each storyboard is associated with a sorting adjustment control, and the sorting adjustment control is used to drag the corresponding storyboard to realize the sorting adjustment of the multiple storyboards. After multiple storyboard texts and multiple storyboards are displayed on the storyboard generation interface, it is also possible to change the sorting of the multiple storyboards from the current arrangement order to the target arrangement order in response to a sorting adjustment operation on the multiple storyboards triggered by the sorting adjustment control. Thus, based on the multiple storyboard texts and the multiple storyboards, in response to a video generation instruction, a video of the target video type including the multiple storyboards generated based on the target arrangement order is displayed.

[0172] It should be noted that as described above, the target video is composed of multiple storyboards. Therefore, the arrangement order of the multiple storyboards also indicates the combination order of the multiple storyboards. Thus, different arrangement orders of the multiple storyboards result in different combination orders of the multiple storyboards, thereby resulting in different contents of the target video.

[0173] It should be noted that the sorting of multiple storyboards can be achieved by dragging the corresponding storyboard text by dragging the sorting adjustment control, or by performing a triggering operation on the sorting adjustment control to display an input box in the associated area of each storyboard, so as to respond to the input operation on the input box and display the entered serial numbers for each storyboard, thereby changing the arrangement order of multiple storyboards from the current arrangement order to the target arrangement order; or by performing a triggering operation on the sorting adjustment control to display at least one sorting element, such as when the storyboard includes virtual objects, the duration of the virtual object in each storyboard, or the number of words spoken, etc., so as to respond to the selection operation on the target sorting element in at least one sorting element, and change the arrangement order of multiple storyboards from the current arrangement order to the target arrangement order determined based on the target sorting element. For example, based on the duration of the virtual object in each storyboard, multiple storyboards are sorted from long to short.

[0174] Exemplarily, continue to refer to Figure 13 , the control indicated in the dashed box 1305 is a sorting adjustment control for sorting multiple storyboards by dragging the corresponding storyboard text by dragging the sorting adjustment control. Multiple storyboards are combined in a top-to-bottom order to form a target video, so as to respond to the sorting adjustment operation on multiple storyboards triggered based on the sorting adjustment control and change the arrangement order of multiple storyboards.

[0175] In actual implementation, some of the multiple storyboards can also be updated. Specifically, a storyboard update control for updating the content of the storyboard is also displayed in the storyboard generation interface, and the storyboard update control can be one or more;

[0176] When there are multiple storyboard update controls, the storyboard update controls correspond to the storyboards one by one. After displaying multiple storyboard texts and multiple storyboards in the storyboard generation interface, it is also possible to, based on the storyboard update control, respond to the content update operation on the second target storyboard among multiple storyboards and update the second target storyboard to a first new storyboard, and the content of the first new storyboard is associated with the content of the second target storyboard.

[0177] It should be noted that the content of the first new storyboard is also generated based on the content of the storyboard text of the corresponding storyboard, so it is associated with the content of the second target storyboard.

[0178] Exemplarily, continue to refer to Figure 13 , 1306 indicates the storyboard update control of one of the storyboards, and the storyboard update controls of other storyboards are similar; thus, in response to the triggering operation on the storyboard update control corresponding to one of the storyboards, the corresponding storyboard is updated, or the storyboard update controls corresponding to each storyboard can be triggered in sequence to complete the update of all storyboards.

[0179] When there is only one storyboard update control, the storyboard update control is used to update multiple storyboards with one click. Specifically, after multiple storyboard texts and multiple storyboards are displayed on the storyboard generation interface, it is also possible to update the multiple storyboards to multiple third new storyboards in response to a trigger operation on the storyboard update control; among them, there is a corresponding relationship between the multiple storyboards and the multiple third new storyboards, and the content of the third new storyboard is associated with the content of the corresponding storyboard.

[0180] It should be noted that the content of the third new storyboard is also generated based on the content of the storyboard text of the corresponding storyboard, so it is associated with the content of the corresponding storyboard; at the same time, when there are multiple storyboard update controls, it is still possible to display the storyboard update control for updating multiple storyboards with one click. In this regard, the embodiments of the present application do not make any limitations.

[0181] Exemplarily, continue to refer to Figure 13 , what 1307 indicates is the storyboard update control, so as to update all storyboards in response to a trigger operation on the storyboard update control.

[0182] In actual implementation, there is an arrangement order for multiple storyboards. Among the multiple storyboards, there is an upward association control for the third target storyboard. The upward association control is used to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard. The third target storyboard is any storyboard other than the first storyboard among the multiple storyboards; after multiple storyboard texts and multiple storyboards are displayed on the storyboard generation interface, it is also possible to update the third target storyboard to a second new storyboard in response to a trigger operation on the upward association control. Specifically, in response to a trigger operation on the target upward association control among one or more upward association controls, the storyboard associated with the target upward association control is updated to a second new storyboard; among them, the second new storyboard is obtained by updating the content of the storyboard associated with the target upward association control based on the content of the previous storyboard of the storyboard associated with the target upward association control.

[0183] It should be noted that each storyboard is associated with an upward association control, but only the third target storyboard can activate, that is, trigger the upward association control; the upward association control is used to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard, that is, to derive the content of the current third target storyboard according to the content of the previous storyboard. In this way, the relevance of the content and effects between storyboards can be improved. For example, if there is a group of birds flying by in the previous storyboard, then based on the content of the previous storyboard, the content of the third target storyboard is updated so that the third target storyboard also includes this group of birds.

[0184] Meanwhile, when the upward association control is triggered, the corresponding third target storyboard can only be associated upward by default and cannot be associated downward, and the storyboard ranked first cannot be associated upward. In addition, for two adjacent storyboards, if both can enable the upward association control but cannot be enabled simultaneously, the storyboard ranked later can only be associated after the storyboard ranked earlier is associated, that is, after the content is updated, so as to prevent the storyboard ranked later from performing multiple content updates and wasting resources.

[0185] Exemplarily, continue to refer to Figure 13 , the one indicated by 1308 is the upward association control of one of the third target storyboards, and the upward association controls of other third target storyboards are similar; thus, in response to the triggering operation for the upward association control corresponding to this third target storyboard, based on the content of the previous storyboard, this third target storyboard is updated to the second new storyboard, or the upward association controls corresponding to each third target storyboard can be triggered in sequence to complete the association of all storyboards.

[0186] In actual implementation, the function of the upward association control can also be introduced. Specifically, it is determined that the operation is performed by the target object; before updating the third target storyboard to the second new storyboard for the triggering operation of the upward association control, if the current account generates a long video type video for the first time, function introduction information can also be displayed; wherein, the function introduction information is used to introduce the function of the upward association control.

[0187] It should be noted that the function introduction information can be displayed before the storyboard generation interface is displayed, or can be displayed when the storyboard generation interface is displayed. In this regard, the embodiments of the present application do not make a limitation.

[0188] Exemplarily, refer to Figure 14 , Figure 14 is a schematic diagram of the function introduction information provided by the embodiments of the present application. Based on Figure 14 , before the storyboard generation interface is displayed, function introduction information as shown in Figure 14 can also be displayed, so as to respond to the determination instruction for the function introduction control triggered by the control indicated by 1401 and display the storyboard generation interface.

[0189] It should be noted that when a long video type video is generated, the corresponding account is detected, and when the detection result indicates that the corresponding account does not carry the target identifier, it is determined that the corresponding account generates a long video type video for the first time, and after the long video is generated, the corresponding account is marked with the target identifier; when the detection result indicates that the corresponding account carries the target identifier, it is determined that the corresponding account is not the first time to generate a long video type video. Among them, the target identifier is used to indicate whether the corresponding account is the first time to generate a long video type video.

[0190] In actual implementation, the number of candidate storyboards corresponding to each storyboard text can be multiple. After multiple storyboard texts and multiple storyboards are displayed on the storyboard generation interface, it is also possible to display multiple storyboard texts on the storyboard generation interface and display multiple candidate storyboards corresponding to each storyboard text; for each storyboard text, in response to a selection operation on a target candidate storyboard among the multiple candidate storyboards, control the target candidate storyboard to be in a selected state, and determine the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

[0191] It should be noted that each candidate storyboard is also associated with a preview control, so that in response to a trigger operation on the preview control of the target candidate storyboard, the corresponding target candidate storyboard is previewed, so as to select a suitable fourth target storyboard from the multiple candidate storyboards.

[0192] Exemplarily, refer to Figure 15 , Figure 15 is a schematic diagram of multiple candidate storyboards provided by an embodiment of the present application. Based on Figure 15 , the three candidate storyboards indicated by the dashed box 1501 correspond to each storyboard text. Therefore, for each storyboard text, a fourth target storyboard is selected from the three candidate storyboards to determine multiple storyboards.

[0193] In actual implementation, the storyboard text can also be modified. Specifically, in response to a text modification operation on the target storyboard text among the multiple storyboard texts, the modified target storyboard text is displayed. Thus, in the subsequent process, based on the modified target storyboard text, in response to a video generation instruction, a video of the target video type including the modified target storyboard text is displayed.

[0194] It should be noted that after the storyboard text is modified, a trigger operation needs to be performed on the storyboard update control corresponding to the corresponding storyboard text to display the storyboard generated based on the modified storyboard text, or the storyboard generated based on the modified storyboard text may not be displayed, but directly in response to a video generation instruction, a video of the target video type including the modified target storyboard text is displayed. As Figure 13 shown, after the storyboard text is modified, a trigger operation is directly performed on the second determination control as described above to receive a video generation instruction.

[0195] In actual implementation, the storyboard generation interface includes at least one generation prompt message. The generation prompt message has a corresponding relationship with the storyboard text and is used to indicate whether the storyboard corresponding to the corresponding storyboard text is successfully generated; thus, based on the fifth target storyboard, in response to a video generation instruction, a video of the target video type including the fifth target storyboard is displayed; wherein, the fifth target storyboard is the storyboard indicated by the generation prompt message as successfully generated.

[0196] It should be noted that when multiple candidate storyboards are generated based on the storyboard text, as long as there is a successfully generated candidate storyboard, the corresponding generation prompt information indicates that the storyboard corresponding to the corresponding storyboard text is successfully generated. Only when all candidate storyboards fail to be generated, the corresponding generation prompt information will indicate that the storyboard corresponding to the corresponding storyboard text is not successfully generated.

[0197] In some embodiments, the target video type is a long video type, and a video of the long video type includes multiple storyboards. After displaying the input target content in response to an input operation on the content input area, a storyboard generation interface can also be directly displayed in response to a determination instruction for the input target content; wherein, the storyboard generation interface includes multiple storyboard texts and multiple storyboards, each storyboard text corresponds to a storyboard, and the storyboard is generated based on the corresponding storyboard text; thereby, based on the multiple storyboard texts and multiple storyboards, in response to a video generation instruction, a video of the target video type including multiple storyboards is displayed.

[0198] It should be noted that the relevant content about the storyboard generation interface here is similar to the relevant content of the storyboard generation interface displayed after the storyboard disassembly interface as described above. For this, the embodiments of the present application will not be elaborated.

[0199] Step 104, in response to a video generation instruction, display a video of the target video type generated based on the target content.

[0200] It should be noted that as described above, the video type includes long videos and short videos, so the video of the target video type generated can also be a long video or a short video.

[0201] In some embodiments, the process of displaying a video of the target video type generated based on the target content in response to a video generation instruction can be that, in response to the video generation instruction, the video of the target video type generated based on the target content is displayed on the video display interface, that is, the target video; wherein, the video display interface further includes at least one of the following: detailed information of the generated video, a third editing control, and a first new video generation control; wherein, the third editing control is used to regenerate a video of the target video type based on the target content, and the first new video generation control is used to re-enter the target content in the content input area and generate a video of the target video type based on the re-entered target content.

[0202] It should be noted that the function of the third editing control is similar to that of the first editing control described above. The detailed information includes at least one of the duration, size, content description, prompt information on whether audio is included, key image frames, and resolution of the target video. Among them, the key image frames are used to indicate the cover of the target video, or pictures input when generating the target video, etc.; the content description is used to indicate the text input when generating the target video, or the text content obtained by analyzing the video content of the target video after generating the target video; and whether audio is included can indicate whether audio is input when generating the target video template, or can indicate whether audio is generated when generating the target video. In this regard, the embodiments of the present application do not make any limitations.

[0203] In some embodiments, the video display interface includes a third editing control, the target video type includes long videos, and long videos include multiple scenes. The detailed information in the detailed interface further includes the duration of each scene and the scene text corresponding to the content of each scene. Exemplarily, referring to Figure 16 , Figure 16 is a schematic diagram of the video display interface of a long video provided by the embodiments of the present application. Based on Figure 16 , the dashed box 1601 indicates the detailed information, 1602 indicates the third editing control, and 1603 indicates the first new video generation control.

[0204] At the same time, as described above, when the target video is a long video, in response to a trigger operation on the third editing control, the last interface when generating the long video can be displayed, which can be a video generation interface, or a scene disassembly interface (such as directly generating the target video based on the scene disassembly interface), or a scene generation interface (such as directly generating the target video based on the scene generation interface, or generating the target video based on the scene generation interface after displaying the scene generation interface based on the scene disassembly interface).

[0205] Based on this, taking the last interface when generating the long video displayed as the scene generation interface as an example, as Figure 13 shown, the target content corresponding to the target video is used to indicate the scene text, the target content corresponding to the target video is multiple scene texts, the scene generation interface further includes multiple scenes, each scene text corresponds to a scene, and the scene is generated based on the corresponding scene text; thus, in response to a video generation instruction, after displaying the video of the target video type generated based on the target content on the video display interface, it is also possible to, in response to a trigger operation on the third editing control, such as Figure 13The storyboard generation interface shown, and display multiple storyboard texts and multiple storyboards determined based on the target content on the storyboard generation interface; based on the storyboard generation interface, in response to editing operations on the multiple storyboard texts and multiple storyboards, display the edited multiple storyboard texts and the multiple storyboards; in response to a video generation instruction, display a long video type video generated based on the edited multiple storyboard texts and multiple storyboards.

[0206] Or, taking the last interface when generating a long video as the storyboard disassembly interface as an example, as Figure 12 shown, the target content corresponding to the target video is used to indicate the storyboard text, the target content corresponding to the target video is multiple storyboard texts, each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; thus, after displaying a target video type video generated based on the target content on the video display interface in response to a video generation instruction, it is also possible to, in response to a triggering operation on the third editing control, such as Figure 12 the storyboard disassembly interface shown, and display multiple storyboard texts determined based on the target content on the storyboard disassembly interface; based on the storyboard disassembly interface, in response to editing operations on the multiple storyboard texts, display the edited multiple storyboard texts; in response to a video generation instruction, display a long video type video generated based on the edited multiple storyboard texts.

[0207] It should be noted that the editing operations triggered based on the storyboard generation interface can be, as described above, operations triggered based on various controls in the storyboard generation interface such as update controls, storyboard update controls, sorting adjustment controls, parameter setting controls, or can also be editing operations on the storyboard text as described above; similarly, the editing operations triggered based on the storyboard disassembly interface can be, as described above, operations triggered based on various controls in the storyboard disassembly interface such as sorting adjustment controls, parameter setting controls, or can also be editing operations on the storyboard text as described above.

[0208] In actual implementation, when a triggering operation on the first new video generation control is received, then display Figure 4 the video generation interface shown in b in

[0209] In some other embodiments, the video display interface includes a third editing control, and the type of the target video includes short videos; refer to Figure 17 , Figure 17 is a schematic diagram of the video display interface of the short video provided by the embodiments of the present application. Based on Figure 17, the details are indicated by the dashed box 1701, the third editing control is indicated by 1702, and the first new video generation control is indicated by 1703; meanwhile, as described above, when the target video is a short video, in response to a trigger operation on the third editing control, a video generation interface is displayed. Specifically, in response to a video generation instruction, after a video of the target video type generated based on the target content is displayed on the video display interface, it is also possible to, in response to a trigger operation on the third editing control, display a video generation interface including a content input area, and the target content is displayed in the content input area; in response to an editing operation on the target content, a new target content obtained by editing the target content is displayed; in response to a video generation instruction, a video of the short video type generated based on the new target content is displayed.

[0210] It should be noted that the editing operation on the target content in the content input area can be an editing operation on the input text, or an editing operation on the input audio or picture. For example, deleting the original audio and / or picture, or uploading new audio and / or picture again, etc. Different target contents result in different objects for the editing operation.

[0211] In actual implementation, when a trigger operation on the first new video generation control is received, then the Figure 4 video generation interface shown in a of is displayed.

[0212] In some embodiments, the process of displaying a video of the target video type generated based on the target content in response to a video generation instruction can be that, in response to a video generation instruction, at least one of a first progress prompt message, a fourth editing control, a first new video generation control, and a generation cancellation control is displayed; wherein, the first progress prompt message is used to prompt the generation progress of the video of the target video type, the fourth editing control is used to generate a video of the target video type again based on the target content, the first new video generation control is used to re-enter the target content in the content input area and generate a video of the target video type based on the re-entered target content, and the generation cancellation control is used to terminate the generation process of the video of the target video type; when the first progress prompt message indicates that the generation of the video of the target video type is completed, the first progress prompt message is cancelled and a video of the target video type generated based on the target content is displayed.

[0213] It should be noted that when the first progress prompt message is displayed, it indicates that the video of the target video type is in the process of being generated. When the generation progress of the video of the target video type indicated by the first progress prompt message reaches 100%, it indicates that the video of the target video type has been generated. In addition, the fourth editing control is similar to the third editing control described above. At the same time, when the video of the target video type is a long video or a short video, the relevant content of the fourth editing control is also similar to the relevant content of the third editing control. Therefore, the embodiments of the present application will not elaborate on this.

[0214] Exemplarily, refer to Figure 18 , Figure 18 which is a schematic diagram of the first progress prompt message provided by the embodiments of the present application. Based on Figure 18 , in response to the video generation instruction, the first progress prompt message indicated by the dashed box 1801 in Figure 18 , the fourth editing control indicated by 1802, the first new video generation control indicated by 1803, and the generation cancellation control indicated by 1804 are displayed.

[0215] In some embodiments, the process of displaying the video of the target video type generated based on the target content in response to the video generation instruction may be to display at least one of the second waiting prompt message, the sixth editing control, the first new video generation control, and the waiting cancellation control in response to the video generation instruction. The second waiting prompt message is used to indicate the duration to wait before starting to generate the video of the target video type. The sixth editing control is used to regenerate the video of the target video type based on the target content. The first new video generation control is used to re-enter the target content in the content input area and generate the video of the target video type based on the re-entered target content. The waiting cancellation control is used to terminate the waiting process required to generate the video of the target video type. When the second waiting prompt message indicates that the duration to wait before starting to generate the video of the target video type is zero, the second waiting prompt message is cancelled and the video of the target video type generated based on the target content is displayed.

[0216] It should be noted that when the second waiting prompt message is displayed, it indicates that the video of the target video type is in the process of waiting to be generated, that is, the video of the target video type has not started to be generated and is in the queue state. When the second waiting prompt message indicates that the duration to wait before starting to generate the video of the target video type is zero, it indicates that the waiting for the video of the target video type has ended, that is, there is no video being generated in front and there is no need to queue. In addition, the sixth editing control is similar to the third editing control described above. At the same time, when the target video is a long video or a short video, the relevant content of the sixth editing control is also similar to the relevant content of the third editing control. Therefore, the embodiments of the present application will not elaborate on this.

[0217] Exemplarily, refer toFigure 19 , Figure 19 is a schematic diagram of the second waiting prompt information provided by an embodiment of the present application. Based on Figure 19 , in response to a video generation instruction, the second waiting prompt information indicated by the dashed box 1901 in Figure 19 , the sixth editing control indicated by 1902, the first new video generation control indicated by 1903, and the waiting cancellation control indicated by 1904 are displayed.

[0218] It should be noted that in addition to directly displaying the first progress prompt information or the second waiting prompt information, at least one of the second waiting prompt information, the sixth editing control, and the first new video generation control can also be displayed first. Thus, when the second waiting prompt information indicates that the waiting duration for starting to generate a video of the target video type is zero, after canceling the display of the second waiting prompt information, at least one of the first progress prompt information, the fourth editing control, and the first new video generation control is displayed. Furthermore, when the first progress prompt information indicates that the target video generation is completed, the display of the first progress prompt information is canceled, and a video of the target video type generated based on the target content is displayed. In this regard, the embodiments of the present application do not make any limitations.

[0219] In some embodiments, after displaying the video generation interface, a video generation prompt information can also be displayed; wherein, the video generation prompt information is used to indicate the number of videos that can be generated within the target time; thus, the process of displaying a video of the target video type generated based on the target content in response to a video generation instruction can be that, in response to the video generation instruction, when the video generation prompt information indicates that the number of videos that can be generated within the target time is not zero, a video of the target video type generated based on the target content is displayed.

[0220] It should be noted that the target time here is used to indicate the target period, which can be preset, such as one day, one week, etc.; for the process of displaying the video generation prompt information, it can be directly displaying the video generation prompt information, or it can be displaying the input target content in response to an input operation on the content input area and then displaying the video generation prompt information, or it can also be displaying the video generation prompt information after receiving the triggering operation on the first determination control in the activated state as described above. In this regard, the embodiments of the present application do not make any limitations.

[0221] Meanwhile, the video prompt information can be displayed at an associated position of the content input area, such as one of the upper position, lower position, left position, and right position of the content input area, or it can also be displayed at an associated position of the first determination control, such as one of the upper position, lower position, left position, and right position of the first determination control.

[0222] For example, directly display video generation prompt messages, such as "There are 10 remaining production tasks today", indicating that the target number of videos that can be generated within one day is 10; or "The video generation quota for today has been used up, please try again tomorrow", indicating that the target number of videos that can be generated within one day is zero. In this way, when responding to a video generation instruction, when the video generation prompt message indicates that the number of videos that can be generated within the target time is not zero, display a video of the target video type generated based on the target content.

[0223] In some embodiments, the video generation interface further includes a video display control, which is used to display at least one asset video. The at least one asset video includes at least one of a successfully generated historical success video, a failed historical failure video, a video being generated, and a video waiting to be generated; thus, after responding to a video generation instruction and displaying a video of the target video type generated based on the target content, it is also possible to respond to a trigger operation on the video display control and display at least one asset video including a video of the target video type.

[0224] It should be noted that the video display control is used to display an asset video display interface, and the asset video display interface further includes at least two video controls, and different video controls are used to display different types of asset videos; in response to the target video control among the at least two video controls being in a selected state, display at least one asset video of the target type.

[0225] It should be noted that the arrangement order of the successfully generated historical success video, the failed historical failure video, the video being generated, and the video waiting to be generated among the at least one asset video can be preset, such as in the order of waiting video, video being generated, failed video, and successful video, and are displayed in sequence from top to bottom.

[0226] Exemplarily, refer to Figure 20 , Figure 20 is a schematic diagram of the asset video display interface provided by an embodiment of the present application. Based on Figure 20 , in response to the long video control being in a selected state, display nine asset videos of the long video type. Among them, the first, second, and third asset videos are waiting videos, the fourth and fifth asset videos are videos being synthesized, the sixth and seventh asset videos are failed videos, and the eighth and ninth asset videos are successful videos.

[0227] It should be noted that at the associated position of each asset video, at least one of the duration, generation time, content description, and key image frame of the corresponding asset video will also be displayed; at the same time, a refresh control will also be displayed in the asset video display interface, so as to respond to a trigger operation on the refresh control and refresh at least one asset video in the asset video display interface.

[0228] Next, the different asset videos will be described separately.

[0229] In some embodiments, the asset video includes a success video, and thus the playback control corresponding to the success video is displayed; in response to a trigger operation on the playback control, the success video is played; when the success video finishes playing, in response to a trigger operation on the success video, a details interface for displaying the details information of the success video is displayed; wherein, the details interface further includes a fifth editing control and a first new video generation control, the fifth editing control is used to regenerate a video of the video type corresponding to the success video based on the content of the success video, and the first new video generation control is used to re-enter target content in the content input area and generate a video of the video type corresponding to the success video based on the re-entered target content.

[0230] It should be noted that when the asset video is a success video, the key image frame is used to indicate the cover of the success video (such as can be preset), or the picture input when generating the success video, etc.; the details information includes at least one of the duration, size, content description, prompt information on whether audio is included, resolution, and key image frame of the success video. When the success video is a long video including multiple scenes, the details information may further include the duration of each scene and the scene text corresponding to the content of each scene; at the same time, the fifth editing control is similar to the third editing control described above. Also, when the target video is a long video or a short video, the relevant content of the fifth editing control is similar to the relevant content of the third editing control. For this, the embodiments of the present application will not elaborate.

[0231] In some embodiments, the asset video includes a failure video, the key image frame is used to indicate the cover of the success video (such as can be preset), or the picture input when generating the success video, etc., and the failure video is associated with a regeneration control, and the regeneration control is used to regenerate the failure video; thus, after displaying at least one asset video including the target video in response to a trigger operation on the video display control, it is also possible to, in response to a trigger operation on the regeneration control associated with the target failure video among at least one failure video, regenerate the failure video; when the regeneration is successful, switch the failure video in at least one asset video to the regenerated video; when the regeneration fails, display a regeneration failure prompt message, and the regeneration failure prompt message is used to prompt that the regeneration of the failure video fails.

[0232] Exemplarily, such as Figure 20As shown, in response to a trigger operation for the regeneration control associated with the seventh failed video as indicated in 2001, the corresponding failed video is regenerated; when the regeneration is successful, the corresponding failed video is switched to the regenerated video; when the regeneration fails, a regeneration failure prompt message is displayed.

[0233] It should be noted that for a failed video, when the failed video is a long video, the reason for the generation failure will also be displayed at the associated position, such as the failure of shot generation or the failure of shot synthesis, etc.; at the same time, when the regeneration is successful, while switching the failed video to the regenerated video, a playback control will also be displayed, so that the regenerated video can be played. Moreover, after switching the failed video to the regenerated video, the switched video can be regarded as a successful video, so as to respond to the trigger operation for this successful video and implement the relevant processes of the asset video being a successful video as described above. Regarding this, the embodiments of the present application will not be elaborated.

[0234] It should be noted that when the number of at least one asset video generated is multiple, if there is a target video with successful generation, that is, among the at least one asset videos with successful generation, it is regarded as a successful video, and among the at least one asset videos with failed generation, it is regarded as a failed video.

[0235] In some embodiments, the asset video includes a video being generated. The key image frame is used to indicate the cover of the successful video (such as can be preset), or the picture input when generating the successful video, etc. Thus, in the associated area of the video being generated, a second progress prompt message is displayed. The second progress prompt message is used to prompt the generation progress of the video being generated, and the generation progress indicated by the progress prompt message increases with the passage of time; in response to a trigger operation for the video being generated, a generation cancellation control is displayed, and the generation cancellation control is used to terminate the generation process of the video being generated.

[0236] Exemplarily, as Figure 20 shown, in the associated areas of the fourth and fifth videos being generated, the second progress prompt message as indicated in 2002 is displayed, where the second progress prompt message is used to prompt the generation progress of each video being generated.

[0237] It should be noted that in response to a trigger operation for the video being generated, the interface as Figure 18 shown is displayed, where, as described above, the interface includes the first progress prompt message indicated by the dashed box 1801, the fourth editing control indicated by 1802, the first new video generation control indicated by 1803, and the generation cancellation control indicated by 1804.

[0238] In some embodiments, the asset video includes a waiting video. The key image frame is used to indicate the cover of the successful video (such as a pre-set one), or the picture input when generating the successful video, etc. Thus, in the associated area of the waiting video, a first waiting prompt message is displayed. The first waiting prompt message is used to indicate that the corresponding video is waiting to be generated; in response to a trigger operation on the waiting video, a waiting cancellation control is displayed, and the waiting cancellation control is used to terminate the waiting process of the waiting video.

[0239] Exemplarily, as Figure 20 shown, in the associated areas of the first, second, and third waiting videos, the first waiting prompt message indicated by 2003 is displayed, where the first waiting prompt message is used to indicate that the corresponding video is waiting to be generated.

[0240] It should be noted that for the waiting video, when the waiting video is a long video, waiting process information will also be displayed at the associated position, such as the storyboard to be generated, or the storyboard is being generated, or waiting to be synthesized, etc.

[0241] When the waiting process information indicates that the corresponding waiting video is in the state of waiting for the storyboard to be generated, in response to a trigger operation on the waiting video, the storyboard disassembly interface or the video generation interface as described above is displayed; when the waiting process information indicates that the corresponding waiting video is in the state of the storyboard being generated, in response to a trigger operation on the waiting video, the storyboard generation interface as described above is displayed; when the waiting process information indicates that the corresponding waiting video is in the state of waiting to be synthesized, in response to a trigger operation on the waiting video, the interface as Figure 19 shown is displayed, where, as described above, the interface includes a second waiting prompt message indicated by the dashed box 1901, a sixth editing control indicated by 1902, a first new video generation control indicated by 1903, and a waiting cancellation control indicated by 1904.

[0242] In actual implementation, each asset video is also associated with a deletion control. Thus, in response to a trigger operation on the deletion control corresponding to the target asset video in at least one asset video, a deletion prompt message is displayed, and the deletion prompt message is used to determine whether to delete the target asset video; in response to a confirmation instruction triggered based on the deletion prompt message, the target asset video is deleted.

[0243] Exemplarily, continue to refer to Figure 20 and refer to Figure 21 , Figure 21 is a schematic diagram of the deletion prompt message provided by an embodiment of the present application. Based on Figure 20 , in response to a trigger operation on the deletion control indicated by 2004 associated with the seventh failed video, the deletion prompt message as Figure 21 indicated is displayed, and thus, in response to a confirmation instruction triggered based on the deletion prompt message, the target asset video is deleted.

[0244] In some embodiments, after displaying the input target content in response to an input operation on the content input area, in response to a video generation instruction, when the generation of a video of a target video type fails, at least one of a generation failure prompt message, a regenerate control, and a backtracking control may be displayed; wherein, the generation failure prompt message is used to indicate that the generation of a video of the target video type fails and the reason for the generation failure, the regenerate control is used to regenerate the video of the target video type, and the backtracking control is used to return to the previous step.

[0245] Exemplarily, referring to Figure 22 , Figure 22 is a schematic diagram of the generation failure prompt message, the regenerate control, and the backtracking control provided by an embodiment of the present application. Based on Figure 22 , what is indicated by the dashed box 2201 is the generation failure prompt message, what is indicated by 2202 is the regenerate control, and what is indicated by 2203 is the backtracking control.

[0246] It should be noted that the storyboard disassembly interface, the storyboard generation interface, and the interface for displaying the second waiting prompt message and the first process prompt message may all include a backtracking control, so as to return to the previous step with respect to the backtracking control; however, after triggering the backtracking control displayed on the interface where the generation failure prompt message is displayed, the content in the previous step returned can be modified. For example, when the generation of a long video fails, in response to a triggering operation on the backtracking control, the storyboard disassembly interface or the storyboard generation interface is displayed. At this time, the storyboard text and / or the storyboard can be modified based on the corresponding storyboard disassembly interface and the storyboard generation interface, so as to regenerate the long video again based on the modified storyboard text and / or the storyboard;

[0247] For the situation where it is not a generation failure, after triggering the backtracking control displayed on the corresponding interface, the content in the previous step returned cannot be modified. For example, on the interface for displaying the second waiting prompt message, a backtracking control is displayed. In response to a triggering operation on the backtracking control, the interface of the previous step is displayed, but the content of this interface cannot be modified.

[0248] In some embodiments, the video generation interface further includes a style setting control, and the style setting control is used to set the video style of the generated video; after displaying the video generation interface, in response to a triggering operation on the style setting control, at least one video style may be displayed; in response to a selection operation on a target video style among the at least one video style, the target video style is controlled to be in a selected state; thus, in the process of displaying a video of a target video type generated based on the target content in response to a video generation instruction, it may be that in response to a video generation instruction, a video of a target video type with a target video style generated based on the target content is displayed.

[0249] It should be noted that the video style here can be a cartoon style, a science fiction style, a realistic style, etc.; by way of example, see Figure 23 , Figure 23 which is a schematic diagram of at least one video style provided by an embodiment of the present application. Based on Figure 23 , in response to a trigger operation on the style setting control indicated by 2301, at least one video style indicated by the dashed box 2302 is displayed, so that in response to a selection operation on the target video style among at least one video style, the target video style is controlled to be in a selected state, and then a video of the target video type with the target video style is generated.

[0250] In actual implementation, after controlling the target video style to be in a selected state in response to a selection operation on the target video style among at least one video style, it is also possible to, in response to a confirmation instruction for the target video style in the selected state, display an identifier of the target video style on the style setting control, and the identifier is associated with a style cancellation control; wherein, the identifier of the target video style is used to identify that the generated video of the target video type has the target video style; in response to a trigger operation on the style cancellation control, the identifier of the target video style is cancelled from display on the style setting control; thus, the process of displaying a video of the target video type with the target video style generated based on the target content in response to a video generation instruction can be to display a video of the target video type with the default style generated based on the target content in response to the video generation instruction.

[0251] It should be noted that the identifier of the target video style can refer to the name of the target video style. After the identifier of the target video style is cancelled from display on the style setting control, the original video style control will be displayed, so that if a video generation instruction is received at this time, a video of the target video type with the default style will be generated. By way of example, see Figure 24 , Figure 24 which is a schematic diagram of the identifier of the target video style provided by an embodiment of the present application. Based on Figure 24 a, in response to a confirmation instruction for the target video style in the selected state, an identifier of the target video style, that is, the cartoon style, is displayed on the style setting control indicated by 2401, wherein the style cancellation control associated with the identifier is shown as 2402, so that in response to a trigger operation on the style cancellation control, the identifier of the target video style is cancelled from display on the style setting control, as shown in Figure 24 2403 in b of

[0252] In some embodiments, when there are multiple videos of the generated target video type, the process of displaying the videos of the target video type generated based on the target content is to automatically play each video of the target video type in sequence until the last video of the target video type is played.

[0253] In some embodiments, after responding to a video generation instruction and displaying the video of the target video type generated based on the target content, it is also possible to display operation controls for the target video. The operation controls include at least one of the following: a like control for liking the target video, a share control for sharing the target video, a download control for downloading the target video, a dislike control for disliking the target video, a play control for playing the target video, and a full-screen control for full-screen displaying the target video; in response to a trigger operation on the target operation control, perform the operation indicated by the target operation control on the target video, where the target operation control is one of the like control, share control, download control, dislike control, play control, and full-screen control.

[0254] Exemplarily, referring to Figure 25 , Figure 25 is a schematic diagram of the operation controls provided by the embodiments of the present application. Based on Figure 25 , what is indicated by the dashed box 2501 are respectively the like control, dislike control, share control, and download control, what is indicated by 2502 is the play control, and what is indicated by 2503 is the full-screen control.

[0255] In some embodiments, after responding to a video generation instruction and displaying the video of the target video type generated based on the target content, it is also possible to display at least one interaction control for the generated video. The interaction control is one of the following controls: a local editing control for locally editing the generated video, a canvas expansion control for editing the size of the generated video, a motion brush control for changing static objects in the generated video into dynamic objects, a camera movement parameter setting control for setting the camera movement parameters of the generated video, a background removal control for removing the background of the generated video, an object erasing control for erasing target objects in the generated video, a scene detection control for splitting the video into different video segments based on different scenes in the generated video, a depth of field setting control for setting the depth of field effect for the generated video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the generated video, an action sequence generation control for setting an action sequence for an object in the generated video, and a second new video generation control for generating a new video based on the audio data in the generated video; in response to a trigger operation on the target interaction control in the at least one interaction control, perform the interaction operation indicated by the target interaction control on the video of the generated target video type.

[0256] In actual implementation, when the target interaction control is a local editing control, in response to a triggering operation on the local editing control, a local editing box and a text input box are displayed. In response to an input operation based on the text input box, the input text content is displayed, and the text content is used to describe the editing task. In response to a video editing instruction, the target video after local editing is displayed, where the target video is a video obtained by editing the area indicated by the local editing box in the generated video based on the text content.

[0257] Exemplarily, refer to Figure 26 , Figure 26 which is a schematic diagram of interacting with the generated video based on the local editing control provided by an embodiment of the present application. Based on Figure 26 , when the target interaction control is a local editing control, in response to a triggering operation on the local editing control, a local editing box indicated by 2601 in a as shown in Figure 26 and a text input box indicated by 2602 are displayed. Then, in response to an input operation based on the text input box, the input text content indicated by 2603 in b as shown in Figure 26 is displayed, so that in response to a video editing instruction, the target video after local editing is displayed.

[0258] It should be noted that when the local editing box is displayed, the size of the local editing box can also be adjusted. At the same time, an editing effective period can also be set. Specifically, after the local input box is displayed, a forward time input box and a backward time input box are displayed. When the user inputs a target period in the forward time input box, by default, starting from the time corresponding to the image frame displayed by the local editing box, the video within the forward target period has an editing effect. When the user inputs a target period in the backward time input box, by default, starting from the time corresponding to the image frame displayed by the local editing box, the video within the backward target period has an editing effect. When there is no input content in the forward time input box and the backward time input box, by default, the entire video has an editing effect.

[0259] It should be noted that the input text content can be adding a target object in the local editing box, or erasing a certain object in the local input box, etc.

[0260] In actual implementation, when the target interaction control is a canvas expansion control, in response to a triggering operation on the canvas expansion control, canvases of at least one size are displayed. In response to a selection operation on the canvas of the target size, the generated video is displayed on the canvas of the target size. Thus, on the canvas of the target size, in response to a size adjustment operation on the generated video, the target video is displayed, and the target video is the generated video after size adjustment, and the size of the target video is not greater than the target view. In this way, based on the canvas expansion control, the convenience of adjusting the size of the video is improved.

[0261] Exemplarily, referring to Figure 27 , Figure 27 is a schematic diagram of interacting with the generated video based on the canvas expansion control provided by an embodiment of the present application. Based on Figure 27 , when the target interaction control is the canvas expansion control, in response to a selection operation for the canvas of the target size, the generated video indicated by 2702 is displayed on the canvas of the target size indicated by 2701.

[0262] In actual implementation, when the target interaction control is the action sequence generation control, in response to a trigger operation for the action sequence generation control, at least one candidate action sequence is displayed. In response to a trigger operation for the target candidate action sequence among the at least one candidate action sequences, the selected target candidate action sequence is displayed. Then, in response to a video editing instruction, the target video is displayed, where the target video is the video obtained by setting the object in the generated video to the target candidate action sequence.

[0263] It should be noted that the action sequence is used to indicate a continuous action video, such as a dance video, etc. And setting the object in the generated video to the target candidate action sequence means making the corresponding object execute the actions indicated by the target candidate action sequence, such as making the corresponding object dance with the dance actions indicated by the dance video.

[0264] It should be noted that when displaying at least one candidate action sequence, a text input box is also displayed. Thus, in response to an input operation based on the text input box, the input text content is displayed, and the text content is used to describe the editing task. Thus, in response to a video editing instruction, the generated target video also indicates the video obtained by setting the object in the generated video to the target candidate action sequence based on the text content.

[0265] In actual implementation, when the target interaction control is the motion brush control, in response to a trigger operation for the motion brush control, the brush control is displayed. The brush control can move on the image frame of the generated video and is used to select the target object in a static state in the corresponding image frame, so as to change the state of the selected target object from static to dynamic; in response to a selection operation triggered by the brush control, the selected target object in a static state is displayed on the corresponding image frame; in response to a determination instruction for the target object, the state of the selected target object is changed from static to dynamic; in response to a video editing instruction, the target video is displayed, where the target video is the video including the target object whose state has changed from static to dynamic.

[0266] In actual implementation, when the target interaction control is a camera movement parameter setting control, in response to a trigger operation on the camera movement parameter setting control, a camera movement parameter setting area is displayed. The parameter setting area is used to set the camera movement parameters of the generated video. Based on the camera movement parameter setting area, in response to a camera movement parameter setting operation, the set camera movement parameters for the generated video are displayed. Among them, the camera movement parameters include the movement direction and movement speed of the camera, etc.

[0267] In actual implementation, when the target interaction control is a background removal control, in response to a trigger operation on the background removal control, the background of the generated video is removed to obtain a target video. Among them, the background of the video can be, for example, the sky, mountains, etc.

[0268] In actual implementation, when the target interaction control is an object erasing control, in response to a trigger operation on the object erasing control, an erasing tool in a draggable state is displayed. The erasing tool can be moved on the image frames of the generated video to select the target object included in the corresponding image frame, so as to delete the selected target object. In response to a selection operation triggered by the erasing tool, the selected target object is displayed on the corresponding image frame. In response to a determination instruction for the target object, the selected target object is deleted. In response to a video editing instruction, a target video is displayed. The target video is a video including the deleted target object.

[0269] In actual implementation, when the target interaction control is a scene detection control, in response to a trigger operation on the scene detection control, based on at least one scene included in the generated video, the generated video is split into at least one video segment. Among them, each video segment corresponds to one scene.

[0270] In actual implementation, when the target interaction control is a depth of field setting control, in response to a depth of field setting operation triggered by the depth of field setting control, the set target depth of field parameters are displayed. Thus, in response to a determination instruction for the set target depth of field parameters, a target video with the target depth of field parameters is displayed.

[0271] In actual implementation, when the target interaction control is a frame rate adjustment control, in response to a frame rate adjustment operation triggered by the frame rate adjustment control, the adjusted target frame rate is displayed. Thus, in response to a determination instruction for the adjusted target frame rate, a target video with the target frame rate is displayed

[0272] In actual implementation, when the target interaction control is a second new video generation control, in response to a trigger operation on the second new video generation control, based on the audio data of the generated video, a target video is generated. The content of the target video is associated with the audio data.

[0273] It should be noted that, in addition to interacting based on the generated video, other videos or audios can also be selected for interaction. Specifically, at least one interaction control is displayed. In response to a trigger operation on a target interaction control among the at least one interaction control, at least one candidate video is displayed. In response to a selection operation on a target candidate video among the at least one candidate video, the target candidate video is controlled to be in a selected state. In response to a determination instruction for the target candidate video in the selected state, for the target candidate video, an interaction operation indicated by the target interaction control is executed. Among them, when the target interaction control is a second new video generation control, the at least one candidate video also includes audio. In response to a determination instruction for the target audio or the target candidate video in the selected state, a video generated based on the audio data of the target audio or the target candidate video is displayed.

[0274] In actual implementation, when performing an interaction operation on a video based on an interaction control, as described above, a text input box will be displayed. Thus, while editing the video based on the interaction control, it is also possible to respond to an input operation based on the text input box and display the input text content, where the text content is used to describe the editing task. In response to a video editing instruction, a target video is displayed, where the target video is a video obtained based on the text content and the interaction control. For example, content such as "warm up the picture" can be input in the text input box to change the video picture filter effect, or content such as "delete XX object" can be input to delete the XX object in the video.

[0275] Applying the above embodiments of the present application, first, at least one video generation control for generating videos of different video types is displayed. Then, based on the at least one video generation control, in response to a target video type among at least two video types being in a selected state, a content input area corresponding to the target video type is displayed, and based on the content input area, target content is input. Thus, based on the input target content, a target video of the video type corresponding to the target video generation control is generated. In this way, not only is a target video of the target video type corresponding to the input target content generated, such that the video content of the target video meets the user's needs, but also, based on the video generation control, a target video of the target video type is generated, such that the type of the video also meets the user's needs. Compared with the solution where the user still needs to edit the video after it is generated, the present application reduces the editing operations that the user needs to perform on the generated video, not only increasing the human-computer interaction efficiency but also improving the video generation efficiency.

[0276] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0277] There are various types of videos, such as short videos, long videos, etc. The video parameters corresponding to different types of videos are different, and the video generation process is also different. However, when generating videos using related technologies, most of them directly generate videos based on the user's demand for video content, that is, only considering the user's demand for video content while ignoring the user's demand for video type. In this way, after the video is generated, the user still needs to edit the video so that the edited video meets the user's requirements for video type on the basis of meeting the user's requirements for video content. This not only makes the human-computer interaction efficiency too low but also further reduces the video generation efficiency.

[0278] Based on this, the embodiments of the present application provide a video generation method. On the one hand, it comprehensively combines multi-modal capabilities. In addition to supporting pure text-to-video, it also supports image-to-video, text-and-image-to-video, video-to-video, image+video-to-video, etc. On the other hand, by combining the ability to split text into storyboard scripts and the ability to generate and splice short videos, short videos of 2 to 4 seconds can be generated by inputting text. Also, after inputting text (when the text is modified later, the storyboard will change accordingly), the model automatically splits the text into multiple storyboard scripts. These multiple storyboard scripts support editing operations such as addition, deletion, and modification, and also support customizing the association between different storyboards, thereby supporting the generation of longer videos. At the same time, the intermediate steps and information during the long video generation process will be displayed, supporting the user to adjust at any time. In addition, through the laboratory tab (extended editing), different high-order gameplay functions can be entered, supporting functions such as video mask selection and editing, canvas expansion, and generating videos from skeleton points.

[0279] Next, the technical solution of the present application will be described from the product side. The product side of the technical solution of the present application includes two parts. The first part is creating videos, and the second part is high-order gameplay of videos. For the process of creating videos, it is further divided into two processes, namely the process of creating short videos and the process of creating long videos.

[0280] For the process of creating short videos, as Figure 4 shown, goodcases (video templates) are displayed in a waterfall flow form below the input box, including videos and text. When the target video template is in the selected state, the content description of the target video template and the corresponding secondary creation button (the first editing control) are displayed. Then, when a click operation on the secondary creation button is received, the content description of the target video template is filled into the input box.

[0281] In actual implementation, in response to the user's input operation, control Figure 4The immediate generation control (OK control) in

[0282] It should be noted that it can also respond to the user's triggering operation on the parameter setting control in Figure 4 to display the parameter setting area as shown in Figure 11 so as to set the parameters of the video in response to the parameter setting operation triggered based on the parameter setting area. Among them, if the user does not perform the parameter setting operation, the default parameter configuration is used. For example, the default for pictures is none, the default for special effects is none, the default video quality is 896*512, the default ratio is 7:4, the default video duration is 4s, and the default for audio is none.

[0283] In actual implementation, in response to the user's input operation, the immediate generation control is controlled to switch from the unactivated state to the activated state, and after setting the parameters of the video in response to the parameter setting operation triggered based on the parameter setting area, when receiving the triggering operation on the activated immediate generation control, the first waiting prompt message as shown in Figure 19 is displayed. The first waiting prompt message is used to indicate that the short video to be created is in the queuing state. This state is the waiting queue state from when the immediate generation control is clicked until the generation starts, and the corresponding state is waiting in the queue. At the same time, the estimated remaining task quantity will also be displayed.

[0284] In actual implementation, after the queuing state ends, the first progress prompt message as shown in Figure 18 is displayed. The first progress prompt message is used to indicate that the short video to be created is in the generating state. This state is the generating process state from the start of generation to the return of the result, and the video generation progress will also be displayed.

[0285] It should be noted that in the intermediate states, namely the queuing state and the generating state, interface modification operations are not allowed, including text, pictures, parameter configuration, etc.; but users are allowed to create new short videos, such as in response to the click operation on the secondary creation control or the create new video control in Figure 18 or 19. When the user clicks the secondary creation control, the interface shown in Figure 8 b is displayed, and the input content corresponding to the video to be created is filled back into the input box as the new input content, so as to generate a video associated with the video to be created based on the new input content. When the user clicks the triggering operation of the create new video control, the interface shown in Figure 8The interface shown in a, where the input box in the interface is blank, enabling the user to re-enter new content and then generate a video based on the new input content. Among them, when the user clicks on the Figure 18 or the secondary creation control or the create new video control in 19, the currently ongoing task is saved to the asset.

[0286] In actual implementation, three videos are generated each time during the video generation process, and one of them is supported to be selected for preview in the main window; at the same time, the generated videos can also be liked, disliked, shared, and downloaded.

[0287] It should be noted that when all videos fail to be generated, the generation failure interface shown in Figure 22 is displayed. Among them, the generation failure interface includes a regenerate control for regenerating. When the user clicks the regenerate control, the video is regenerated; at the same time, the failed-generated video is also saved as an asset. When part of the generation fails, the successfully generated videos are returned normally, and the failed videos are shown as generation failed at the front end. At the same time, for the videos with partial generation failure, since there are successfully generated videos, this asset is considered to be successfully generated and does not support regeneration, that is, as long as one video is successfully generated, this asset is considered to be in a successfully generated state.

[0288] In actual implementation, the video style can also be set. As shown in Figure 23 a style setting control is displayed. In response to a trigger operation on the style setting control, at least one selectable video style is shown. In response to a selection operation on the target video style among the at least one selectable video styles, a video corresponding to the target video style is generated.

[0289] For the process of creating a long video, it needs to be generated in stages, specifically including four stages. Next, the four stages will be described separately.

[0290] 1) Stage 1: Input content plus configure global parameters.

[0291] As shown in Figure 9 c and d, when no input operation from the user is received, the next step cannot be clicked, that is, the next step button is grayed out (in an inactive state). When an input operation from the user is received, the next step button is clickable (in an active state); among them, the input operation can be at least one of inputting text, pictures, and audio.

[0292] In actual implementation, it will also show as shown in Figure 11The parameter configuration area shown, so as to respond to the parameter setting operation triggered based on the parameter setting area and set the parameters of the video. Among them, if the user does not perform the parameter setting operation, the default parameter configuration is used. For example, the default video duration is 6s, the default video quality is 896*512, the default ratio is 7:4, and there is no audio.

[0293] 2) Phase two: Generate segmented storyboard descriptions and allow users to edit them again.

[0294] In actual implementation, as shown in 12, generate storyboard descriptions (display the storyboard disassembly interface), including the corresponding descriptions and parameter settings of this segment of the video. Each segment of the description supports secondary editing, that is, editing text, adjusting parameter settings (adding special effects), adding padding images (key image frames), etc. The storyboard supports users to adjust the storyboard order; at the same time, users can also customize and add new storyboards and corresponding texts. After adding a new storyboard, they can call AI to automatically generate a storyboard script or customize the input; finally, after all adjustments are completed, in response to the click operation on the generation control, enter the storyboard generation interface;

[0295] It should be noted that for the case where the storyboard generation fails, when all generations fail, a generation failure message (generation failure prompt message) is displayed, and the video is supported to be regenerated.

[0296] 3) Phase three: Generate storyboards.

[0297] In actual implementation, as Figure 13 shown, generate storyboard shots (display the storyboard generation interface). Each shot includes the text description of this storyboard, parameter configuration, and 3 corresponding generated storyboard videos (candidate storyboards). Each storyboard supports secondary editing, that is, editing text, adjusting parameter settings, adding padding images (key image frames), supporting the selection of whether the next-level storyboard is associated with the previous level. The parameter adjustment settings include adjusting special effects and uploading pictures. The order of storyboards can be adjusted between different storyboards; at the same time, after adjustment, it supports regenerating the corresponding storyboard.

[0298] It should be noted that for the process of users selecting storyboards (multiple selections are supported) and the specific videos under the corresponding storyboards (single selection), it is default to select all storyboards, and select the first video under each storyboard. If no storyboard is selected, the button for synthesizing the video (the second confirmation control) cannot be clicked; at the same time, it supports regenerating all storyboards with one key. Among them, users can also upload an existing short video to generate a video (video of the target video type) together with the storyboards on the interface.

[0299] In actual implementation, when some storyboards are successfully generated, the successfully generated storyboards are returned normally, and the failed-generated storyboards are displayed at the front end; when there is at least 1 successfully generated candidate storyboard under the corresponding storyboard text, it is considered a successful generation.

[0300] 4) Stage 4: Synthesize long video.

[0301] In actual implementation, after the user clicks on the composite video control, the storyboard, sequence information and related information selected by the user will be transmitted to the background for video synthesis; at the same time, the waiting state and the generation failure situation are the same as the process of generating a short video; in addition, the generated video can also be liked, disliked, shared and downloaded.

[0302] 5) The information of the completed previous stage can be viewed by returning (based on the backtracking control), but editing is not supported.

[0303] It should be noted that if there is overlapping information in phase 2 and phase 3, if it is modified in phase 3 but regeneration is not triggered, it will be out of sync in phase 2. If regeneration is triggered, it will be synchronized to phase 2.

[0304] For advanced gameplay functions, advanced gameplay functions include but are not limited to video support for style conversion and regeneration (style setting controls), video canvas expansion (canvas expansion controls), addition, deletion and modification of elements in the video (local editing controls), such as adding specific subjects by circling specific areas and editing prompts, motion brushes, camera parameter settings, video generation time extension, multi-image video generation, skeleton point driven video generation (action sequence generation controls), video generation, such as generating new videos through 3 modes (i.e. video + picture, video + style selection, video + text), storyboard function (such as after generating a video, you can choose to add a new scene, re-fill in the prompt words, add a new plot to the entire video, and you can edit each plot Individual settings, newly added scenes do not support uploading pictures again, and you can also delete a scene), video editing tools (background removal controls), object erasing tools (object erasing controls), color grading (this function only needs to enter descriptive text to easily process video filter effects), super slow motion tools (frame rate adjustment controls, this function can reduce the frame rate of your video footage and convert it into a smooth slow motion video), video depth of field tools (depth of field setting controls. This function can automatically process the video into a picture with depth of field effect), scene detection (scene detection controls, this function can automatically divide the footage of your uploaded video into multiple clips) and audio generation video (the second new video generation control).

[0305] In actual implementation, all videos that have undergone advanced gameplay are imported into a unified asset tab, and the function source of each video is identified, supporting quick jumps from each function tab; each function stores its own function assets; at the same time, for the successfully generated assets under the asset tab, an advanced editing shortcut is added, and clicking the shortcut entry can directly display the AI editing function interface (i.e. skipping the video upload step).

[0306] In actual implementation, the process of high-order gameplay specifically includes selecting / uploading videos, editing, generating, and storing assets. For the process of uploading videos, it supports pulling video assets and local uploading, and single selection is supported; pulling existing assets means displaying all successfully generated videos under the asset tab, allowing users to select; while local uploading means supporting local video uploading, and after uploading, cropping the duration is supported. Specifically, the cropping duration can only be 2, 3, or 4 seconds as an integer, and custom cropping of non-integer durations such as 2.5 seconds is not supported.

[0307] In actual implementation, for the functions in high-order gameplay, the input source is always a video. After uploading a video, the sub-function operations under AI editing are independent of each other, and each sub-function has a separate generation button.

[0308] For the process of local video editing, specifically, as Figure 26 shown, during the video playback process, users are supported to select the required area by framing. When framing, the video automatically pauses. Then, after the framing is completed, a prompt word editing box and a selector will pop up. Among them, the editing box supports filling in prompt words, and the selector supports selecting effects such as taking effect forward, taking effect backward, and taking effect for all. If not selected, it defaults to taking effect for all; among them, taking effect forward means that the video from the start of the video to the current time includes the effect of the prompt word change, taking effect backward means that the video from the current time to the end of the video includes the effect of the prompt word change, and taking effect for all means that the entire video includes the effect of the prompt word change; then, conventional parameter settings are carried out, such as supporting the setting of special effects, video quality, sound effects, and styles, and the setting logic reuses the video generation process.

[0309] For the process of AI canvas expansion, as Figure 27 shown, after uploading a video, the selection of the expansion ratio is supported, such as selecting one of the ratios of 16:9, 9:16, or 1:1. After selecting the ratio, a canvas of the corresponding size is displayed, and the video is shown in the middle of the canvas; here, if the original video size is larger than the canvas size, the original video needs to be scaled proportionally to keep the entire video within the canvas; then, the size and position of the video can be adjusted, that is, supporting the adjustment of the video size and dragging to adjust the position in the canvas, but it must be ensured that the entire video is within the canvas range and does not exceed the canvas; then, conventional parameter settings are carried out, such as supporting the setting of special effects, video quality, sound effects, and styles, and the setting logic reuses the video generation process.

[0310] In actual implementation, for the process of generating a video during the expansion of the AI canvas, in response to a click operation on the video generation control, coordinate information is converted based on the video quality, ratio, and selected area and sent to the server. Then, based on the data returned by the server, the status of the video is determined, namely the queuing and generation process status. For the queuing status, this status is the waiting queue status from after clicking to generate until starting to enter the generation process, and the corresponding status is waiting in the queue. At the same time, the front end needs to display the estimated remaining task quantity; for the generating status, this status is the generation process status from the start of generation to the return of the result.

[0311] In actual implementation, after going through the queuing and generation process status, the generated video returned by the server is received and the generated video is added to the asset tab. In this way, the videos generated by the new high-level gameplay are added under the asset tab, and thus the videos can also be further edited with non-high-level gameplay, such as the video generation process described above.

[0312] It should be noted that after generating the video, it is also possible to like, dislike, share, and download the generated video; at the same time, when the generation fails, re-generation of the video is also supported.

[0313] For the process of generating a video driven by bone points, the specific process of generating a video includes selecting an action template such as a dance action template, uploading an image, text description, and parameter settings. Specifically, first, at least one dance video template is displayed, and then one dance video template is selected from at least one dance video template; then, the user performs input operations, such as uploading a local image and / or adding a text description, and then performs conventional parameter settings, such as supporting the setting of special effects, ratio, video quality, sound effects, and style. The setting logic re-uses the video generation process;

[0314] For the process of generating a video during the generation of a video driven by bone points, in response to a click operation on the video generation control, the selected dance template and the information input by the user are sent to the server. Then, based on the data returned by the server, the status of the video is determined, namely the queuing and generation process status. For the queuing status, this status is the waiting queue status from after clicking to generate until starting to enter the generation process, and the corresponding status is waiting in the queue. At the same time, the front end needs to display the estimated remaining task quantity; for the generating status, this status is the generation process status from the start of generation to the return of the result.

[0315] In actual implementation, after going through the queuing and generation process status, the generated video returned by the server is received and the generated video is added to the asset tab. In this way, the videos generated by the new high-level gameplay are added under the asset tab, and thus the videos can also be further edited with non-high-level gameplay, such as the video generation process described above.

[0316] It should be noted that after the video is generated, it is also possible to like, dislike, share, and download the generated video; at the same time, when the generation fails, it also supports regenerating the video.

[0317] In actual implementation, for the videos in the assets (asset videos), sharing and downloading can be performed. Among them, only the successfully generated videos support sharing, and the videos during the generation process do not support sharing. At the same time, the shared content includes the generated video, the corresponding parameter configuration, text, and padding image information of the video.

[0318] In actual implementation, it is also possible to continue generating videos based on the videos in the assets. Specifically, a continue generation button (regenerate control) is displayed. After clicking, it jumps to the editor interface, and at the same time, all information is filled back to the editor to support the user to continue editing and generating.

[0319] In actual implementation, it is also possible to perform secondary creation based on the videos in the assets. Specifically, for the generated videos, a generate again button (fifth edit control) is displayed. After clicking, it directly jumps to the editor corresponding to the long / short video; the corresponding text, pictures, and parameter configuration information are filled back to support the user to directly click to generate; the video generated at this time is a new video, and the generated video is stored as a new asset video in the asset module without overwriting the original asset video. For the secondary creation of long videos, the information of phases one, two, and three is filled back, and the process directly goes to phase three.

[0320] In actual implementation, it is also possible to delete the videos in the assets. Specifically, both the generating / completed assets support deletion; the deletion status is passed from the front end to the back end. The back end determines that if it is an asset that has been synthesized / started to be generated, the data is directly deleted; if it is an asset video to be produced, the video generation request is cancelled.

[0321] Next, the technical solution of this application will be described from the technical side.

[0322] In actual implementation, refer to Figure 28 , Figure 28 which is a schematic structural diagram of the video generation model provided by the embodiments of this application. Based on Figure 28 , the video general model adopted by the embodiments of this application is a conditional video generation model based on the diffusion model. Different from the classical text-to-image diffusion model SD model (stablediffusion[1]), the 3D Unet network designed in the embodiments of this application mainly includes as Figure 28The main network indicated by the dashed box 2802 and the conditional control network indicated by the dashed box 2801. The conditional control network can encode and model various input reference conditions (such as text, images, videos, skeletons, etc.), and send the encoded reference condition information into the main network to guide the main network to learn. The main network mainly learns to predict the noise situation at each time step, so as to subsequently predict denoising from the noise data to obtain the final clear video frame. In the main network and the conditional control network, based on the 2D Unet used in the SD model, a temporal modeling module is additionally added to learn the temporal-related motion information in video generation.

[0323] In actual implementation, during the training process, video-text training data is first input, and the training videos are composed into a tensor of B*T*C*H*W (denoted as z0), where B is the training batch size, T is the number of frames of the generated video, and C, H, W are the channels, height, and width of the video frame; among them, in some latent variable-based diffusion models, C, H, W represent the channels, height, and width of the latent variable; then noise is input, and according to the common noise addition strategy and the number of noise addition steps t, the input video tensor is noise-added and sent into the model, that is

[0324] z t = α t z0 + δ t ε... Equation (1);

[0325] where ∈ is Gaussian noise, and α, σ are noise addition coefficients.

[0326] Then video spatio-temporal domain modeling is performed. Specifically, the noisy input z t enters the main network (denoted as f θ ). After that, the conv convolution in each 3D residual block and the transformer module in the Attention block will perform modeling learning on the input space and time series. Then conditional modeling is carried out. Common encoding modules such as CLIP are used to extract the embedding features of the text, and then attention calculation is performed in the cross attention module in the Attention block with the original input to guide the noise prediction of the 3D Unet; the image, video, and skeleton node information is conditionally encoded through the conditional control module, and then the output of different 3D residual modules is controlled by the condition and spliced with the information in the main network to guide the noise prediction of the 3D Unet. Finally, noise prediction is performed. Specifically, after the noisy input z t and various conditional information are input into the 3D Unet of the present application, the noise ∈ added at the noise addition step t is predicted θ, the loss function is the MSE Loss with the input real noise ∈, and the network parameters are continuously updated by minimizing the loss, that is

[0327] ε θ = f θ (z t , t, c) …… Formula (2);

[0328]

[0329] where L is the loss function.

[0330] In actual implementation, after the model training is completed, sampling can be performed. The sampling process uses the existing conditional information (text, image, video, skeleton, etc.) to continuously predict the noise to be removed in the random noise; the random noise is z t = B * T * C * H * W, and the input form of the condition c is the same as that in the training stage;

[0331] Then, denoising is performed to finally generate a video. Specifically, the constructed z t , that is, the condition c, is input into the trained 3D Unet, and the noise is removed through multiple rounds of predicted noise, and the final clear noise-free video frames are obtained. Then, the final video is obtained by combining them at a certain frame rate.

[0332] In actual implementation, for the above high-order gameplay, that is, the video AI editing model, the main structure of the video AI editing model is similar to the video general model described above, but there are some differences in the input of the conditional control module. Referring to the product interface, video AI editing requires the user to upload a video of their own or the generated video, and then specify the position of the local area to be edited. The local area generally specifies the upper left and lower right coordinates of a rectangular box.

[0333] In actual implementation, during the training and sampling processes, refer to Figure 29 , Figure 29 is a schematic diagram of the sampling process provided by the embodiment of the present application. Based on Figure 29 , the conditional control encoding module will input the 0-1 binary map obtained by covering the area specified by the user in the original video (filled with 0) and the Gaussian noise after connecting them according to the channel position into the conditional control module to guide the learning of the main model. Then, the remaining training and sampling processes can refer to the above description.

[0334] For the skeletal-guided dance video model in the above high-order gameplay, that is, the process of generating videos driven by skeletal points, the skeletal-guided dance video model adds two types of conditional control modules to the general video model. The first type is the conditional control module for the skeletal sequence, and the second type is the conditional control module for the given image. The skeletal sequence is several dancing skeletal templates provided by the platform, and the given image is an image containing a human body uploaded by the user.

[0335] See Figure 30 , Figure 30 is a schematic diagram of the processing process of the skeletal-guided dance video model provided by the embodiments of the present application. Based on Figure 30 , the skeletal-guided dance video model supports two functions. The first function is to generate videos by adding text and skeletal information. The text prompt here is mainly a description of the generated video characters and background, and the actions will be generated according to the skeletal sequence; the skeletal information uses a lightweight convolutional network for conditional control, and the text is directly encoded using the text encoder of the general model. The second function is to generate videos by adding pictures and skeletal information, and the appearance of the characters in the generated video is the same as the given picture; the picture information is introduced in two forms, the control network and the image adapter, and then the rest of the training and sampling processes can refer to the foregoing.

[0336] Thus, the present application proposes a relatively general 3D video generation model based on various conditional controls. It not only ensures that during the video generation process, it can well learn the modeling of spatial and temporal motion (ensuring clear video quality, reasonable motion, and no distortion), but also can uniformly use various conditional information (text, image, video, skeleton, etc.) to guide the final video generation process.

[0337] Applying the above embodiments of the present application, first display at least one video generation control for generating videos of different video types, and then based on the at least one video generation control, in response to the target video type in at least two video types being in a selected state, display a content input area corresponding to the target video type, and based on the content input area, input target content, so as to generate a target video of the video type corresponding to the target video generation control based on the input target content. In this way, not only is a target video of the target video type corresponding to the target content input by the user generated, so that the video content of the target video meets the user's needs, but also, based on the video generation control, a target video of the target video type is generated, so that the type of the video also meets the user's needs. Compared with the solution where the user still needs to edit the video after generating the video, the present application reduces the editing operation that the user needs to perform on the generated video, not only increasing the human-computer interaction efficiency, but also improving the video generation efficiency.

[0338] The following continues to describe the exemplary structure of the software module implementation of the video generation device 455 provided in the embodiments of the present application. In some embodiments, as Figure 2 shown, the software module in the video generation device 455 stored in the memory 450 may include:

[0339] A first display module 4551, configured to display a video generation interface, where the video generation interface includes at least two video generation controls, and different video generation controls are used to generate videos of different video types;

[0340] A second display module 4552, configured to display a content input area in response to a target video generation control among the at least two video generation controls being in a selected state;

[0341] A third display module 4553, configured to display the input target content in response to an input operation on the content input area;

[0342] A fourth display module 4554, configured to display a target video generated based on the target content in response to a video generation instruction, where the type of the target video is the video type corresponding to the target video generation control.

[0343] In some embodiments, the number of the video generation controls is multiple, and different video generation controls correspond to different video types. The first display module 4551 is further configured to display a video generation interface including multiple video generation controls in response to an open instruction for the video generation interface; the device further includes a determination module, and the determination module is configured to determine that the target video type is in a selected state in response to a video generation control corresponding to the target video type in the video generation interface being in a selected state.

[0344] In some embodiments, the device further includes a parameter setting module. The parameter setting module is further configured to display a parameter setting area in the video generation interface, where the parameter setting area is used to set parameters of the generated video; based on the parameter setting area, in response to a parameter setting operation, display the video parameters set for the video of the target video type; the fourth display module 4554 is further configured to display the video of the target video type with the video parameters generated based on the target content in response to a video generation instruction.

[0345] In some embodiments, the target video type is a long video type, and the videos of the long video type include multiple storyboards; the apparatus further includes a fifth display module, configured to display a storyboard disassembly interface in response to a determination instruction for the input target content, and display multiple storyboard texts determined based on the target content in the storyboard disassembly interface; wherein each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; the fourth display module 4554 is further configured to display a video of the target video type including storyboards corresponding to the respective storyboard texts in response to a video generation instruction triggered based on the multiple storyboard texts.

[0346] In some embodiments, the apparatus further includes a storyboard text quantity adjustment module, configured to display a storyboard text quantity adjustment control, where the storyboard text quantity adjustment control is used to add a new storyboard text; and in response to a storyboard text addition operation triggered based on the storyboard text quantity adjustment control, display the added new storyboard text in the storyboard disassembly interface.

[0347] In some embodiments, a first determination control is further displayed in the storyboard disassembly interface, and the apparatus further includes a sixth display module, configured to display a storyboard generation interface in response to a trigger operation for the first determination control; in the storyboard generation interface, display the storyboards corresponding to each storyboard text, and display a second determination control; and in response to a trigger operation for the second determination control, receive the video generation instruction.

[0348] In some embodiments, each storyboard is associated with a parameter setting control, where the parameter setting control is used to set the storyboard parameters of the corresponding storyboard; the apparatus further includes a storyboard parameter setting module, configured to display the target storyboard parameters set for the target storyboard in response to a parameter setting operation triggered based on the parameter setting control associated with the target storyboard; and in response to a determination instruction for the target storyboard parameters, update the storyboard parameters of the target storyboard to the target storyboard parameters.

[0349] In some embodiments, the target video type is a long video type, and the videos of the long video type include multiple storyboards; the apparatus further includes a seventh display module, configured to display a storyboard generation interface in response to a determination instruction for the input target content, and display multiple storyboard texts and multiple storyboards in the storyboard generation interface, where each storyboard text corresponds to a storyboard, and the storyboard is generated based on the corresponding storyboard text; the fourth display module 4554 is further configured to display a video of the target video type including the multiple storyboards in response to a video generation instruction based on the multiple storyboard texts and the multiple storyboards.

[0350] In some embodiments, the device further includes a selection module, configured to control the first target sub-shot among the multiple sub-shots to be in a selected state in response to a selection operation for the first target sub-shot among the multiple sub-shots; the fourth display module 4554 is further configured to display a video of a target video type including the first target sub-shot in response to a video generation instruction based on the first target sub-shot in the selected state and the sub-shot text of the first target sub-shot.

[0351] In some embodiments, the multiple sub-shots have an arrangement order, and the videos formed by the multiple sub-shots with different arrangement orders are different. A sorting adjustment control is further displayed in the sub-shot generation interface; the device further includes a sorting adjustment module, configured to change the sorting of the multiple sub-shots from the current arrangement order to a target arrangement order in response to a sorting adjustment operation for the multiple sub-shots triggered based on the sorting adjustment control; the fourth display module 4554 is further configured to display a video of a target video type including the multiple sub-shots generated based on the target arrangement order in response to a video generation instruction based on the multiple sub-shot texts and the multiple sub-shots.

[0352] In some embodiments, a sub-shot update control for updating the content of the sub-shot is further displayed in the sub-shot generation interface; the device further includes a sub-shot update module, configured to update the second target sub-shot to a first new sub-shot based on the sub-shot update control in response to a content update operation for the second target sub-shot among the multiple sub-shots, and the content of the first new sub-shot is associated with the content of the second target sub-shot.

[0353] In some embodiments, the multiple sub-shots have an arrangement order, and there is an upward association control for the third target sub-shot among the multiple sub-shots, and the upward association control is configured to update the content of the third target sub-shot based on the content of the previous sub-shot of the third target sub-shot. The third target sub-shot is any one of the multiple sub-shots except the first sub-shot; the device further includes an association module, configured to update the third target sub-shot to a second new sub-shot in response to a trigger operation for the upward association control.

[0354] In some embodiments, the device further includes an introduction module, configured to display function introduction information if the long video type of video is generated for the first time by the current account; wherein, the function introduction information is used to introduce the function of the upward association control.

[0355] In some embodiments, the seventh display module is further configured to display a plurality of storyboard texts on the storyboard generation interface and display a plurality of candidate storyboards corresponding to each of the storyboard texts; for each of the storyboard texts, in response to a selection operation on a target candidate storyboard among the plurality of candidate storyboards, control the target candidate storyboard to be in a selected state, and determine the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

[0356] In some embodiments, the apparatus further includes an eighth display module, where the eighth display module is configured to display at least one video template; in response to a selection operation on a target video template among the at least one video template, display a template content description of the target video template and a corresponding first editing control; in response to a triggering operation on the first editing control, determine the triggering operation as an input operation for the content input area; the third display module 4553 is further configured to, in response to a triggering operation on the first editing control, determine the triggering operation as the content input operation, display the template content description of the target video template in the content input area, and determine the displayed template content description as the input target content.

[0357] In some embodiments, the apparatus further includes a ninth display module, where the ninth display module is configured to, in response to a triggering operation on a target video template among the at least one video template, display a details interface including details information of the target video template; where the details interface further includes at least one of a second editing control and a first new video generation control, the second editing control is configured to regenerate a video of the target video type based on the template content description, and the first new video generation control is configured to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

[0358] In some embodiments, the fourth display module 4554 is further configured to, in response to the video generation instruction, display a video of the target video type generated based on the target content on a video display interface; where the video display interface further includes at least one of the following: details information of the generated video, a third editing control, and a first new video generation control; where the third editing control is configured to regenerate a video of the target video type based on the target content, and the first new video generation control is configured to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

[0359] In some embodiments, the video display interface includes the third editing control, the target video type is a long video type, and the long video includes multiple scenes; the device further includes a tenth display module, which is configured to respond to a trigger operation on the third editing control, display a scene generation interface, and display multiple scene texts and multiple scenes determined based on the target content on the scene generation interface; based on the scene generation interface, in response to an editing operation on the multiple scene texts and multiple scenes, display the edited multiple scene texts and the multiple scenes; in response to a video generation instruction, display a long video type video generated based on the edited multiple scene texts and multiple scenes.

[0360] In some embodiments, the video display interface includes the third editing control, and the target video type is a short video type; the device further includes an eleventh display module, which is configured to respond to a trigger operation on the third editing control, display a video generation interface including a content input area, and the target content is displayed in the content input area; in response to an editing operation on the target content, display the new target content obtained by editing the target content; in response to a video generation instruction, display a short video type video generated based on the new target content.

[0361] In some embodiments, the fourth display module 4554 is further configured to respond to a video generation instruction and display a first progress prompt message, where the first progress prompt message is used to prompt the generation progress of the target video type video; when the first progress prompt message indicates that the target video type video generation is completed, cancel the display of the progress prompt message and display the target video type video generated based on the target content.

[0362] In some embodiments, the device further includes a video generation prompt module, which is configured to display a video generation prompt message, and the video generation prompt message is used to prompt the number of videos that can be generated within a target time; the fourth display module 4554 is further configured to respond to a video generation instruction, and when the video generation prompt message indicates that the number of videos that can be generated within the target time is not zero, display the target video type video generated based on the target content.

[0363] In some embodiments, the device further includes a twelfth display module, configured to display a sharing link of other received videos; in response to a display instruction triggered based on the sharing link, if the current account has the playback permission for the other videos, display a playback interface of the other videos, and display detailed information of the other videos and a first new video generation control on the playback interface; the first display module 4551 is further configured to display the video generation interface in response to a trigger operation on the first new video generation control.

[0364] In some embodiments, the video generation interface further includes a video display control, configured to display at least one asset video, where the at least one asset video includes at least one of a successfully generated video that was successfully generated in the past, a failed video that failed to be generated in the past, a video being generated that is currently being generated, and a video waiting to be generated that is waiting to be generated; the device further includes an asset video display module, configured to display at least one asset video including videos of the target video type in response to a trigger operation on the video display control.

[0365] In some embodiments, the asset video includes a failed video, and the failed video is associated with a regeneration control, configured to regenerate the failed video; the device further includes a regeneration module, configured to regenerate the failed video in response to a trigger operation on the regeneration control associated with the target failed video; when the regeneration is successful, switch the failed video in the at least one asset video to the regenerated video; when the regeneration fails, display a regeneration failure prompt message.

[0366] In some embodiments, the at least one asset video includes a video waiting to be generated, and the device further includes a waiting prompt module, configured to display a first waiting prompt message in an associated area of the video waiting to be generated, where the first waiting prompt message is used to indicate that the corresponding video is waiting to be generated; in response to a trigger operation on the video waiting to be generated, display a waiting cancellation control, configured to terminate the waiting process of the video waiting to be generated.

[0367] In some embodiments, the video generation interface further includes a style setting control for setting the video style of the generated video. The device further includes a style setting module configured to, in response to a trigger operation on the style setting control, display at least one video style; in response to a selection operation on a target video style among the at least one video style, control the target video style to be in a selected state; and the fourth display module 4554 is further configured to, in response to a video generation instruction, display a video of a target video type having the target video style generated based on the target content.

[0368] In some embodiments, the device further includes a style identifier display module configured to, in response to a determination instruction on the target video style in a selected state, display an identifier of the target video style on the style setting control, where the identifier is associated with a style cancellation control; wherein the identifier of the target video style is used to identify that the generated video of the target video type has the target video style; in response to a trigger operation on the style cancellation control, cancel the display of the identifier of the target video style on the style setting control; and the fourth display module 4554 is further configured to, in response to a video generation instruction, display a video of a target video type having a default style generated based on the target content.

[0369] In some embodiments, the fourth display module 4554 is further configured to, in response to a video generation instruction, display a second waiting prompt message for indicating the duration to wait until the generation of the video of the target video type starts; when the waiting prompt message indicates that the duration to wait until the generation of the video of the target video type starts is zero, cancel the display of the second waiting prompt message and display the video of the target video type generated based on the target content.

[0370] In some embodiments, the device further includes an interaction module, which is configured to display at least one interaction control for the generated video. The interaction control is one of the following controls: a local editing control for locally editing the generated video, a canvas expansion control for editing the size of the generated video, a motion brush control for changing static objects in the generated video into dynamic objects, a shooting parameter setting control for setting shooting parameters of the generated video, a background removal control for removing the background of the generated video, an object erasing control for erasing target objects in the generated video, a scene detection control for splitting the video into different video segments based on different scenes in the generated video, a depth of field setting control for setting a depth of field effect for the generated video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the generated video, an action sequence generation control for setting an action sequence for an object in the generated video, and a second new video generation control for generating a new video based on the audio data in the generated video; in response to a trigger operation on a target interaction control among at least one of the interaction controls, an interaction operation indicated by the target interaction control is performed on the video of the generated target video type.

[0371] An embodiment of the present application provides a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the video generation method described above in the embodiments of the present application.

[0372] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where the computer-executable instructions, when executed by a processor, cause the processor to execute the video generation method provided in the embodiments of the present application. For example, as Figure 3 shown in the video generation method.

[0373] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a CD-ROM, etc.; it may also be various devices including one or any combination of the above memories.

[0374] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0375] As an example, the computer-executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program in question, or, stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).

[0376] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected via a communication network.

[0377] In summary, the embodiments of the present application have the following beneficial effects:

[0378] (1) Not only is a video of a target video type corresponding to the target content input by the user generated, such that the video content of the target video meets the user's needs, but also, based on the video generation control, a video of the target video type is generated, such that the type of the video also meets the user's needs. Compared with the solution where the user still needs to edit the video after it is generated, the present application reduces the editing operations that the user needs to perform on the generated video, not only increasing the human-computer interaction efficiency, but also improving the video generation efficiency.

[0379] (2) A relatively general 3D video generation model based on various conditional controls is proposed. Not only is it ensured that during the video generation process, the modeling of space and temporal motion can be well learned (ensuring clear video image quality, reasonable motion, and no distortion); at the same time, various conditional information (text, image, video, skeleton, etc.) can be uniformly used to guide the final video generation process.

[0380] (3) Based on the upward association control, the relevance of the content and effects between storyboards is enhanced.

[0381] It should be noted that in the embodiments of the present application, relevant data such as obtaining the operation data of users and the target content input by users are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0382] As described above, the above are only embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. A video generation method, characterized in that, The method includes: Displaying a video generation interface, the video generation interface including at least one video generation control for generating videos of at least two video types; Based on the at least one video generation control, in response to a target video type among the at least two video types being in a selected state, displaying a content input area corresponding to the target video type; In response to a content input operation in the content input area, displaying the input target content; In response to a video generation instruction, displaying a video of the target video type generated based on the target content.

2. The method according to claim 1, wherein The number of the video generation controls is multiple, and different video generation controls correspond to different video types. The displaying of the video generation interface includes: In response to an open instruction for the video generation interface, displaying a video generation interface including multiple video generation controls; The method further includes: In response to a video generation control corresponding to the target video type in the video generation interface being in a selected state, determining that the target video type is in a selected state.

3. The method according to claim 1, characterized in that The method further includes: Displaying a parameter setting area in the video generation interface, the parameter setting area being used for setting parameters of the generated video; Based on the parameter setting area, in response to a parameter setting operation, displaying the video parameters set for the video of the target video type; The in response to a video generation instruction, displaying a video of the target video type generated based on the target content includes: In response to a video generation instruction, based on the target content, displaying a video of the target video type with the video parameters.

4. The method according to claim 1, wherein The target video type is a long video type, and the video of the long video type includes multiple storyboards; After the in response to a content input operation in the content input area, displaying the input target content, the method further includes: In response to a determination instruction for the input target content, displaying a storyboard disassembly interface and displaying multiple storyboard texts determined based on the target content in the storyboard disassembly interface; wherein each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; The in response to a video generation instruction, displaying a video of the target video type generated based on the target content includes: In response to a video generation instruction triggered based on the multiple storyboard texts, displaying a video of the target video type generated including storyboards corresponding to the respective storyboard texts.

5. The method according to claim 4, wherein The method further includes: Displaying a storyboard text quantity adjustment control for adding a new storyboard text; In response to a storyboard text addition operation triggered based on the storyboard text quantity adjustment control, displaying the added new storyboard text in the storyboard disassembly interface.

6. The method according to claim 1, characterized in that The target video type is a long video type, and the video of the long video type includes multiple storyboards; After the in response to an input operation for the content input area, displaying the input target content, the method further includes: In response to a determination instruction for the input target content, a storyboard generation interface is displayed, and a plurality of storyboard texts and a plurality of storyboards are displayed on the storyboard generation interface. Each of the storyboard texts corresponds to a storyboard, and the storyboard is generated based on the corresponding storyboard text. In response to a video generation instruction, a video of a target video type generated based on the target content is displayed, including: Based on the plurality of storyboard texts and the plurality of storyboards, in response to a video generation instruction, a video of a target video type including the plurality of storyboards is displayed.

7. The method according to claim 6, wherein After displaying the plurality of storyboard texts and the plurality of storyboards on the storyboard generation interface, the method further includes: In response to a selection operation for a first target storyboard among the plurality of storyboards, controlling the first target storyboard to be in a selected state. Based on the plurality of storyboard texts and the plurality of storyboards, in response to a video generation instruction, a video of a target video type including the plurality of storyboards is displayed, including: Based on the first target storyboard in the selected state and the storyboard text of the first target storyboard, in response to a video generation instruction, a video of a target video type including the first target storyboard is displayed.

8. The method according to claim 6, wherein The plurality of storyboards have an arrangement order, and videos formed by the plurality of storyboards with different arrangement orders are different. A sorting adjustment control is also displayed on the storyboard generation interface. After displaying the plurality of storyboard texts and the plurality of storyboards on the storyboard generation interface, the method further includes: In response to a sorting adjustment operation for the plurality of storyboards triggered based on the sorting adjustment control, changing the sorting of the plurality of storyboards from the current arrangement order to a target arrangement order. Based on the plurality of storyboard texts and the plurality of storyboards, in response to a video generation instruction, a video of a target video type including the plurality of storyboards is displayed, including: Based on the plurality of storyboard texts and the plurality of storyboards, in response to a video generation instruction, a video of a target video type including the plurality of storyboards generated based on the target arrangement order is displayed.

9. The method according to claim 6, wherein A storyboard update control for updating the content of the storyboard is also displayed on the storyboard generation interface. After displaying the plurality of storyboard texts and the plurality of storyboards on the storyboard generation interface, the method further includes: Based on the storyboard update control, in response to a content update operation for a second target storyboard among the plurality of storyboards, updating the second target storyboard to a first new storyboard, and the content of the first new storyboard is associated with the content of the second target storyboard.

10. The method according to claim 6, characterized in that, The plurality of storyboards have an arrangement order, and there is an upward association control for a third target storyboard among the plurality of storyboards. The upward association control is used to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard. The third target storyboard is any one of the plurality of storyboards except the first storyboard. After displaying the plurality of storyboard texts and the plurality of storyboards on the storyboard generation interface, the method further includes: In response to a trigger operation for the upward association control, updating the third target storyboard to a second new storyboard.

11. The method according to claim 10, wherein Before updating the third target storyboard to the second new storyboard for the trigger operation on the upward associated control, the method further includes: When the current account generates a video of the long video type for the first time, display function introduction information; Among them, the function introduction information is used to introduce the function of the upward associated control.

12. The method according to claim 6, wherein, The displaying of multiple storyboard texts and multiple storyboards in the storyboard generation interface includes: Display multiple storyboard texts in the storyboard generation interface, and display multiple candidate storyboards corresponding to each of the storyboard texts; For each of the storyboard texts, in response to a selection operation on a target candidate storyboard among the multiple candidate storyboards, control the target candidate storyboard to be in a selected state, and determine the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

13. The method according to claim 1, wherein The method further includes: Display at least one video template; In response to a selection operation on a target video template among the at least one video template, display a template content description of the target video template and a corresponding first editing control; In response to a trigger operation on the first editing control, determine the trigger operation as an input operation for the content input area; In response to a content input operation in the content input area, display the input target content, including: In response to a trigger operation on the first editing control, determine the trigger operation as the content input operation, in the content input area, display the template content description of the target video template, and determine the displayed template content description as the input target content.

14. The method according to claim 13, wherein After displaying the at least one video template, the method further includes: In response to a trigger operation on a target video template among the at least one video template, display a details interface including details information of the target video template; Among them, the details interface further includes at least one of a second editing control and a first new video generation control. The second editing control is used to generate a video of the target video type again based on the template content description, and the first new video generation control is used to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

15. The method according to claim 1, wherein The displaying of a video of the target video type generated based on the target content in response to a video generation instruction includes: In response to the video generation instruction, display a video of the target video type generated based on the target content on a video display interface; Among them, the video display interface further includes at least one of the following: details information of the generated video, a third editing control, and a first new video generation control; Among them, the third editing control is used to generate a video of the target video type again based on the target content, and the first new video generation control is used to re-enter target content in the content input area and generate a video of the target video type based on the re-entered target content.

16. The method according to claim 15, characterized in that, The video display interface includes the third editing control, the target video type is a long video type, and the long video type of video includes multiple storyboards; After displaying a video of the target video type generated based on the target content on the video display interface in response to the video generation instruction, the method further includes: In response to a trigger operation on the third editing control, display a storyboard generation interface, and display a plurality of storyboard texts and a plurality of storyboards determined based on the target content on the storyboard generation interface; Based on the storyboard generation interface, in response to an editing operation on the plurality of storyboard texts and the plurality of storyboards, display the edited plurality of storyboard texts and the plurality of storyboards; In response to the video generation instruction, display a long video type video generated based on the edited plurality of storyboard texts and the plurality of storyboards.

17. The method according to claim 15, characterized in that, The video display interface includes the third editing control, and the target video type is a short video type; after displaying a video of the target video type generated based on the target content on the video display interface in response to the video generation instruction, the method further includes: In response to a trigger operation on the third editing control, display a video generation interface including a content input area, and the target content is displayed in the content input area; In response to an editing operation on the target content, display a new target content obtained by editing the target content; In response to the video generation instruction, display a short video type video generated based on the new target content.

18. The method according to claim 1, wherein The video generation interface further includes a video display control for displaying at least one asset video, and the at least one asset video includes at least one of a successful video that was successfully generated historically, a failed video that failed to be generated historically, a generating video that is currently being generated, and a waiting video that is waiting to be generated; After displaying a video of the target video type generated based on the target content on the video display interface in response to the video generation instruction, the method further includes: In response to a trigger operation on the video display control, display at least one asset video including a video of the target video type.

19. The method according to claim 18, wherein The asset video includes a failed video, and the failed video is associated with a regeneration control for regenerating the failed video; After displaying at least one asset video including a video of the target video type in response to a trigger operation on the video display control, the method further includes: In response to a trigger operation on the regeneration control associated with the target failed video, regenerate the failed video; When the regeneration is successful, switch the failed video in the at least one asset video to the regenerated video; When the regeneration fails, display a regeneration failure prompt message.

20. The method according to claim 1, characterized in that, The video generation interface further includes a style setting control for setting the video style of the generated video; After displaying the video generation interface, the method further includes: In response to a trigger operation on the style setting control, display at least one video style; In response to a selection operation on a target video style in the at least one video style, control the target video style to be in a selected state; In response to a video generation instruction, displaying a video of a target video type generated based on the target content, including: In response to a video generation instruction, displaying a video of a target video type having the target video style generated based on the target content.

21. The method according to claim 1, wherein After, in response to a video generation instruction, displaying a video of a target video type generated based on the target content, the method further includes: Displaying at least one interaction control for the generated video, where the interaction control is one of the following controls: a local editing control for locally editing the generated video, a canvas expansion control for editing the size of the generated video, a motion brush control for changing static objects in the generated video into dynamic objects, a camera movement parameter setting control for setting the camera movement parameters of the generated video, a background removal control for removing the background of the generated video, an object erasing control for erasing a target object in the generated video, a scene detection control for splitting the video into different video segments based on different scenes in the generated video, a depth of field setting control for setting a depth of field effect for the generated video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the generated video, an action sequence generation control for setting an action sequence for an object in the generated video, and a second new video generation control for generating a new video based on the audio data in the generated video; In response to a trigger operation on a target interaction control among at least one of the interaction controls, performing an interaction operation indicated by the target interaction control on the generated video of the target video type.

22. A video generation device, characterized in that, The apparatus includes: A first display module, configured to display a video generation interface, where the video generation interface includes at least one video generation control for generating videos of at least two video types; A second display module, configured to, based on the at least one video generation control, in response to a target video type among the at least two video types being in a selected state, display a content input area corresponding to the target video type; A third display module, configured to, in response to a content input operation in the content input area, display the input target content; A fourth display module, configured to, in response to a video generation instruction, display a video of the target video type generated based on the target content.

23. An electronic device, characterized in that, Including: A memory, configured to store computer-executable instructions; A processor, configured to, when executing the computer-executable instructions stored in the memory, implement the video generation method according to any one of claims 1 to 21.

24. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored and used to cause the processor to implement the video generation method according to any one of claims 1 to 21 when executed.

25. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the processor, the video generation method according to any one of claims 1 to 21 is implemented.

Citation Information

Cited By

  • Media content processing method and device, equipment, storage medium and program product

    CN121284345A

  • Multi-modal content generation method and device, intelligent agent and electronic equipment

    CN121541807A