Video generation method, apparatus, electronic device, storage medium, and program product

By displaying video generation controls in the video generation interface and responding to user selection of the target video type, the problem of low video generation efficiency in the prior art is solved, and efficient human-computer interaction and hardware resource utilization are achieved.

WO2025152644A1PCT designated stage expired Publication Date: 2025-07-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2024/137383
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2024-12-06
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The prior art fails to effectively combine the user's demand for video types when generating videos, resulting in low video generation efficiency, low human-computer interaction efficiency and low hardware processing resource utilization.

Method used

A video generation method is provided, by displaying a video generation control in a video generation interface, displaying a corresponding content input area in response to a video type selected by the user, and generating a video of a target video type based on the input content.

Benefits of technology

It improves video generation efficiency, improves human-computer interaction efficiency and hardware resource utilization, and meets users' dual needs for video content and types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137383_24072025_PF_FP_ABST
    Figure CN2024137383_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a video generation method, an apparatus, an electronic device, a computer-readable storage medium and a computer program product. The method comprises: in a video generation interface, displaying at least one video generation control, the at least one video generation control being used for generating videos of at least two video types; on the basis of the at least one video generation control and in response to a target video type among the at least two video types being in a selected state, displaying a content input area corresponding to the target video type; in response to a content input operation executed in the content input area, displaying an input target content; and, in response to a video generation instruction, displaying a first video of the target video type generated on the basis of the target content.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method, device, electronic device, storage medium, and program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The embodiments of this application are based on and claim the priority of Chinese patent application with application number 2024100580053 and application date January 15, 2024. The entire contents of the Chinese patent application are hereby introduced into the embodiments of this application as a reference. Technical Field

[0003] The present application relates to the field of Internet technology, and in particular to a video generation method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0004] There are many types of videos, such as short videos and long videos. Different types of videos have different corresponding video parameters and different video generation processes. However, when generating videos, related technologies mostly directly generate videos based on the user's demand for video content, that is, they only consider the user's demand for video content, but ignore the user's demand for video type. In this way, the user also needs to edit the video to generate the final video, so that the final video meets the user's requirements for video content and the user's requirements for video type. In this way, not only the video generation efficiency is low, but also the human-computer interaction efficiency and the utilization rate of hardware processing resources are low. Summary of the Invention

[0005] The embodiments of the present application provide a video generation method, device, electronic device, computer-readable storage medium, and computer program product, which can improve video generation efficiency, human-computer interaction efficiency, and utilization of hardware processing resources.

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] The present invention provides a method for generating a video, including:

[0008] In the video generation interface, at least one video generation control is displayed, wherein the at least one video generation control is used to generate videos of at least two video types;

[0009] Generating a control based on the at least one video, and in response to a target video type being selected from the at least two video types, displaying a content input area corresponding to the target video type;

[0010] In response to a content input operation performed in the content input area, displaying the input target content;

[0011] In response to a video generation instruction, a first video of the target video type generated based on the target content is displayed.

[0012] The present invention provides a video generation device, including:

[0013] A first display module is configured to display at least one video generation control in a video generation interface, wherein the at least one video generation control is used to generate videos of at least two video types;

[0014] a second display module configured to generate a control based on the at least one video, and in response to a target video type being selected from the at least two video types, display a content input area corresponding to the target video type;

[0015] a third display module configured to display the input target content in response to a content input operation performed in the content input area;

[0016] The fourth display module is configured to display a first video of the target video type generated based on the target content in response to a video generation instruction.

[0017] An embodiment of the present application provides an electronic device, including:

[0018] a memory for storing computer-executable instructions;

[0019] The processor is configured to implement the video generation method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute the instructions to implement the video generation method provided in the embodiment of the present application.

[0021] The present invention provides a computer program product including computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the video generation method provided in the present invention.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] First, at least one video generation control for generating videos of different video types is displayed. Then, based on the at least one video generation control, in response to a target video type being selected from at least two video types, a content input area corresponding to the target video type is displayed, and based on the content input area, target content is input, thereby generating a first video of the target video type based on the input target content. In this way, not only is a video of the target video type corresponding to the target content generated based on the target content input by the user, so that the video content of the target video meets the user's needs, but also, based on the video generation control, a video of the target video type is generated, so that the type of the video also meets the user's needs. Compared with a solution in which the user needs to edit the video to generate the final video, the present application reduces the editing operations that the user needs to perform on the video, which not only improves the efficiency of video generation, but also improves the efficiency of human-computer interaction and the hardware resource utilization of the electronic device. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG1 is a schematic diagram of the architecture of a video generation system 100 provided in an embodiment of the present application;

[0025] FIG2 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0026] FIG3 is a flow chart of a video generation method according to an embodiment of the present application;

[0027] FIG4 is a schematic diagram of a display content input area provided in an embodiment of the present application;

[0028] FIG5 is a schematic diagram of guidance information provided in an embodiment of the present application;

[0029] FIG6 is a schematic diagram of another video playback interface provided in an embodiment of the present application;

[0030] FIG7 is a schematic diagram of permission prompt information provided in this embodiment;

[0031] FIG8 is a schematic diagram of at least one video template provided in an embodiment of the present application;

[0032] FIG9 is a schematic diagram of a details interface including detailed information of a target video template provided in an embodiment of the present application;

[0033] FIG10 is a schematic diagram of a first determination control provided in an embodiment of the present application;

[0034] FIG11 is a schematic diagram of a display parameter setting area provided in an embodiment of the present application;

[0035] FIG12 is a schematic diagram of a mirror disassembly interface provided in an embodiment of the present application;

[0036] FIG13 is a schematic diagram of a storyboard generation interface provided in an embodiment of the present application;

[0037] FIG14 is a schematic diagram of function introduction information provided in an embodiment of the present application;

[0038] FIG15 is a schematic diagram of at least two candidate storyboards provided in an embodiment of the present application;

[0039] FIG16 is a schematic diagram of a video display interface for a long video provided in an embodiment of the present application;

[0040] FIG17 is a schematic diagram of a video display interface for a short video provided in an embodiment of the present application;

[0041] FIG18 is a schematic diagram of first progress prompt information provided in an embodiment of the present application;

[0042] FIG19 is a schematic diagram of a second waiting prompt information provided in an embodiment of the present application;

[0043] FIG20 is a schematic diagram of an asset video display interface provided in an embodiment of the present application;

[0044] FIG21 is a schematic diagram of deletion prompt information provided in an embodiment of the present application;

[0045] FIG22 is a schematic diagram of generation failure prompt information, regeneration control, and backtracking control provided in an embodiment of the present application;

[0046] FIG23 is a schematic diagram of at least one video style provided by an embodiment of the present application;

[0047] FIG24 is a schematic diagram of identifying a target video style according to an embodiment of the present application;

[0048] FIG25 is a schematic diagram of an operation control provided in an embodiment of the present application;

[0049] FIG26 is a schematic diagram of interacting with a first video based on a local editing control according to an embodiment of the present application;

[0050] FIG27 is a schematic diagram of interacting with a first video based on a canvas extension control provided by an embodiment of the present application;

[0051] FIG28 is a schematic diagram of the structure of a video generation model provided in an embodiment of the present application;

[0052] FIG29 is a schematic diagram of a sampling process provided in an embodiment of the present application;

[0053] Figure 30 is a schematic diagram of the processing process of the skeleton-guided dance video model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting the embodiments of the present application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0055] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0056] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0058] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0059] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0060] 2) Client, also known as user end, refers to the program corresponding to the server that provides local services to users. Except for some applications that can only run locally, it is generally installed on the terminal and needs to cooperate with the server to run. That is, there must be corresponding servers and service programs in the network to provide corresponding services. In this way, a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application, such as the virtual scene client (such as the game client).

[0061] 3) Storyboard, also known as shot-by-shot script, is used to indicate a series of pictures that are broken down into a script or video in chronological order. Each picture, or each storyboard, represents a shot, including the shot's field of view, angle, movement, etc.

[0062] 4) Shot movement, also known as motion shot, is used to indicate the shooting method performed by moving the camera position, changing the lens optical axis, or changing the lens focal length in a shot. The pictures captured by this shooting method are called moving pictures. For example: push shots, pull shots, pan shots, move shots, follow shots, lift shots and comprehensive motion shots formed by push, pull, shake, move, follow shots, lift shots and comprehensive motion shots, etc.

[0063] Refer to Figure 1, which is a schematic diagram of the architecture of the video generation system 100 provided in an embodiment of the present application, a terminal (terminal 400 is shown as an example), and the terminal 400 is connected to the server 200 through the network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two, using wireless or wired links to realize data transmission.

[0064] The server 200 is configured to send interface data corresponding to a video generation interface including at least one video generation control to the terminal 400;

[0065] Terminal 400 is used to receive interface data corresponding to a video generation interface including at least one video generation control; based on the interface data, display the video generation interface; display at least one video generation control in the video generation interface, and the at least one video generation control is used to generate videos of at least two video types; based on the at least one video generation control, in response to a target video type being selected from at least two video types, display a content input area corresponding to the target video type; in response to a content input operation performed in the content input area, display the input target content; in response to a video generation instruction, display a first video of the target video type generated based on the target content.

[0066] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, an intelligent voice interaction device, a smart home appliance, a virtual reality device, a vehicle-mounted terminal, an aircraft, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device, an intelligent speaker, and a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0067] Next, an electronic device for implementing the video generation method provided by an embodiment of the present application is described. Referring to Figure 2, Figure 2 is a structural diagram of an electronic device provided by an embodiment of the present application. The electronic device can be a server or a terminal. Taking the electronic device as the terminal shown in Figure 1 as an example, the electronic device shown in Figure 2 includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together through a bus system 440. It can be understood that the bus system 440 is configured to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 440 in Figure 2.

[0068] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0069] The user interface 430 includes one or more output devices 431 that enable display of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0070] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0071] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0072] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0073] Operating system 451, including system programs configured to handle various basic system services and perform hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0074] A network communication module 452 is configured to reach other electronic devices via one or more (wired or wireless) network interfaces 420 , exemplary network interfaces 420 including Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB);

[0075] a presentation module 453 configured to enable display of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0076] The input processing module 454 is configured to detect user input or interaction from one or more input devices 432 and to translate the detected input or interaction.

[0077] In some embodiments, the apparatus provided in the embodiments of the present application can be implemented using software. FIG2 shows a video generation device 455 stored in a memory 450 . This device 455 can be software in the form of a program or plug-in, and includes the following software modules: a first display module 4551 , a second display module 4552 , a third display module 4553 , and a fourth display module 4554 . These modules are logical and can be arbitrarily combined or further analyzed based on the functions they implement. The functions of each module will be described below.

[0078] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the video generation device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video generation method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0079] In some embodiments, the terminal or server can implement the video generation method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (Application, APP), that is, a local client, that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP, a web browser APP; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of client, module or plug-in.

[0080] Based on the above description of the video generation system and electronic device provided in the embodiment of the present application, the video generation method provided in the embodiment of the present application is described below. In actual implementation, the video generation method provided in the embodiment of the present application can be implemented independently by a terminal or a server, or by a terminal and a server in collaboration. The video generation method provided in the embodiment of the present application is described by taking the terminal 400 in Figure 1 as an example of executing the video generation method provided in the embodiment of the present application independently. Referring to Figure 3, Figure 3 is a schematic flow chart of the video generation method provided in the embodiment of the present application. Next, the steps shown in Figure 3 will be described.

[0081] Step 101: The terminal displays at least one video generation control in a video generation interface, where the at least one video generation control is used to generate videos of at least two video types.

[0082] In actual implementation, the terminal is provided with a client that supports video generation. When the user opens the client on the terminal and the terminal runs the client, the terminal can display a video generation interface based on the client and display at least one video generation control in the video generation interface. The at least one video generation control is used to generate videos of at least two video types; wherein the video types include long videos and short videos.

[0083] In some embodiments, the number of video generation controls may be at least two, and different video generation controls correspond to different video types, so that the process of displaying the video generation interface may be, in response to an opening instruction for the video generation interface, displaying a video generation interface including at least two video generation controls; thereby, in response to the video generation control corresponding to the target video type in the video generation interface being in a selected state, determining that the target video type is in a selected state.

[0084] By applying the above embodiment, by selecting different video generation controls, different video types can be controlled to be in a selected state, thereby generating videos of the corresponding video types in the subsequent process; in this way, not only the video generation efficiency is improved, but also the human-computer interaction efficiency and the hardware resource utilization of the electronic device are improved.

[0085] In other embodiments, the number of video generation controls is one, so the process of displaying the video generation interface may be to display a video generation interface including a video generation control in response to an opening instruction for the video generation interface, and then to display at least two video type selection controls in response to a trigger operation for the video generation control, with different video type selection controls being used to select different video types, and to display a video generation interface of the target video type indicated by the target video type selection control in response to a selection operation for a target video type selection control among the at least two video type selection controls, and to determine that the target video type indicated by the target video type selection control is in a selected state.

[0086] It should be noted that when the video generation interface is displayed and there is one video generation control, after at least two video type selection controls are displayed in response to a trigger operation on the video generation control, the corresponding video type is determined to be in the selected state depending on which video type selection control the user triggers.

[0087] Step 102 : Based on at least one video generation control, in response to a target video type being selected from at least two video types, displaying a content input area corresponding to the target video type.

[0088] When a video generation interface is displayed and there are multiple video generation controls, one of the video generation controls may be automatically selected, or one of the video generation controls may be controlled to be in a selected state, that is, the target video type is in a selected state, based on a user's selection operation; for example, a blank video generation interface and at least two video generation controls in an unselected state are first displayed, and then, in response to a selection operation for a target video generation control among the at least two video generation controls, the target video generation control among the at least two video generation controls is controlled to be in a selected state, thereby displaying a content input area; or one of the at least two video generation controls is pre-set to be in a selected state, so that when the video generation interface is displayed, the video generation interface including the target video generation control in a selected state is first displayed, thereby displaying the content input area;

[0089] Based on this, in response to the video generation control corresponding to the target video type in the video generation interface being in a selected state, in the process of determining that the target video type is in a selected state, the target video generation control in the selected state can be determined based on the user's selection operation after displaying a blank video generation interface, or it can be determined in advance, that is, directly when the video generation interface is displayed, or it can be determined based on the user's selection operation on another video generation control that is not in a selected state after displaying the video generation control that is pre-set to be in a selected state. The embodiments of the present application do not limit this.

[0090] For example, referring to Figure 4, Figure 4 is a schematic diagram of a content input area displayed according to an embodiment of the present application. Based on a in Figure 4, when the target video generation control is a short video generation control, in response to the short video generation control indicated by 401 being in a selected state, a content input area as indicated by 402 is displayed; or, based on b in Figure 4, when the target video generation control is a long video generation control, in response to the long video generation control indicated by 403 being in a selected state, a content input area as indicated by 404 is displayed.

[0091] In some embodiments, based on at least one video generation control, in response to a target video type being in a selected state among at least two video types, while displaying the content input area corresponding to the target video type, guidance information is also displayed in the upper layer of the content input area in the form of a floating layer, and the guidance information is used to guide input operations based on the content input area. For example, referring to Figure 5, Figure 5 is a schematic diagram of guidance information provided in an embodiment of the present application. Based on Figure 5, when the target video generation control is a short video generation control, in response to the short video generation control indicated by 501 being in a selected state, while displaying the content input area indicated by 502, guidance information as indicated by 503 is displayed in the form of a floating layer.

[0092] In some embodiments, the user can directly click on the function control for displaying the video generation interface when running the client to display the video generation interface, that is, the video generation interface is displayed in response to the trigger operation of the function control for displaying the video generation interface; in other embodiments, the video generation interface can also be displayed based on the link shared by other users. Before displaying the video generation interface, it is also possible to respond to the display instruction for other videos triggered based on the shared link. If the current account has the playback permission for other videos, the playback interface of other videos is displayed, and the detailed information of other videos and the fourth new video generation control are displayed on the playback interface; thus, the process of displaying the video generation interface can be to display the video generation interface in response to the trigger operation for the fourth new video generation control.

[0093] It should be noted that the fourth new video generation control is used to jump from the playback interface to the video generation interface including a blank content input area, and generate a new video based on the video generation interface, that is, in response to the trigger operation for the fourth new video generation control, after the video generation interface including a blank content input area is displayed, the target content is entered in the blank content input area, and a video of the target video type is generated based on the input target content; the detailed information includes at least one of the duration, size, content description, prompt information on whether audio is included, resolution, and key image frames of other videos, wherein the key image frames are used to indicate the cover of other videos, or the pictures entered when generating other videos, etc.; the content description is used to indicate the text entered when generating other videos, or to indicate the text content obtained by analyzing the video content of other videos after generating other videos; and whether audio is included can indicate whether audio was entered when generating other videos, or can indicate whether audio was generated when generating other videos. This is not limited in the embodiments of the present application.

[0094] For example, refer to Figure 6, which is a schematic diagram of the playback interface of other videos provided in an embodiment of the present application. When the other video is a long video, the displayed details interface is as shown in Figure 6a, wherein the details information of the other video is shown in the dotted box 601, and in response to the triggering operation of the fourth new video generation control indicated by 602, the video generation interface as shown in Figure 4b is displayed; when the other video is a short video, the displayed playback interface is as shown in Figure 6b, wherein the details information of the other video is shown in the dotted box 603, and in response to the triggering operation of the fourth new video generation control indicated by 604, the video generation interface as shown in Figure 4a is displayed.

[0095] In actual implementation, when the current account does not have the permission to play other videos, permission prompt information is displayed; wherein, the permission prompt information is used to prompt the target object that the target object does not have the permission to play other videos.

[0096] It should be noted that whether the current account has permission may be set by the account that generates other videos, or may be determined by information such as the level of the current account. This embodiment of the present application does not limit this.

[0097] For example, referring to FIG7 , FIG7 is a schematic diagram of the permission prompt information provided by the present embodiment. Based on FIG7 , when the current account does not have the permission to play other videos, the permission prompt information shown in FIG7 is displayed.

[0098] In some embodiments, in a video generation interface, while displaying at least one video generation control, at least one video template can also be displayed; in response to a selection operation on a target video template in at least one video template, a template content description of the target video template and a corresponding first editing control are displayed; wherein the first editing control is used to generate a first video of the target video type based on the template content description; in response to a trigger operation on the first editing control, the trigger operation is determined as an input operation on the content input area; thus, in response to a content input operation in the content input area, the process of displaying the input target content may be, in response to a trigger operation on the first editing control, displaying the template content description of the target video template in the content input area, and determining the displayed template content description as the input target content.

[0099] It should be noted that the video template is a pre-generated video, and the first editing control is used to generate the first video of the target video type based on the template content description, which means that the template content description of the target video template is backfilled into the content input area, and based on the template content description of the target video template, the first video associated with the target video template is generated; for at least one video template, it can be uploaded by other users or uploaded by the current user; the selection operation of the target video template in at least one video template is used to indicate that the target video template is in a selected state, rather than clicking on the target video template. For example, the user can move the cursor controlled by the mouse to the top of the target video template, that is, when the cursor controlled by the user moves to the top of the target video template, it can be considered that the selection operation for the target video template is triggered.

[0100] Then, in response to the selection operation of the target video template in at least one video template, the template content description of the target video template and the corresponding first editing control are displayed on the target video template; wherein, when the click operation on the first editing control is performed, the template content description of the target video template is backfilled into the content input area, that is, in response to the triggering operation on the first editing control, a video generation interface including a content input area carrying the template content description of the target video template is displayed.

[0101] For example, referring to FIG8 , FIG8 is a schematic diagram of at least one video template provided in an embodiment of the present application. Based on FIG8 a , in response to a selection operation on a target video template indicated by 801, a template content description of the target video template as indicated by a dotted box 802 and a corresponding first editing control as indicated by 803 are displayed; then, in response to a click operation on the first editing control, when the target video template is a short video, an interface as shown in FIG8 b is displayed, wherein 804 in FIG8 b indicates a content input area carrying the template content description of the target video template;

[0102] When the target video template is a long video, the last interface when generating the long video is displayed, for example, the last interface is the video generation interface as shown in FIG8 c, wherein 805 in FIG8 c indicates a content input area carrying the template content description of the target video template; or the last interface is the storyboard disassembly interface mentioned later (such as when the target video, i.e., the target video type, is generated directly based on the storyboard disassembly interface, the last interface is the storyboard disassembly interface), or the last interface is the storyboard generation interface mentioned later (such as when the target video is generated directly based on the storyboard generation interface, or after the storyboard generation interface is displayed based on the storyboard disassembly interface, when the target video is generated based on the storyboard generation interface, the last interface is the storyboard generation interface). This embodiment of the application does not limit this.

[0103] By applying the above embodiment, based on the first editing control displayed for the selection operation of the target video template in at least one video template, a video associated with the target video template can be quickly generated. In this way, if the user is satisfied with the target video template, a video associated with the target video template can be directly generated. This not only improves the efficiency of human-computer interaction, but also improves the efficiency of video generation.

[0104] In other embodiments, in a video generation interface, while displaying at least one video generation control, at least one video template may also be displayed; in response to a trigger operation on a target video template in at least one video template, a detail interface including detailed information of the target video template is displayed; wherein the detail interface further includes at least one of a second editing control and a first new video generation control, the second editing control being used to regenerate a first video of the target video type based on the template content description, and the first new video generation control being used to re-enter the target content in the content input area and generate the first video of the target video type based on the re-entered target content; in response to a trigger operation on the second editing control, the user jumps from the detail interface to a video generation interface including a content input area carrying a template content description of the target video template; in response to a trigger operation on the first new video generation control, the user jumps from the detail interface to a video generation interface including a blank content input area;

[0105] Thus, in the subsequent process, in response to the content input operation performed in the content input area, the input target content is displayed. When a trigger operation for the second editing control is received, the trigger operation is determined as a content input operation, and the template content description displayed in the content input area is determined as the input target content; when a trigger operation for the first new video generation control is received, the input target content is displayed in response to the content input operation performed in the blank content input area.

[0106] It should be noted that the function of the second editing control is similar to that of the above-mentioned first editing control, and the first new video generation control is used to re-input the target content in the content input area, and generate the first video of the target video type based on the re-input target content. It means that when a trigger operation for the first new video generation control is received, it jumps from the details interface to the video generation interface including a blank content input area, so that the user can input the target content in the blank content input area by himself, and generate the first video of the target video type based on the input target content; at the same time, the trigger operation for the target video template in at least one video template can be a click operation on the target video template; the details information includes the target video At least one of the template's duration, size, content description, prompt information indicating whether audio is included, key image frames, and resolution; wherein the key image frames are used to indicate the cover of the target video template, or are pictures input when generating the target video template, etc.; the content description is used to indicate the text input when generating the target video template (that is, the content input in the content input area when generating the target video template), or is used to indicate the text content obtained by analyzing the video content of the target video template after generating the target video template; and whether audio is included can indicate whether audio was input when generating the target video template, or can indicate whether audio was generated when generating the target video template. This is not limited in the embodiments of the present application.

[0107] Exemplarily, continue to refer to Figure 8 and Figure 9, which is a schematic diagram of a details interface including detailed information of a target video template provided in an embodiment of the present application. Based on Figure 8a, when the target video template is a short video, in response to a trigger operation on the target video template indicated by 801, a details interface as shown in Figure 9a is displayed, wherein the dotted box 901 indicates the detailed information of the target video template, 902 indicates the second editing control, and 903 indicates the first new video generation control, thereby responding to a trigger operation on the second editing control, an interface as shown in Figure 8b is displayed, or, in response to a trigger operation on the first new video control, the video generation interface shown in Figure 4a is displayed;

[0108] Based on a of FIG8 , when the target video template is a long video, in response to the trigger operation for the target video template indicated by 801, the detail interface shown in b of FIG9 is displayed, wherein the dotted box 904 indicates the detail information of the target video template, 905 indicates the second editing control, and 906 indicates the first new video generation control, thereby responding to the trigger operation for the first new video control, the video generation interface shown in b of FIG4 is displayed; or, in response to the trigger operation for the second editing control, the last interface when generating a long video is displayed, for example, the last interface is shown in c of FIG8 The video generation interface, wherein 805 in FIG8 c indicates a content input area carrying a template content description of the target video template; or the last interface is the storyboard disassembly interface mentioned later (e.g., when the target video, i.e., the target video type, is generated directly based on the storyboard disassembly interface, the last interface is the storyboard disassembly interface), or the last interface is the storyboard generation interface mentioned later (e.g., when the target video is generated directly based on the storyboard generation interface, or after the storyboard generation interface is displayed based on the storyboard disassembly interface, when the target video is generated based on the storyboard generation interface, the last interface is the storyboard generation interface). This embodiment of the application does not limit this.

[0109] By applying the above embodiment, if the user is satisfied with the detail information corresponding to the target video template, a video associated with the target video template can be quickly generated based on the second editing control in the detail interface. In this way, not only the human-computer interaction efficiency is improved, but also the video generation efficiency is improved; at the same time, if the user is not satisfied with the detail information corresponding to the target video template, a new video can be quickly generated based on the first new video control in the detail interface. Therefore, not only the human-computer interaction efficiency is improved, but also the video generation efficiency is improved; in this way, based on the second editing control and the first new video control, not only the human-computer interaction efficiency and the video generation efficiency are improved, but also different choices are given to users, thereby improving the user experience.

[0110] Step 103 : In response to the content input operation performed in the content input area, the input target content is displayed.

[0111] It should be noted that the content input operation in the content input area can indicate the input of at least one of pictures, text, and audio, so that the target content can be at least one of the corresponding pictures, text, and audio; wherein, the associated area of ​​the content input area can also display at least one input content editing control, wherein at least one input content editing control is for the input pictures and audio, including a preview control, a re-upload control, and a delete control, after displaying the input picture or audio in response to the input operation on the content input area; in response to the trigger operation on the preview control, the input picture and audio are previewed, or in response to the re-upload operation triggered by the re-upload control, the re-input picture or audio is displayed, or in response to the trigger operation on the delete control, the input picture or audio is deleted. For example, referring to Figure 4 , based on a in Figure 4 , the dotted box 405 indicates three input content editing controls, wherein the first control is a preview control, the second control is a re-upload control, and the third control is a delete control. Correspondingly, based on b in Figure 4 , the dotted box 406 indicates three input content editing controls, wherein the first control is a preview control, the second control is a re-upload control, and the third control is a delete control.

[0112] In some embodiments, in response to a target video type being selected among at least two video types, a first determination control in an inactive state is displayed, and the first determination control is used to confirm the target content entered in the content input area; thereby, in response to a content input operation performed in the content input area, after the entered target content is displayed, the first determination control can be switched from an inactive state to an active state; and in response to a trigger operation on the activated first determination control, a video generation instruction is received.

[0113] It should be noted that the inactive state is used to indicate that the first determination control is in a gray state and cannot be triggered, and the active state is used to indicate that the first determination control is in a normal state and can be triggered. For example, referring to FIG10 , FIG10 is a schematic diagram of the first determination control provided in an embodiment of the present application. Based on FIG10 a , when the target video generation control is a short video generation control, in response to the short video type being in a selected state, the first determination control in an inactive state as indicated by 1001 in FIG10 a is displayed. After the input target content is displayed in response to the content input operation in the content input area, the first determination control is switched from an inactive state to an active state as shown in FIG10 b 1002 , thereby receiving a video generation instruction in response to the trigger operation on the first determination control in an activated state.

[0114] Accordingly, based on c in Figure 10, when the target video generation control is a long video generation control, in response to the long video type being in the selected state, the first determination control in the inactivated state as indicated by 1003 in c in Figure 10 is displayed, and in response to the content input operation in the content input area, after the input target content is displayed, as shown in 1004 in d in Figure 10, the first determination control is switched from the inactivated state to the activated state, thereby receiving the video generation instruction in response to the trigger operation for the first determination control in the activated state.

[0115] In some embodiments, while displaying the content input area, a parameter setting area is also displayed in the video generation interface, and the parameter setting area is used to set the parameters of the first video (the video to be generated); based on the parameter setting area, in response to the parameter setting operation, the video parameters set for the first video of the target video type are displayed; thereby, in response to the video generation instruction, based on the target content, the generated first video of the target video type with video parameters is displayed. For example, when the parameter setting of the first video is completed, in response to the trigger operation of the first determination control in the activated state, the video generation instruction is received; thereby, in the subsequent process, in response to the video generation instruction, the first video of the target video type with target parameters is generated; wherein, the parameters of the video include the duration, quality, ratio, resolution, etc. of the video.

[0116] For example, referring to Figure 11, Figure 11 is a schematic diagram of the display parameter setting area provided in an embodiment of the present application. Based on a in Figure 11, when the target video generation control is a short video generation control, the parameter setting area indicated by the dotted box 1101 is displayed; or, based on b in Figure 11, when the target video generation control is a long video generation control, the parameter setting area indicated by the dotted box 1102 is displayed; thereby, when the parameter setting of the target video is completed, in response to the trigger operation of the first determination control in the activated state, the video generation instruction is received, and the corresponding short video or long video is generated.

[0117] It should be noted that before setting the parameters, the default parameters will be displayed. For example, when the target video generation control is a short video generation control, the default parameters may be that the special effects are none by default, the video quality is 896*512 by default, the ratio is 7:4 by default, the video is 4s by default, and the audio is none by default. When the target video generation control is a long video generation control, the default parameters may be that the video length is 6s by default, the video quality is 896*512 by default, the ratio is 7:4 by default, and the audio is none. When no parameter setting is performed and the trigger operation for the first determination control in the activated state is directly executed, the parameters of the first video of the target video type generated in the subsequent process are the default video parameters. If parameter setting is performed, the parameters of the first video of the target video type generated in the subsequent process are the target parameters.

[0118] By applying the above embodiment, the video parameters of the generated first video can be set in the parameter setting area to generate a video that better meets the needs of the user, which not only enriches the diversity of the generated video, but also improves the effect of the generated video and the user experience.

[0119] In actual implementation, the video generation interface includes a parameter setting control, which is used to set the parameters of the generated first video, that is, the parameter setting area can be displayed based on the triggering of the parameter setting control, and in response to the triggering operation on the parameter setting control, the parameter setting area is displayed; based on the parameter setting area, in response to the parameter setting operation on the first video, the target parameters set for the first video are displayed; thus, in response to the video generation instruction, the process of displaying the first video generated based on the target content can be that when the setting of the parameters of the first video is completed, in response to the video generation instruction, the first video generated based on the target content is displayed; thus, in the subsequent process, in response to the video generation instruction, the first video with the target parameters is generated. For example, referring to Figure 4, based on a in Figure 4, the control indicated by 407 is the parameter setting control, and correspondingly, based on b in Figure 4, the control indicated by 408 is the parameter setting control.

[0120] It should be noted that, as mentioned above, video types include long videos and short videos, and there can be multiple video generation controls. The process of receiving a video generation instruction varies for different video types. Next, the process of receiving a video generation instruction for different video types is described.

[0121] This is for the case where the target video generation control includes a short video generation control, and the target video type includes a short video type.

[0122] In actual implementation, when the first video is a short video, the video generation instruction can be received directly based on the process described above, such as when the parameter setting of the first video is completed, in response to the trigger operation of the first determination control in the activated state, or directly in response to the trigger operation of the first determination control in the activated state.

[0123] This is for the case where the target video generation control includes a long video generation control and the target video type includes a long video type.

[0124] In some embodiments, when the first video is a long video, the long video includes multiple storyboards. After displaying the input target content in response to the content input operation performed in the content input area, a storyboard disassembly interface can also be displayed in response to a determination instruction for the target content, and at least two storyboard texts determined based on the target content can be displayed in the storyboard disassembly interface; wherein each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; thereby, in a subsequent process, in response to a video generation instruction triggered based on at least two storyboard texts, the generated first video of the long video type including the storyboards corresponding to each storyboard text is displayed.

[0125] It should be noted that a storyboard here is used to indicate a video clip, and multiple video clips are combined to form the final first video; after the input target content is displayed, the target content is parsed to obtain the text content corresponding to the target content, and then the text content is split to obtain multiple text paragraphs, thereby determining the multiple text paragraphs as multiple storyboard texts, and then displaying them in the storyboard disassembly interface.

[0126] At the same time, the confirmation instruction for at least two storyboard texts can be triggered by a second confirmation control. The storyboard disassembly interface also includes a second confirmation control, so that after the storyboard disassembly interface displays at least two storyboard texts determined based on the target content, it can also respond to the trigger operation of the second confirmation control and receive the confirmation instruction for at least two storyboard texts.

[0127] For example, referring to FIG12 , FIG12 is a schematic diagram of a storyboard disassembly interface provided in an embodiment of the present application. Based on FIG12 , the dotted box 1201 indicates four storyboard texts, each of which is used to describe the content of the corresponding storyboard. 1202 indicates a second determination control, and in response to a trigger operation on the second determination control, determination instructions for at least two storyboard texts are received.

[0128] By applying the above embodiment, the target content is decomposed into at least two storyboard texts, and then corresponding storyboards are generated based on the storyboard texts, thereby generating a first video of the long video type based on the at least two storyboards; in this way, the user can better understand the generated video, and it is also convenient for the user to adjust the generated video based on the storyboard in the subsequent process.

[0129] During actual implementation, it is also possible to display a storyboard text quantity adjustment control, which is used to adjust the quantity of storyboard texts, such as adding new storyboard texts or deleting existing storyboard texts; in response to a storyboard text quantity adjustment operation triggered based on the storyboard text quantity adjustment control, such as a storyboard text addition operation or a storyboard text deletion operation, in the storyboard disassembly interface, the storyboard text with the adjusted quantity, such as the added new storyboard text or the deleted storyboard text, is displayed; in response to a determination operation triggered based on at least two storyboard texts including the new storyboard text, a video generation instruction is received.

[0130] As an example, taking adding new text as an example, specifically, in response to the trigger operation of the storyboard text quantity adjustment control, a text input box is displayed in the associated area of ​​at least two storyboard texts; in response to the input operation on the text input box, the input text is displayed; in response to the confirmation instruction on the input text, the added new storyboard text is displayed, and the new storyboard text is associated with the input text.

[0131] It should be noted that the associated area of ​​at least two storyboard texts can be, for example, the area below at least two storyboard texts, and the confirmation instruction for the input text can be triggered by a confirmation control, which is not limited in this embodiment of the present application.

[0132] By applying the above embodiment, new storyboard texts are added or the original storyboard texts are deleted based on the storyboard text quantity adjustment control. In this way, the user can edit the generated video in more detail, generate a video that better meets the user's needs, and improve the effect of the generated video and user experience.

[0133] In actual implementation, the parameters of the storyboard of the corresponding storyboard text can also be set. Each storyboard text is associated with a parameter setting control, and the parameter setting control is used to set the storyboard parameters of the storyboard corresponding to the corresponding storyboard text; thus, after displaying the storyboard disassembly interface, it is also possible to display the target storyboard parameters set for the target storyboard text in response to the parameter setting operation triggered by the parameter setting control associated with the target storyboard text; in response to the determination instruction for the target storyboard parameters, the storyboard parameters of the storyboard corresponding to the target storyboard text are updated to the target storyboard parameters; thus, in the subsequent process, the generated first video, that is, the first video of the long video type, includes the target storyboard parameters as the target storyboard parameters.

[0134] It should be noted that the target storyboard text here may be part of the storyboard text in at least two storyboard texts, or may refer to all the storyboard texts, and this embodiment of the present application does not limit this; at the same time, when the trigger operation for the parameter setting control is executed, the parameter setting area is displayed, thereby implementing the parameter setting operation based on the parameter setting area. The parameter setting area here is as described above, and this embodiment of the present application does not elaborate on this.

[0135] For example, referring to FIG. 12 , the control indicated by the dashed box 1203 is a parameter setting control. In response to a triggering operation on the parameter setting control corresponding to a storyboard text, a parameter setting area is displayed. Then, in response to a parameter setting operation triggered by the parameter setting area, the target parameters set for the target storyboard text are displayed. In this way, the parameters of the storyboards corresponding to the storyboard texts are sequentially set, so that the ultimately generated first video conforms to the set parameters.

[0136] In actual implementation, the arrangement order of the storyboard texts can also be changed. At least two storyboard texts have an arrangement order. The content of the video composed of the storyboards corresponding to at least two storyboard texts with different arrangement orders is different. Each storyboard text is associated with a sorting adjustment control; thus, after the storyboard disassembly interface is displayed, the at least two storyboard texts can also be changed from the current arrangement order to the target arrangement order in response to the sorting adjustment operation for at least two storyboard texts triggered based on the sorting adjustment control; thus, in the subsequent process, based on the at least two storyboard texts, in response to the video generation instruction, a long video type video including at least two storyboards generated based on the target arrangement order is displayed.

[0137] It should be noted that, as mentioned above, the first video is composed of at least two storyboards, and there is a corresponding relationship between the storyboards and the storyboard texts. Therefore, the arrangement order of the at least two storyboard texts also indicates the combination order of the at least two storyboards. Therefore, the arrangement order of the at least two storyboards is different, which makes the combination order of the at least two storyboards different, resulting in different content of the first video.

[0138] It should be noted that the sorting of at least two storyboard texts can be achieved by dragging the corresponding storyboard text by dragging the sort adjustment control, or by executing a trigger operation on the sort adjustment control, an input box is displayed in the associated area of ​​each storyboard text, and then in response to the input operation on the input box, the input serial number for each storyboard text is displayed, thereby changing the at least two storyboard texts from the current sorting order to the target sorting order.

[0139] Exemplarily, continuing to refer to Figure 12, the control indicated in the dotted box 1204 is a sorting adjustment control for dragging the corresponding storyboard text to achieve the sorting of at least two storyboard texts. The storyboards corresponding to the at least two storyboard texts are combined in order from top to bottom to form a target video, thereby responding to the sorting adjustment operation on the at least two storyboard texts triggered by the sorting adjustment control, changing the arrangement order of the at least two storyboard texts, and then receiving a video generation instruction based on the changed order.

[0140] In actual implementation, part of the storyboard text can also be selected to generate a first video. In response to the selection operation of one or at least two target storyboard texts among at least two storyboard texts, the target storyboard text is controlled to be in a selected state; thereby, in the subsequent process, based on the target storyboard text in the selected state, in response to the video generation instruction, a first video of the long video type including the target storyboard text is displayed.

[0141] In actual implementation, the storyboard text can also be modified. In response to the text modification operation on the target storyboard text in at least two storyboard texts, the modified target storyboard text is displayed, so that in the subsequent process, based on the modified target storyboard text, in response to the video generation instruction, a first video of the long video type including the modified target storyboard text is displayed.

[0142] In actual implementation, as described above, a second confirmation control is also displayed in the storyboard disassembly interface. When the second confirmation control is triggered, a video generation instruction can be received, that is, the first video is directly generated; or a storyboard generation interface can be displayed, and the process of receiving the video generation instruction in response to the confirmation instruction for at least two storyboard texts can be, in response to the triggering operation for the second confirmation control, displaying the storyboard generation interface; in the storyboard generation interface, displaying the generated storyboard corresponding to each storyboard text, and displaying the third confirmation control; in response to the triggering operation for the third confirmation control, receiving the video generation instruction.

[0143] For example, continue to refer to Figure 12 and Figure 13, Figure 13 is a schematic diagram of the storyboard generation interface provided by an embodiment of the present application. As shown in Figure 12, 1202 indicates a second determination control, so that in response to the determination instruction for at least two storyboard texts triggered based on the second determination control, the storyboard generation interface as indicated in Figure 13 is displayed. In the storyboard generation interface, the storyboard corresponding to each storyboard text generated as indicated by the dotted box 1301 is displayed, and the third determination control indicated by 1302 is displayed, so that in response to the triggering operation of the third determination control, the video generation instruction is received.

[0144] It should be noted that the storyboard generation interface includes at least two storyboard texts and at least two storyboards, each storyboard text corresponds to a storyboard, and the storyboard is generated based on the corresponding storyboard text; thus, based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, a video of the target video type including at least two storyboards is displayed.

[0145] In actual implementation, each storyboard is associated with a parameter setting control, which is used to set the storyboard parameters of the corresponding storyboard; after displaying at least two storyboard texts and at least two storyboards on the storyboard generation interface, it is also possible to display the target storyboard parameters set for the target storyboard in response to a parameter setting operation triggered by a parameter setting control associated with the target storyboard; in response to a determination instruction for the target storyboard parameters, the storyboard parameters of the target storyboard are updated to the target storyboard parameters; thereby, in the subsequent process, the storyboard parameters of the target storyboard included in the first video of the generated long video type are the target storyboard parameters.

[0146] It should be noted that the target mirror here can be part of the mirrors among at least two mirrors, or it can refer to all the mirrors. This embodiment of the present application does not limit this. At the same time, when the trigger operation for the parameter setting control is executed, the parameter setting area is displayed, so that the parameter setting operation is implemented based on the parameter setting area. The parameter setting area here is as described above, and this embodiment of the present application does not elaborate on this.

[0147] For example, referring to FIG. 13 , 1303 indicates a parameter setting control for one of the storyboards, and the parameter setting controls for the other storyboards are similar. Thus, in response to a triggering operation on the parameter setting control corresponding to one of the storyboards, a parameter setting area is displayed, and then, in response to a parameter setting operation triggered by the parameter setting area, target storyboard parameters set for the target storyboard are displayed. In this way, the storyboard parameters for each storyboard are set sequentially, so that the ultimately generated first video, i.e., the first video of the long video type, conforms to the set storyboard parameters.

[0148] During actual implementation, it is also possible to select part of the storyboards to generate the first video. After displaying at least two storyboard texts and at least two storyboards on the storyboard generation interface, it is also possible to control the first target storyboard to be in a selected state in response to a selection operation for the first target storyboard among the at least two storyboards; thereby, in the subsequent process, based on the first target storyboard in the selected state and the storyboard text of the first target storyboard, in response to the video generation instruction, the generated first video of the long video type including the first target storyboard is displayed.

[0149] It should be noted that the number of the first target mirrors here can be one, or can be at least two, such as two or three, and this embodiment of the present application does not limit this.

[0150] For example, continuing to refer to Figure 13, based on the control indicated by the dotted box 1304, in response to the selection operation for the 1st, 2nd, and 4th frames, the 1st, 2nd, and 4th frames are controlled to be in a selected state, wherein "√" indicates that the corresponding frame is in a selected state, thereby generating a first video including the 1st, 2nd, and 4th frames.

[0151] By applying the above embodiment, by selecting a first target storyboard from at least two storyboards, a first video of a long video type including the first target storyboard is generated. In this way, the user can select the storyboards in the first video by himself, thereby generating a video that better meets the user's needs, which not only improves the effect of the generated video and the user experience, but also improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0152] In actual implementation, the arrangement order of the storyboards can also be changed. At least two storyboards have an arrangement order. The first videos corresponding to the storyboards of at least two storyboards with different arrangement orders are different. Each storyboard is associated with a sorting adjustment control, and the sorting adjustment control is used to drag the corresponding storyboard to achieve the sorting adjustment of the at least two storyboards; after at least two storyboard texts and at least two storyboards are displayed on the storyboard generation interface, the sorting of the at least two storyboards can be changed from the current arrangement order to the target arrangement order in response to the sorting adjustment operation for the at least two storyboards triggered based on the sorting adjustment control; thereby, based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, a first video of the long video type including at least two storyboards generated based on the target arrangement order is displayed.

[0153] It should be noted that, as mentioned above, the first video is composed of at least two storyboards. Therefore, the arrangement order of the at least two storyboards also indicates the combination order of the at least two storyboards. Therefore, the arrangement order of the at least two storyboards is different, which makes the combination order of the at least two storyboards also different, resulting in different content of the first video.

[0154] It should be noted that the sorting of at least two storyboards can be achieved by dragging the corresponding storyboard text by dragging the sort adjustment control, or by performing a trigger operation on the sort adjustment control to display an input box in the associated area of ​​each storyboard, so that in response to the input operation on the input box, the serial number entered for each storyboard is displayed, thereby changing the at least two storyboards from the current arrangement order to the target arrangement order; or by performing a trigger operation on the sort adjustment control to display at least one sorting element, such as when the storyboard includes a virtual object, the duration of the virtual object in each storyboard, or the number of words spoken, etc., so that in response to the selection operation of the target sorting element in at least one sorting element, the at least two storyboards are changed from the current arrangement order to the target arrangement order determined based on the target sorting element, for example, based on the duration of the virtual object in each storyboard, multiple storyboards are sorted from long to short.

[0155] Exemplarily, continuing to refer to Figure 13, the control indicated in the dotted box 1305 is a sorting adjustment control for dragging the corresponding storyboard text by dragging the sorting adjustment control to achieve the sorting of at least two storyboards. The at least two storyboards are combined in order from top to bottom to form a first video, thereby changing the arrangement order of the at least two storyboards in response to the sorting adjustment operation for the at least two storyboards triggered based on the sorting adjustment control.

[0156] By applying the above embodiment, a first video generated based on a target arrangement order is generated by changing the arrangement order of at least two storyboards. In this way, the diversity of the generated videos is enriched, so that videos that better meet user needs can be generated, which not only improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device, but also improves the effect of the generated video.

[0157] In actual implementation, part of the at least two storyboards may be updated. The storyboard generation interface further displays a storyboard update control for updating the content of the storyboard. There may be one or at least two storyboard update controls.

[0158] When there are at least two storyboard update controls, the storyboard update controls correspond one-to-one to the storyboards. After the storyboard generation interface displays at least two storyboard texts and at least two storyboards, it is also possible, based on the storyboard update controls, to respond to a content update operation for a second target storyboard among the at least two storyboards, to update the second target storyboard to a first new storyboard, and the content of the first new storyboard is associated with the content of the second target storyboard.

[0159] It should be noted that the content of the first new storyboard is also generated based on the content of the storyboard text of the corresponding storyboard, and is therefore associated with the content of the second target storyboard.

[0160] For example, referring to FIG13 , 1306 indicates the storyboard update control of one of the storyboards, and the storyboard update controls of other storyboards are similar; thus, in response to the triggering operation of the storyboard update control corresponding to one of the storyboards, the corresponding storyboard is updated, and the storyboard update controls corresponding to each storyboard can also be triggered in turn to complete the update of all storyboards.

[0161] When there is one storyboard update control, the storyboard update control is used to update at least two storyboards with one click. After the storyboard generation interface displays at least two storyboard texts and at least two storyboards, it is also possible to update the at least two storyboards to at least two third new storyboards in response to a triggering operation on the storyboard update control; wherein, there is a corresponding relationship between the at least two storyboards and the at least two third new storyboards, and the content of the third new storyboards is associated with the content of the corresponding storyboards.

[0162] It should be noted that the content of the third new storyboard is also generated based on the content of the storyboard text of the corresponding storyboard, and is therefore associated with the content of the corresponding storyboard; at the same time, when there are at least two storyboard update controls, a storyboard update control for performing one-click update of at least two storyboards can still be displayed, and this embodiment of the present application does not limit this.

[0163] Exemplarily, referring to FIG. 13 , 1307 indicates a storyboard update control, so that in response to a triggering operation on the storyboard update control, all storyboards are updated.

[0164] In actual applications, the storyboard is updated to generate a first video including the updated storyboard. In this way, if the user is not satisfied with the current storyboard, the current storyboard can be updated to generate a storyboard that better meets the user's needs. In the subsequent process, not only the effect of the generated video is improved, but also the human-computer interaction efficiency and the hardware resource utilization of the electronic device are improved.

[0165] In actual implementation, at least two storyboards are arranged in an order, and a third target storyboard among the at least two storyboards has an upward association control, the upward association control being used to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard, and the third target storyboard is any storyboard among the at least two storyboards except the first storyboard; after the storyboard generation interface displays the at least two storyboard texts and at least two storyboards, the third target storyboard can also be updated to a second new storyboard in response to a triggering operation on the upward association control. If there is at least one third target storyboard, there is also at least one upward association control, then, in response to a triggering operation on a target upward association control among the at least one upward association control, the storyboard associated with the target upward association control is updated to a second new storyboard; wherein the second new storyboard is obtained by updating the content of the storyboard associated with the target upward association control based on the content of the previous storyboard associated with the target upward association control.

[0166] It should be noted that each storyboard is associated with an upward association control, but only the third target storyboard can be turned on, that is, the upward association control can be triggered; the upward association control is used to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard, that is, based on the content of the previous storyboard, the content of the current third target storyboard is derived. In this way, the correlation between the content and effects between the storyboards can be improved. For example, if there is a group of birds flying by in the previous storyboard, the content of the third target storyboard is updated based on the content of the previous storyboard, so that the third target storyboard also includes this group of birds.

[0167] At the same time, when the upward association control is triggered, the corresponding third target storyboard can only be upwardly associated by default, and cannot be downwardly associated, and the storyboard ranked first cannot be upwardly associated; in addition, for two adjacent storyboards, if the upward association control can be turned on, but not at the same time, it is necessary to wait until the storyboard ranked earlier in the two storyboards is associated, that is, after the content is updated, the storyboard ranked later can be associated; in this way, it is prevented that the storyboard ranked later is updated multiple times, wasting resources.

[0168] For example, continuing to refer to Figure 13, 1308 indicates the upward association control of one of the third target storyboards, and the upward association controls of other third target storyboards are similar; thus, in response to the triggering operation of the upward association control corresponding to the third target storyboard, based on the content of the previous storyboard, the third target storyboard is updated to the second new storyboard, and the upward association controls corresponding to each third target storyboard can also be triggered in turn to complete the association of all storyboards.

[0169] By applying the above embodiment, the content of the storyboard associated with the target upward association control can be updated based on the content of the previous storyboard associated with the target upward association control. This not only improves the correlation between the content and effects between the storyboards, but also improves the editing efficiency of the storyboards.

[0170] During actual implementation, the functions of the upward-linked control can also be introduced. In response to the trigger operation for the upward-linked control, before updating the third target storyboard to the second new storyboard, if the account that performs the trigger operation generates the first video of the long video type for the first time, function introduction information can be displayed; wherein, the function introduction information is used to introduce the functions of the upward-linked control.

[0171] It should be noted that the function introduction information can be displayed before the storyboard generation interface is displayed, or it can be displayed when the storyboard generation interface is displayed. This embodiment of the present application does not limit this.

[0172] For example, referring to Figure 14, Figure 14 is a schematic diagram of the function introduction information provided in an embodiment of the present application. Based on Figure 14, before displaying the storyboard generation interface, the function introduction information as shown in Figure 14 can also be displayed, thereby displaying the storyboard generation interface in response to the determination instruction for the function introduction control triggered by the control indicated by 1401.

[0173] It should be noted that when a long video is generated, the corresponding account is tested. If the test result indicates that the corresponding account does not carry the target identifier, it is determined that this is the first time the corresponding account has generated a long video of this type. After the long video is generated, the corresponding account is annotated with the target identifier. If the test result indicates that the corresponding account carries the target identifier, it is determined that this is not the first time the corresponding account has generated a long video of this type. The target identifier is used to indicate whether this is the first time the corresponding account has generated a long video of this type.

[0174] In actual applications, when a user generates the first video of a long video type for the first time, the function of the upward-associated control is introduced based on the function introduction information, which reduces the difficulty of using the upward-associated control. This not only improves the user experience, but also avoids the incorrect use of the upward-associated control, thereby improving the efficiency of human-computer interaction.

[0175] In actual implementation, the number of candidate storyboards corresponding to each storyboard text may be at least two, and the process of displaying at least two storyboard texts and at least two storyboards on the storyboard generation interface may be, that is, displaying at least two storyboard texts on the storyboard generation interface, and displaying at least two candidate storyboards corresponding to each storyboard text; for each storyboard text, in response to a selection operation on a target candidate storyboard from the at least two candidate storyboards, controlling the target candidate storyboard to be in a selected state, and determining the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

[0176] It should be noted that each candidate storyboard is also associated with a preview control, so that in response to a trigger operation on a target preview control among the at least two preview controls, the candidate storyboard corresponding to the target preview control is previewed, and then a selection operation for the target candidate storyboard among the at least two candidate storyboards is received, which is triggered based on the preview result, so as to select a suitable target candidate storyboard from the at least two candidate storyboards.

[0177] For example, referring to FIG15 , FIG15 is a schematic diagram of at least two candidate storyboards provided in an embodiment of the present application. Based on FIG15 , the dotted box 1501 indicates three candidate storyboards corresponding to each storyboard text. Thus, for each storyboard text, a target candidate storyboard is selected from the three candidate storyboards, thereby determining at least two storyboards including the target candidate storyboard.

[0178] By applying the above embodiment, for each storyboard text, there are multiple candidate storyboards, so that the user can select the storyboard that ultimately corresponds to the storyboard text from the multiple candidate storyboards. In this way, the user is given more choices, which not only improves the user experience, but also makes the generated video more in line with the user's needs, that is, improves the effect of the generated video.

[0179] In actual implementation, the storyboard text can also be modified. In response to the text modification operation on the target storyboard text in at least two storyboard texts, the modified target storyboard text is displayed, so that in the subsequent process, based on the modified target storyboard text, in response to the video generation instruction, a first video of the long video type including the modified target storyboard text is displayed.

[0180] It should be noted that, after the storyboard text is modified, the storyboard update control corresponding to the corresponding storyboard text needs to be triggered in order to display the storyboard generated based on the modified storyboard text, or the storyboard generated based on the modified storyboard text may not be displayed, but the first video of the long video type including the modified target storyboard text may be displayed directly in response to the video generation instruction. As shown in FIG13 , after the storyboard text is modified, the trigger operation for the third determination control as described above is directly executed, thereby receiving the video generation instruction.

[0181] In actual implementation, the storyboard generation interface includes at least one generation prompt information, and there is a corresponding relationship between the generation prompt information and the storyboard text, which is used to indicate whether the storyboard corresponding to the corresponding storyboard text is successfully generated; thus, based on the fourth target storyboard, in response to the video generation instruction, the first video of the long video type including the fourth target storyboard is displayed; wherein the fourth target storyboard is the storyboard that is successfully generated as indicated by the generation prompt information.

[0182] It should be noted that when at least two candidate storyboards are generated based on the storyboard text, as long as there is a successfully generated candidate storyboard, the corresponding generation prompt information will indicate that the storyboard corresponding to the corresponding storyboard text has been successfully generated. Only when all candidate storyboards fail to be generated, the corresponding generation prompt information will indicate that the storyboard corresponding to the corresponding storyboard text has not been successfully generated.

[0183] In some embodiments, the target video type is a long video type, and the long video type video includes multiple storyboards. After the input target content is displayed in response to the input operation on the content input area, the storyboard generation interface can also be directly displayed in response to the confirmation instruction for the input target content; wherein the storyboard generation interface includes at least two storyboard texts and at least two storyboards, each storyboard text corresponds to a storyboard, and the storyboard is generated based on the corresponding storyboard text; thereby, based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, a first video of the long video type including at least two storyboards is displayed.

[0184] It should be noted that the relevant content in the storyboard generation interface and the relevant operations performed based on the storyboard generation interface are similar to the relevant content of the storyboard generation interface displayed after the storyboard disassembly interface as described above, and the process of performing relevant operations based on the displayed storyboard generation interface is similar. This embodiment of the present application will not be elaborated on this.

[0185] In actual applications, after triggering a determination instruction for the target content, the storyboard generation interface can be directly displayed, so that based on the storyboard generation interface, a first video of a long video type including at least two storyboards is generated; compared with first displaying the storyboard disassembly interface and then displaying the storyboard generation interface, so as to generate the first video based on the storyboard generation interface, the video generation efficiency is improved.

[0186] Step 104 : In response to the video generation instruction, display a first video of a target video type generated based on the target content.

[0187] It should be noted that, as mentioned above, video types include long videos and short videos, so the first video of the generated target video type can also be a long video or a short video.

[0188] In some embodiments, the process of displaying a first video of a target video type generated based on target content in response to a video generation instruction may be, in response to the video generation instruction, displaying a first video of a target video type generated based on target content in a video display interface; wherein, the video display interface further includes at least one of the following: detailed information of the generated first video, a third editing control, and a second new video generation control; wherein, the third editing control is used to regenerate a second video of the target video type based on the target content, and the second new video generation control is used to re-enter the target content in the content input area, and generate a third video of the target video type based on the re-entered target content.

[0189] It should be noted that the function of the third editing control is similar to that of the first editing control described above, and the function of the second new video generation control is similar to that of the first new video generation control described above; the detailed information includes at least one of the duration, size, content description, prompt information indicating whether audio is included, key image frames, and resolution of the first video; wherein, the key image frame is used to indicate the cover of the first video, or is a picture input when the first video is generated, etc.; the content description is used to indicate the text input when the first video is generated, or is used to indicate the text content obtained by analyzing the video content of the first video after the first video is generated; and whether audio is included can indicate whether audio was input when the first video was generated, or can indicate whether audio was generated when the first video was generated. This is not limited in the embodiments of the present application.

[0190] By applying the above embodiment, if the user is satisfied with the generated first video based on the detailed information of the first video, the second video associated with the first video can be quickly generated based on the third editing control in the video display interface. In this way, not only the human-computer interaction efficiency is improved, but also the video generation efficiency is improved; at the same time, if the user is not satisfied with the first video, the new video can be quickly generated based on the second new video control in the video display interface. Therefore, not only the human-computer interaction efficiency is improved, but also the video generation efficiency is improved; in this way, based on the third editing control and the second new video control, not only the human-computer interaction efficiency and the video generation efficiency are improved, but also different choices are given to users, thereby improving user experience.

[0191] In some embodiments, the video display interface includes a third editing control, the target video type includes a long video, and the long video type video includes at least two storyboards. The detailed information in the details interface also includes the duration of each storyboard and the storyboard text corresponding to the content of each storyboard. For example, referring to FIG16, FIG16 is a schematic diagram of the video display interface of a long video provided in an embodiment of the present application. Based on FIG16, the dotted box 1601 indicates detailed information, the third editing control indicated by 1602, and the second new video generation control indicated by 1603;

[0192] At the same time, when the target video is a long video, in response to the trigger operation for the third editing control, the last interface when generating the long video is displayed, which can be a video generation interface, or a storyboard disassembly interface (such as when the target video is generated directly based on the storyboard disassembly interface), or a storyboard generation interface (such as when the target video is generated directly based on the storyboard generation interface, or after the storyboard generation interface is displayed based on the storyboard disassembly interface, the target video is generated based on the storyboard generation interface).

[0193] Based on this, taking the last interface displayed when generating a long video as the storyboard generation interface as an example, as shown in Figure 13, the target content corresponding to the first video is at least two storyboard texts, and the storyboard generation interface also includes at least two storyboards, each storyboard text corresponds to a storyboard, and the storyboard is generated based on the corresponding storyboard text; thus, in response to the video generation instruction, after displaying the first video of the target video type generated based on the target content in the video display interface, it is also possible to, in response to the trigger operation of the third editing control, display the storyboard generation interface as shown in Figure 13; display at least two storyboard texts determined based on the target content on the storyboard generation interface, and display at least two storyboards determined based on the at least two storyboard texts; based on the storyboard generation interface, in response to the editing operation of the at least two storyboard texts and the at least two storyboards, display the edited at least two storyboard texts and the at least two storyboards; in response to the video generation instruction, display the second video of the long video type generated based on the edited at least two storyboard texts and the at least two storyboards.

[0194] Alternatively, taking the last interface displayed when generating a long video as the storyboard disassembly interface as an example, as shown in Figure 12, the target content corresponding to the first video is at least two storyboard texts, each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; thus, in response to the video generation instruction, after displaying the first video of the target video type generated based on the target content in the video display interface, it is also possible to, in response to the triggering operation of the third editing control, display the storyboard disassembly interface as shown in Figure 12, and display at least two storyboard texts determined based on the target content in the storyboard disassembly interface; based on the storyboard disassembly interface, in response to the editing operation of the at least two storyboard texts, display the edited at least two storyboard texts; in response to the video generation instruction, display the second video of the long video type generated based on the edited at least two storyboard texts.

[0195] It should be noted that the editing operations triggered by the storyboard generation interface can be as described above, operations triggered by various controls in the storyboard generation interface such as update controls, storyboard update controls, sorting adjustment controls, and parameter setting controls, or editing operations on storyboard texts as described above; similarly, the editing operations triggered by the storyboard disassembly interface can be as described above, operations triggered by various controls in the storyboard disassembly interface such as sorting adjustment controls and parameter setting controls, or editing operations on storyboard texts as described above, and the embodiments of the present application do not limit this.

[0196] Applying the above embodiment, when the first video is a long video, if the user is not satisfied with the generated first video, the storyboard generation interface corresponding to the first video can be displayed based on the third editing control, so that at least two storyboards and at least two storyboard texts included in the first video can be edited in the storyboard generation interface, and then a new first video can be generated based on the edited at least two storyboards and at least two storyboard texts; in this way, if the user is not satisfied with the generated first video, based on the third editing control, the user can modify the first video, so that a storyboard that better meets the user's needs can be generated, which not only improves the effect of the generated video, but also improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0197] Among them, when the video display interface includes a second new video generation control and the type of the first video includes a long video, when a trigger operation for the second new video generation control is received, the video generation interface shown in Figure 4b is displayed.

[0198] In other embodiments, the video display interface includes a third editing control, and the type of the first video includes a short video; see Figure 17, Figure 17 is a schematic diagram of the video display interface of the short video provided in an embodiment of the present application. Based on Figure 17, the dotted box 1701 indicates detailed information, 1702 indicates the third editing control, and 1703 indicates the second new video generation control; at the same time, as described above, when the first video is a short video, in response to the trigger operation on the third editing control, the video generation interface is displayed, and in response to the video generation instruction, after the first video of the target video type generated based on the target content is displayed in the video display interface, it can also be displayed in response to the trigger operation on the third editing control, with the target content displayed in the content input area; in response to the editing operation on the target content, the new target content obtained by editing the target content is displayed; in response to the video generation instruction, the second video of the short video type generated based on the new target content is displayed.

[0199] It should be noted that the editing operation on the target content in the content input area can be an editing operation on the input text, or it can be an editing operation on the input audio or picture, such as deleting the original audio and / or picture, or re-uploading new audio and / or picture, etc. The target content is different, and the object of the editing operation is different.

[0200] Among them, when the video display interface includes a second new video generation control and the type of the first video includes a short video, when a trigger operation for the second new video generation control is received, the video generation interface shown in a in Figure 4 is displayed.

[0201] In actual applications, when the first video is a short video, if the user is not satisfied with the generated first video, the user can display a video generation interface including a content input area carrying the target content corresponding to the first video based on the third editing control, so that the target content corresponding to the first video can be edited in the video generation interface, and then a new first video can be generated based on the edited target content; in this way, if the user is not satisfied with the generated first video, based on the third editing control, the user can modify the first video, so that a video that better meets the user's needs can be generated, which not only improves the effect of the generated video, but also improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0202] In some embodiments, in response to a video generation instruction, the process of displaying the first video of the target video type generated based on the target content may be, in response to the video generation instruction, displaying at least one of a first progress prompt information, a fourth editing control, a fourth new video generation control, and a generation cancel control; wherein the first progress prompt information is used to prompt the generation progress of the first video of the target video type, the fourth editing control is used to regenerate the first video of the target video type based on the target content, the fourth new video generation control is used to re-enter the target content in the content input area, and generate the first video of the target video type based on the re-entered target content, and the generation cancel control is used to terminate the generation process of the first video of the target video type; when the first progress prompt information indicates that the generation of the first video of the target video type is complete, the first progress prompt information is canceled, and the first video of the target video type generated based on the target content is displayed.

[0203] It should be noted that when the first progress prompt information is displayed, it indicates that the first video of the target video type is in the process of being generated. When the generation progress of the first video of the target video type indicated by the first progress prompt information is 100%, it indicates that the generation of the first video of the target video type is completed. In addition, the fourth editing control is similar to the third editing control described above, and the fourth new video generation control has a similar function to the second new video generation control described above. At the same time, when the first video of the target video type is a long video or a short video, the relevant content of triggering the fourth editing control is also similar to the relevant content of triggering the third editing control, and accordingly, the relevant content of triggering the fourth new video generation control is also similar to the relevant content of triggering the second new video generation control. This will not be elaborated in the embodiments of the present application.

[0204] For example, referring to Figure 18, Figure 18 is a schematic diagram of the first progress prompt information provided by an embodiment of the present application. Based on Figure 18, in response to the video generation instruction, the first progress prompt information indicated by the dotted box 1801 in Figure 18, the fourth editing control indicated by 1802, the fourth new video generation control indicated by 1803, and the generation cancel control indicated by 1804 are displayed.

[0205] In some embodiments, in response to a video generation instruction, the process of displaying a first video of a target video type generated based on target content may be, in response to the video generation instruction, displaying at least one of a second waiting prompt message, a sixth editing control, a fifth new video generation control, and a waiting cancellation control; wherein the second waiting prompt message is used to indicate the length of time required to wait before starting to generate the first video; the sixth editing control is used to regenerate the first video of the target video type based on the target content, the fifth new video generation control is used to re-enter the target content in the content input area, and generate the first video of the target video type based on the re-entered target content, and the waiting cancellation control is used to terminate the waiting process required to generate the first video of the target video type; if a second waiting prompt message is displayed, when the second waiting prompt message indicates that the length of time required to wait before starting to generate the first video is zero, the display of the second waiting prompt message is canceled; and the first video of the target video type generated based on the target content is displayed.

[0206] It should be noted that when the second waiting prompt information is displayed, it means that the first video of the target video type is waiting to be generated, that is, the first video of the target video type has not yet started to be generated and is in a queue state. When the second waiting prompt information indicates that the waiting time from the start of generating the first video of the target video type is zero, it means that the waiting of the first video of the target video type is over, that is, there is no video being generated ahead and there is no need to queue. In addition, the sixth editing control is similar to the third editing control described above, and the fifth new video generation control is also similar in function to the second new video generation control described above. At the same time, when the target video is a long video or a short video, the relevant content when triggering the sixth editing control is also similar to the relevant content when triggering the third editing control. Correspondingly, the relevant content when triggering the fifth new video generation control is also similar to the relevant content when triggering the second new video generation control. This will not be elaborated in the embodiments of the present application.

[0207] For example, referring to Figure 19, Figure 19 is a schematic diagram of the second waiting prompt information provided in an embodiment of the present application. Based on Figure 19, in response to the video generation instruction, the second waiting prompt information indicated by the dotted box 1901 in Figure 19, the sixth editing control indicated by 1902, the fifth new video generation control indicated by 1903, and the waiting cancellation control indicated by 1904 are displayed.

[0208] It should be noted that, in addition to directly displaying the first progress prompt information or the second waiting prompt information, at least one of the second waiting prompt information, the sixth editing control, and the fifth new video generation control may also be displayed first; if the second waiting prompt message is displayed, when the second waiting prompt information indicates that the waiting time from the start of generating the first video of the target video type is zero, the second waiting prompt information is canceled, and then at least one of the first progress prompt information, the fourth editing control, and the fourth new video generation control is displayed; if the first progress prompt information is displayed, when the first progress prompt information indicates that the generation of the first video is complete, the first progress prompt information is canceled, and the first video of the target video type generated based on the target content is displayed. This embodiment of the present application does not limit this.

[0209] By applying the above embodiment, the second waiting prompt information is displayed to prompt the user how long they need to wait before starting the first video, thereby improving the user experience.

[0210] In some embodiments, after displaying the video generation interface, video generation prompt information may also be displayed; wherein the video generation prompt information is used to indicate the number of first videos that can be generated within the target time; thus, the process of displaying the first video of the target video type generated based on the target content in response to the video generation instruction may be, in response to the video generation instruction, displaying the first video of the target video type generated based on the target content when the video generation prompt information indicates that the number of first videos that can be generated within the target time is not zero.

[0211] It should be noted that the target time here is used to indicate the target period and can be pre-set, such as one day, one week, etc.; the process of displaying the video generation prompt information can be directly displayed, or it can be displayed after the input operation is performed on the content input area and the input target content is displayed, or it can be displayed after receiving the trigger operation of the first determination control in the activated state as described above. This embodiment of the application does not limit this.

[0212] At the same time, the video prompt information can be displayed at an associated position of the content input area, such as one of the upper position, lower position, left position and right position of the content input area, or it can be displayed at an associated position of the first confirmation control, such as one of the upper position, lower position, left position and right position of the first confirmation control.

[0213] For example, a video generation prompt message such as "10 production tasks remaining today" is displayed, indicating that the number of first videos that can be generated within the target time, i.e., one day, is 10; or a message such as "Today's generation task quota has been used up, please try again tomorrow" is displayed, indicating that the number of first videos that can be generated within the target time, i.e., one day, is zero. In this way, when responding to a video generation instruction, when the video generation prompt message indicates that the number of first videos that can be generated within the target time is not zero, a first video of the target video type generated based on the target content is displayed.

[0214] In some embodiments, the video generation interface also includes a video display control, which is used to display at least one asset video, including at least one of a successful video of historical generation success, a failed video of historical generation failure, a generating video in the process of being generated, and a waiting video waiting to be generated; thus, after displaying the first video of the target video type generated based on the target content in response to the video generation instruction, it is also possible to display at least one asset video including the first video of the target video type in response to a trigger operation on the video display control.

[0215] It should be noted that the video display control is used to display the asset video display interface, which also includes at least two video controls. Different video controls are used to display different types of asset videos. In response to the target video control being selected among the at least two video controls, at least one asset video of the target type corresponding to the target video control is displayed.

[0216] It should be noted that the arrangement order of the successful videos with historical generation success, the failed videos with historical generation failure, the generating videos in the generating process, and the waiting videos waiting to be generated in at least one asset video can be pre-set, such as being displayed from top to bottom in the order of waiting videos, generating videos, failed videos, and successful videos.

[0217] For example, refer to Figure 20, which is a schematic diagram of the asset video display interface provided by an embodiment of the present application. Based on Figure 20, in response to the long video control being in the selected state, nine asset videos of the long video type are displayed, among which the first, second, and third asset videos are waiting videos, the fourth and fifth asset videos are synthesized videos, the sixth and seventh asset videos are failed videos, and the eighth and ninth asset videos are successful videos.

[0218] In actual applications, by displaying an asset video display interface including asset videos, the generated video can be edited after it is generated, which not only optimizes the effect of the generated video, but also improves the human-computer interaction efficiency and the hardware resource utilization of electronic equipment.

[0219] It should be noted that at the associated position of each asset video, at least one of the duration, generation time, content description, and key image frame of the corresponding asset video will also be displayed; at the same time, a refresh control will also be displayed in the asset video display interface, so that in response to the triggering operation of the refresh control, at least one asset video in the asset video display interface will be refreshed.

[0220] Next, different asset videos are explained separately.

[0221] In some embodiments, the asset video includes a successful video, thereby displaying the playback control corresponding to the successful video; in response to a trigger operation on the playback control, the successful video is played; when the successful video is finished playing, in response to a trigger operation on the successful video, a detail interface of the detailed information of the successful video is displayed; wherein, the detail interface also includes a fifth editing control and a sixth new video generation control, the fifth editing control is used to regenerate a video of the video type corresponding to the successful video based on the content of the successful video, and the sixth new video generation control is used to re-enter the target content in the content input area, and generate a video of the video type corresponding to the successful video based on the re-entered target content.

[0222] It should be noted that, when the asset video is a successful video, the key image frame is used to indicate the cover of the successful video (such as it can be pre-set), or the picture input when the successful video is generated; the detailed information includes the duration, size, content description, prompt information of whether audio is included, resolution, and at least one of the key image frames of the successful video. When the successful video is a long video including multiple storyboards, the detailed information may also include the duration of each storyboard and the storyboard text corresponding to the content of each storyboard; at the same time, the fifth editing control is similar to the function of the third editing control described above, and the sixth new video generation control is also similar to the function of the second new video generation control described above; at the same time, when the target video is a long video or a short video, the relevant content when the fifth editing control is triggered is also similar to the relevant content when the third editing control is triggered. Correspondingly, the relevant content when the sixth new video generation control is triggered is also similar to the relevant content when the second new video generation control is triggered. This will not be elaborated in the embodiments of the present application.

[0223] In some embodiments, the asset video includes a failed video, and the failed video is associated with a regeneration control, and the regeneration control is used to regenerate the failed video; thus, in response to a trigger operation on the video display control, after displaying at least one asset video including the first video, it is also possible to regenerate the target failed video in response to a trigger operation on the regeneration control associated with the target failed video in the at least one failed video; when the regeneration is successful, the failed video in the at least one asset video is switched to a regenerated fourth video; when the regeneration fails, a regeneration failure prompt message is displayed, and the regeneration failure prompt message is used to prompt that the regeneration of the failed video has failed.

[0224] Exemplarily, as shown in FIG20 , in response to the triggering operation of the regeneration control associated with the seventh failed video as indicated in 2001, the corresponding failed video is regenerated; when the regeneration is successful, the corresponding failed video is switched to the regenerated fourth video; when the regeneration fails, a regeneration failure prompt message is displayed.

[0225] By applying the above embodiment, when the asset video includes a failed video that fails to be generated, the failed video can also be regenerated based on the regeneration control, thereby improving the success rate of video generation, thereby improving the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0226] It should be noted that, for failed videos, when the failed video is a long video, the reason for the generation failure will also be displayed in the associated position, such as the failure of storyboard generation, or the failure of storyboard synthesis, etc.; at the same time, when the regeneration is successful, the failed video will be switched to the regenerated fourth video, and the playback controls will also be displayed so that the regenerated fourth video can be played. Moreover, after the failed video is switched to the regenerated fourth video, the switched fourth video can be regarded as a successful video, thereby responding to the trigger operation for the successful video and realizing the relevant process of the asset video being a successful video as described above. This is not elaborated in the embodiments of the present application.

[0227] In some embodiments, the asset video includes a video being generated, so that a second progress prompt information is displayed in an associated area of ​​the video being generated, and the second progress prompt information is used to prompt the generation progress of the video being generated, and the generation progress indicated by the progress prompt information increases with time; in response to a trigger operation for the video being generated, a generation cancel control is displayed, and the generation cancel control is used to terminate the generation process of the video being generated.

[0228] For example, as shown in FIG20 , in the associated areas of the fourth and fifth generating videos, second progress prompt information as indicated by 2002 is displayed, wherein the second progress prompt information is used to prompt the generating progress of each generating video.

[0229] It should be noted that in response to the trigger operation for the video being generated, an interface as shown in Figure 18 will also be displayed, where, as mentioned above, the interface includes the first progress prompt information indicated by the dotted box 1801, the fourth editing control indicated by 1802, the fourth new video generation control indicated by 1803, and the generation cancel control indicated by 1804.

[0230] In some embodiments, the asset video includes a waiting video, so that a first waiting prompt information is displayed in an associated area of ​​the waiting video, and the first waiting prompt information is used to indicate that the corresponding waiting video is waiting to be generated; in response to a trigger operation for the waiting video triggered based on the first waiting prompt information, a waiting cancel control is displayed, and the waiting cancel control is used to terminate the waiting process of the waiting video.

[0231] Exemplarily, as shown in FIG20 , in the associated areas of the first, second, and third waiting videos, the first waiting prompt information as indicated by 2003 is displayed, wherein the first waiting prompt information is used to indicate that the corresponding waiting video is waiting to be generated.

[0232] In actual applications, by displaying the first waiting prompt information, the user is prompted that the waiting video is waiting to be generated, thereby improving the user experience; at the same time, based on the waiting cancel control, the waiting process of waiting for the video can be terminated. This not only gives the user more choices in the video generation process, further improving the user experience, but also when the user does not want to wait for the video to be generated, the waiting process of the video can be terminated in time, avoiding waste of computer resources.

[0233] It should be noted that for waiting videos, when the waiting video is a long video, the waiting process information will also be displayed in the relevant position, such as storyboards to be generated, storyboards being generated, or waiting to be synthesized;

[0234] When the waiting process information indicates that the corresponding waiting video is in a state of waiting for storyboard generation, in response to the triggering operation on the waiting video, the storyboard disassembly interface or the video generation interface as described above is displayed; when the waiting process information indicates that the corresponding waiting video is in a state of storyboard generation, in response to the triggering operation on the waiting video, the storyboard generation interface as described above is displayed; when the waiting process information indicates that the corresponding waiting video is in a state of waiting for synthesis, in response to the triggering operation on the waiting video, the interface as shown in Figure 19 is displayed, wherein, as described above, the interface includes the second waiting prompt information indicated by the dotted box 1901, the sixth editing control indicated by 1902, the fifth new video generation control indicated by 1903, and the waiting cancellation control indicated by 1904.

[0235] In actual implementation, each asset video is also associated with a deletion control, so that in response to the triggering operation of the deletion control corresponding to the target asset video in at least one asset video, a deletion prompt information is displayed, and the deletion prompt information is used to determine whether to delete the target asset video; in response to the determination instruction triggered based on the deletion prompt information, the target asset video is deleted.

[0236] For example, continue to refer to Figure 20 and Figure 21, Figure 21 is a schematic diagram of the deletion prompt information provided by an embodiment of the present application. Based on Figure 20, in response to the triggering operation of the deletion control indicated by 2004 associated with the seventh failed video, the deletion prompt information as indicated in Figure 21 is displayed, thereby deleting the target asset video in response to the determination instruction triggered based on the deletion prompt information.

[0237] In some embodiments, in response to a content input operation performed in the content input area, after displaying the input target content, it is also possible, in response to a video generation instruction, to display at least one of a generation failure prompt message, a regeneration control, and a backtracking control when the generation of the first video of the target video type fails; wherein the generation failure prompt message is used to indicate that the generation of the first video of the target video type failed and the reasons for the generation failure, the regeneration control is used to regenerate the first video of the target video type, and the backtracking control is used to return to the previous step.

[0238] For example, refer to Figure 22, which is a schematic diagram of the generation failure prompt information, regeneration control and backtracking control provided in an embodiment of the present application. Based on Figure 22, the dotted box 2201 indicates the generation failure prompt information, 2202 indicates the regeneration control, and 2203 indicates the backtracking control.

[0239] It should be noted that the storyboard disassembly interface, the storyboard generation interface, and the interface displaying the second waiting prompt information and the interface displaying the first process prompt information can all include a backtracking control, thereby returning to the previous step based on the backtracking control; however, after triggering the backtracking control displayed on the interface displaying the generation failure prompt information, the content of the previous step returned can be modified. For example, when the long video fails to be generated, the storyboard disassembly interface or the storyboard generation interface is displayed in response to the triggering operation of the backtracking control. At this time, the storyboard text and / or storyboards can be modified based on the corresponding storyboard disassembly interface or the corresponding storyboard generation interface, so that the long video is generated again based on the modified storyboard text and / or storyboards;

[0240] For situations where generation does not fail, after triggering the backtrace control displayed on the corresponding interface, the content of the previous step returned cannot be modified. For example, on the interface displaying the second waiting prompt information, the backtrace control is displayed, and in response to the triggering operation of the backtrace control, the interface of the previous step is displayed, but the content of the interface cannot be modified.

[0241] In some embodiments, the video generation interface also includes a style setting control, which is used to set the video style of the generated first video; in the video generation interface, after displaying at least one video generation control, it can also display at least one video style in response to a trigger operation on the style setting control; in response to a selection operation on a target video style in at least one video style, the target video style is controlled to be in a selected state; thus, the process of displaying the first video of the target video type generated based on the target content in response to the video generation instruction may be, in response to the video generation instruction, displaying the first video of the target video type generated based on the target content with the target video style.

[0242] It should be noted that the video style here can be a cartoon style, science fiction style, realistic style, etc.; for example, refer to Figure 23, Figure 23 is a schematic diagram of at least one video style provided by an embodiment of the present application. Based on Figure 23, in response to the triggering operation of the style setting control indicated by 2301, at least one video style as indicated by the dotted box 2302 is displayed, and thus in response to the selection operation of the target video style in at least one video style, the target video style is controlled to be in a selected state, thereby generating a first video of the target video type with the target video style.

[0243] By applying the above embodiment, the video style of the generated first video can be set through the style setting control, which not only enriches the diversity of the first video, but also improves the video effect of the first video, and also improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0244] In actual implementation, in response to a selection operation for a target video style in at least one video style, after controlling the target video style to be in a selected state, it is also possible to, in response to a determination instruction for the target video style in the selected state, display an identifier of the target video style on the style setting control, the identifier being associated with a style cancellation control; wherein the identifier of the target video style is used to identify that the first video of the generated target video type has the target video style; in response to a triggering operation for the style cancellation control, canceling the display of the identifier of the target video style on the style setting control; thereby, the process of displaying the first video of the target video type with the target video style generated based on the target content in response to the video generation instruction may be, in response to the video generation instruction, displaying the first video of the target video type with a default style generated based on the target content.

[0245] It should be noted that the identifier of the target video style can refer to the name of the target video style. After the identifier of the target video style is canceled on the style setting control, the original video style control will be displayed. Therefore, if a video generation instruction is received at this time, a video of the target video type with a default style will be generated. For example, referring to Figure 24, Figure 24 is a schematic diagram of the identifier of the target video style provided by an embodiment of the present application. Based on Figure 24a, in response to a determination instruction for the target video style in a selected state, the identifier of the target video style, i.e., the cartoon style, is displayed on the style setting control as indicated by 2401, wherein the style cancellation control associated with the identifier is shown as 2402, thereby responding to a triggering operation for the style cancellation control, the identifier of the target video style is canceled on the style setting control, and the original video style control as shown in 2403 in Figure 24b is displayed; thereby, based on the original video style control, in response to the video generation instruction, a first video of the target video type with a default style generated based on the target content is displayed.

[0246] In actual applications, the style cancellation control can be used to cancel the video style selected by the user, avoiding the situation where the user mistakenly selects a video style, thereby improving the fault tolerance of the video generation process. This not only improves the user experience, but also improves the human-computer interaction efficiency and video generation effect.

[0247] In some embodiments, when multiple first videos of the target video type are generated, the process of displaying the first videos of the target video type generated based on the target content is to automatically play the first videos of each target video type in sequence until the first video of the last target video type is played.

[0248] In some embodiments, after displaying a first video of a target video type generated based on target content in response to a video generation instruction, operation controls for the target video may also be displayed, the operation controls including at least one of: a like control for liking the first video, a share control for sharing the first video, a download control for downloading the first video, a dislike control for disliking the first video, a play control for playing the first video, and a full-screen control for displaying the first video in full screen; in response to a trigger operation on the target operation control, the operation indicated by the target operation control is performed on the first video, the target operation control being one of the like control, the share control, the download control, the dislike control, the play control, and the full-screen control.

[0249] For example, refer to Figure 25, which is a schematic diagram of the operating controls provided in an embodiment of the present application. Based on Figure 25, the dotted boxes 2501 indicate the like control, the dislike control, the share control, and the download control, respectively; 2502 indicates the play control; and 2503 indicates the full-screen control.

[0250] In some embodiments, after displaying a first video of a target video type generated based on target content in response to a video generation instruction, at least one interactive control for the first video may also be displayed, the interactive control being one of the following controls: a local editing control for locally editing the first video, a canvas expansion control for editing the size of the first video, a motion brush control for changing a static object in the first video into a dynamic object, a camera parameter setting control for setting camera parameters for the first video, a background removal control for removing the background of the first video, an object erasing control for erasing a target object in the first video, a scene detection control for dividing the video into different video segments based on different scenes in the first video, a depth of field setting control for setting a depth of field effect for the first video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the first video, an action sequence generation control for setting an action sequence for an object in the first video, and a third new video generation control for generating a new video based on audio data in the first video; in response to a triggering operation on a target interactive control among at least one interactive control, the interactive operation indicated by the target interactive control is performed on the generated first video of the target video type.

[0251] In actual implementation, when the target interactive control is a local edit control, in response to a trigger operation on the local edit control, a local edit box and a text input box are displayed; in response to an input operation based on the text input box, the input text content is displayed, and the text content is used to describe the editing task; in response to a video editing instruction, a fifth video that has undergone local editing is displayed, wherein the fifth video is a video obtained by editing the area indicated by the local edit box in the first video based on the text content;

[0252] For example, referring to Figure 26, Figure 26 is a schematic diagram of interacting with the first video based on a local editing control provided in an embodiment of the present application. Based on Figure 26, when the target interactive control is a local editing control, in response to a trigger operation on the local editing control, a local editing box as indicated by 2601 in a of Figure 26 and a text input box as indicated by 2602 are displayed, and then in response to an input operation based on the text input box, the input text content as indicated by 2603 in b of Figure 26 is displayed, thereby displaying a fifth video that has been locally edited in response to the video editing instruction.

[0253] It should be noted that when the local editing box is displayed, the size of the local editing box can also be adjusted; at the same time, the editing effective period can also be set. After the local input box is displayed, the forward time input box and the backward time input box are displayed. When the user enters the target time period in the forward time input box, the video in the forward target time period will have an editing effect by default from the time corresponding to the image frame where the local editing box is located. When the user enters the target time period in the backward time input box, the video in the backward target time period will have an editing effect by default from the time corresponding to the image frame where the local editing box is located. When there is no content entered in the forward time input box and the backward time input box, the entire video will have an editing effect by default.

[0254] It should be noted that the input text content may be to add a target object in the local edit box, or to erase an object in the local input box.

[0255] In actual implementation, when the target interactive control is a canvas expansion control, a canvas of at least one size is displayed in response to a triggering operation on the canvas expansion control. In response to a selection operation on a canvas of the target size, the first video is displayed on the canvas of the target size. Furthermore, in response to a resizing operation on the first video, a sixth video is displayed on the canvas of the target size. The sixth video is the resized first video, and its size is no larger than the first video. This improves the convenience of resizing the first video based on the canvas expansion control.

[0256] For example, referring to FIG27 , FIG27 is a schematic diagram of interacting with the first video based on a canvas expansion control provided in an embodiment of the present application. Based on FIG27 , when the target interactive control is a canvas expansion control, in response to a selection operation on a canvas of a target size, a sixth video as indicated by 2702 is displayed on a canvas of a target size as indicated by 2701.

[0257] In actual implementation, when the target interactive control is an action sequence generation control, in response to a trigger operation on the action sequence generation control, at least one candidate action sequence is displayed, in response to a trigger operation on a target candidate action sequence in at least one candidate action sequence, the target candidate action sequence in a selected state is displayed, and then in response to a video editing instruction, a seventh video is displayed, which is a video obtained by setting the object in the first video to the target candidate action sequence.

[0258] It should be noted that the action sequence is used to indicate a continuous action video, such as a dance video, and setting the object in the first video to the target candidate action sequence refers to controlling the corresponding object to perform the action indicated by the target candidate action sequence, such as making the corresponding object dance with the dance movements indicated by the dance video.

[0259] It should be noted that when displaying at least one candidate action sequence, a text input box will also be displayed, so that in response to the input operation based on the text input box, the input text content is displayed, and the text content is used to describe the editing task, so that in response to the video editing instruction, the seventh video obtained by setting the object in the first video to the target candidate action sequence determined based on the text content is displayed.

[0260] In actual implementation, when the target interactive control is a motion brush control, in response to a trigger operation on the motion brush control, the brush control is displayed, and the brush control can be moved on the image frame of the first video to select a static target object in the corresponding image frame, thereby changing the state of the selected target object from static to dynamic; in response to a selection operation triggered by the brush control, the selected static target object is displayed on the corresponding image frame; in response to a determination instruction for the target object, the state of the selected target object is changed from static to dynamic; in response to a video editing instruction, an eighth video is displayed, and the eighth video is a video obtained by changing the state of the target object in the first video from static to dynamic.

[0261] In actual implementation, when the target interactive control is a camera movement parameter setting control, in response to a trigger operation on the camera movement parameter setting control, a camera movement parameter setting area is displayed, and the parameter setting area is used to set the camera movement parameters of the first video; based on the camera movement parameter setting area, in response to the camera movement parameter setting operation, the camera movement parameters set for the first video are displayed; wherein the camera movement parameters include the camera's movement direction and movement speed, etc.; in response to a determination instruction on the camera movement parameters, a video obtained by adjusting the parameters of the first video based on the set camera movement parameters is displayed.

[0262] In actual implementation, when the target interactive control is a background removal control, in response to the triggering operation of the background removal control, the background of the first video is removed to obtain a ninth video; wherein the background of the first video can be, for example, the sky, a mountain, etc.

[0263] In actual implementation, when the target interactive control is an object erasing control, in response to a triggering operation on the object erasing control, an erasing tool in a draggable state is displayed, and the erasing tool can be moved on the image frame of the first video to select the target object included in the corresponding image frame, thereby deleting the selected target object; in response to a selection operation triggered by the erasing tool, the selected target object is displayed on the corresponding image frame; in response to a confirmation instruction for the target object, the selected target object is deleted; in response to a video editing instruction, the first video with the target object deleted is displayed.

[0264] In actual implementation, when the target interactive control is a scene detection control, in response to a trigger operation on the scene detection control, the first video is split into at least one video segment based on at least one scene included in the first video, wherein each video segment corresponds to a scene.

[0265] In actual implementation, when the target interactive control is a depth of field setting control, in response to the depth of field setting operation triggered by the depth of field setting control, the target depth of field parameters obtained by the setting are displayed, and thus in response to the determination instruction for the target depth of field parameters obtained by the setting, the first video with the target depth of field parameters is displayed, that is, the video obtained by adjusting the parameters of the first video based on the set target depth of field parameters is displayed.

[0266] In actual implementation, when the target interactive control is a frame rate adjustment control, in response to the frame rate adjustment operation triggered by the frame rate adjustment control, the target frame rate obtained by adjustment is displayed, and thus in response to the determination instruction for the target frame rate obtained by adjustment, the first video with the target frame rate is displayed, that is, the video obtained by adjusting the frame rate of the first video based on the set target frame rate is displayed.

[0267] In actual implementation, when the target interactive control is the third new video generation control, in response to the trigger operation for the third new video generation control, a tenth video is generated based on the audio data of the first video, and the content of the tenth video is associated with the audio data.

[0268] It should be noted that, in addition to interacting based on the first video, other videos or audios can also be selected for interaction; that is, at least one interactive control is displayed, and in response to a triggering operation on a target interactive control in at least one interactive control, at least one candidate video is displayed; in response to a selection operation on a target candidate video in at least one candidate video, the target candidate video is controlled to be in a selected state; in response to a determination instruction on the target candidate video in the selected state, the interactive operation indicated by the target interactive control is performed on the target candidate video; wherein, when the target interactive control is a third new video generation control, at least one candidate video also includes audio; thereby, in response to a determination instruction on the target audio or target candidate video in the selected state, a video generated based on the audio data of the target audio or target candidate video is displayed.

[0269] In actual implementation, interactive operations on videos can be triggered not only by controls, but also by text input boxes. As described above, when the target interactive control is a local edit control, in response to the triggering operation on the local edit control, the local edit box and the text input box will be displayed, or the text input box can be displayed while displaying at least one interactive control for the first video; thus, in response to the input operation based on the text input box, the input text content is displayed, and the text content is used to describe the editing task; in response to the video editing instruction, the video obtained by interacting with the first video based on the text content and at least one of the interactive controls is displayed. For example, content such as "change the picture to warmer colors" can be entered in the text input box to change the picture filter effect of the first video, or content such as "delete XX object" can be entered to delete XX object in the first video.

[0270] In practical applications, different interactive operations can be performed on the generated first video based on different interactive controls, which not only optimizes the effect of the generated video, but also improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0271] Applying the above-mentioned embodiment of the present application, at least one video generation control for generating videos of different video types is first displayed, and then based on the at least one video generation control, in response to the target video type being selected in at least two video types, a content input area corresponding to the target video type is displayed, and based on the content input area, the target content is input, thereby generating a first video of the target video type based on the input target content. In this way, not only is a video of the target video type corresponding to the target content generated based on the target content input by the user, so that the video content of the target video meets the user's needs, but also, based on the video generation control, a video of the target video type is generated, so that the type of video also meets the user's needs. Compared with the solution in which the user needs to edit the video to generate the final video, the present application reduces the editing operations that the user needs to perform on the video, which not only improves the efficiency of video generation, but also improves the efficiency of human-computer interaction and the hardware resource utilization of the electronic device.

[0272] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0273] There are many types of videos, such as short videos and long videos. Different types of videos have different corresponding video parameters and different video generation processes. However, when generating videos, related technologies mostly directly generate videos based on the user's demand for video content, that is, they only consider the user's demand for video content, but ignore the user's demand for video type. In this way, after generating the video, the user still needs to edit the video so that the edited video meets the user's requirements for video content and the user's requirements for video type. This not only makes the human-computer interaction efficiency too low, but also further reduces the video generation efficiency.

[0274] Based on this, an embodiment of the present application provides a video generation method. On the one hand, it comprehensively combines multimodal capabilities. In addition to supporting pure text-to-video, it also supports picture-to-video, picture-to-text-to-video, video-to-video, and other capabilities; on the other hand, by combining the text splitting and storyboard script generation capabilities and the short video generation and splicing capabilities, a short video of 2 to 4 seconds can be generated by inputting text, or after inputting text (after the text is subsequently modified, the storyboard will also change with the modified text), the model automatically splits the text into multiple storyboard scripts. Multiple storyboard scripts support editing behaviors such as addition, deletion, and modification, and also support custom selection of the correlation between different storyboards, thereby supporting the generation of longer videos. At the same time, the intermediate steps and information of the long video generation process will be displayed, supporting users to make adjustments at any time; in addition, you can also enter different advanced gameplay functions through the laboratory tab (extended editing), supporting video mask selection editing, canvas expansion, skeleton point video generation and other functions.

[0275] Next, we will explain the technical solution of this application from the product side. The technical solution product side of this application consists of two parts: the first part is video creation, and the second part is interactive video editing. The process of creating videos is divided into two steps, namely the process of creating short videos and the process of creating long videos.

[0276] For the process of creating short videos, as shown in Figure 4, the goodcase (video template) is displayed in a waterfall flow below the input box, including video and text. When the target video template is selected, the content description of the target video template and the corresponding secondary creation button (first editing control) are displayed. Then, when a click operation on the secondary creation button is received, the content description of the target video template is filled in the input box.

[0277] In actual implementation, in response to a user input operation, the immediate generation control (first determination control) in FIG4 is controlled to switch from an inactive state to an active state, wherein the input operation can be at least one of inputting text, images, and audio. Here, when inputting an image, the image is used as a backing image for the video. At the same time, only a single image can be uploaded, and re-uploading, image preview, and image deletion are supported.

[0278] It should be noted that, in response to the user's triggering operation on the parameter setting control as shown in Figure 4, a parameter setting area as shown in Figure 11 can be displayed, so that the parameters of the video can be set in response to the parameter setting operation triggered by the parameter setting area. If the user does not perform the parameter setting operation, the default parameter configuration is used, for example, the picture defaults to none, the special effect defaults to none, the video quality defaults to 896*512, the ratio defaults to 7:4, the video defaults to 4s, and the audio defaults to none.

[0279] In actual implementation, in response to the user's input operation, the control of the immediate generation control is switched from an inactive state to an active state, and in response to the parameter setting operation triggered based on the parameter setting area, after setting the parameters of the video, when a trigger operation is received for the immediate generation control in the activated state, the first waiting prompt information as shown in Figure 19 is displayed. The first waiting prompt information is used to indicate that the short video to be created is in a queue state. This state is the waiting queue state after clicking the immediate generation control to the start of generation. The corresponding state is waiting in the queue, and the estimated countdown task number will also be displayed.

[0280] In actual implementation, when the queuing state ends, the first progress prompt information as shown in Figure 18 is displayed. The first progress prompt information is used to indicate that the short video to be created is in the generation state. This state is the generation process state from the start of generation to the return of results, and the video generation progress will also be displayed.

[0281] It should be noted that in the intermediate state, i.e., the queue state and the generating state, modification operations in the interface are not allowed, including text, pictures, parameter configuration, etc.; but users are allowed to create new videos, such as when creating a new short video, in response to a click operation on the secondary creation control or the create new video control as shown in Figures 18 or 19, such as in response to a click operation on the secondary creation control, an interface as shown in Figure 8 b is displayed, and the input content corresponding to the video to be created is filled back into the input box as new input content, thereby generating a video associated with the video to be created based on the new input content; in response to a click operation on the create new video control, an interface as shown in Figure 8 a is displayed, wherein the input box (content input area) in the interface is blank, so that the user re-enters the new content, and then generates a video based on the input new content. Among them, when a click operation on the secondary creation control or the create new video control as shown in Figures 18 or 19 is received, the currently ongoing task is collected into the assets.

[0282] In actual implementation, three videos will be generated each time during the video generation process, and you can select one of them to preview in the main window; at the same time, you can also perform like operations, dislike operations, share operations, and download operations on the generated video.

[0283] It should be noted that when all videos fail to generate, the generation failure interface shown in Figure 22 is displayed. This interface includes a regeneration control for regeneration. When the user clicks the regeneration control, the video is regenerated; at the same time, the videos that failed to generate are also saved as assets. When some videos fail to generate, the successfully generated videos are returned normally, and the failed videos are displayed as failed on the front end. At the same time, for some videos that failed to generate, since there are successfully generated videos, the asset is considered to have been generated successfully and regeneration is not supported. In other words, as long as at least one video is generated successfully, the asset is considered to have been generated successfully.

[0284] In actual implementation, the video style can also be set. As shown in Figure 23, a style setting control is displayed. In response to a trigger operation on the style setting control, at least one optional video style is displayed. In response to a selection operation on a target video style from at least one optional video style, a video corresponding to the target video style is generated.

[0285] The process of creating long videos needs to be generated in stages, which specifically include four stages. The following describes the four stages one by one.

[0286] 1) Phase 1: Input content and configure global parameters.

[0287] As shown in c and d in Figure 10, when no user input operation is received, the next step cannot be clicked, that is, the next button is grayed out (inactive state). When the user input operation is received, the next button can be clicked (active state); wherein the input operation can be at least one of inputting text, pictures and audio.

[0288] In actual implementation, a parameter configuration area as shown in Figure 11 will also be displayed, so that the parameters of the video can be set in response to the parameter setting operation triggered by the parameter setting area. If the user does not perform a parameter setting operation, the default parameter configuration will be used. For example, the default video length is 6s, the default video quality is 896*512, the default ratio is 7:4, and there is no audio.

[0289] 2) Phase 2: Generate segmented storyboard descriptions, allowing users to edit them.

[0290] In actual implementation, as shown in FIG12 , a storyboard description is generated (a storyboard disassembly interface is displayed), which includes a description of the video segment and corresponding setting parameters. Each description supports secondary editing, i.e., editing text, adjusting parameter settings (adding special effects), adding pad images (key image frames), etc. The storyboard supports users to adjust the storyboard order. At the same time, users can also customize new storyboards and corresponding texts. After adding a storyboard, a storyboard script can be automatically generated, or customized input can be provided. Finally, after all adjustments are completed, in response to a click operation on the Generate control (the second confirmation control), the storyboard generation interface is entered.

[0291] It should be noted that in the case of storyboard generation failure, when all generation fails, a generation failure message (generation failure prompt information) is displayed, and the video can be regenerated.

[0292] 3) Stage 3: Generate storyboards.

[0293] In actual implementation, as shown in Figure 13, storyboard shots are generated (the storyboard generation interface is displayed). Each shot contains the storyboard text description, parameter configuration, and the corresponding three generated storyboard videos (candidate storyboards). Each storyboard supports secondary editing, namely editing text, adjusting parameter settings, adding pad images (key image frames), and supporting the selection of whether the next level storyboard wants to be associated with the previous level. The parameter adjustment settings include adjusting special effects and uploading pictures. The storyboard order between different storyboards can be adjusted; at the same time, the corresponding storyboard can be regenerated after adjustment.

[0294] It should be noted that when users select storyboards (multiple selections are supported) and the specific videos under the corresponding storyboards (single selection), the default is to select all storyboards and select the first video under each storyboard. If no storyboard is selected, the button to combine videos (the third confirmation control) will not be clickable. At the same time, it supports regenerating all storyboards with one click. Users can also upload an existing short video to generate a video (the first video of the target video type) together with the storyboards on the interface.

[0295] In actual implementation, when some storyboards are generated successfully, the successfully generated storyboards are returned normally, and the failed storyboards are displayed on the front end; as long as there are more than or equal to 1 successfully generated candidate storyboards under the corresponding storyboard text, it is considered to be successfully generated.

[0296] 4) Stage 4: Synthesize long video.

[0297] In actual implementation, in response to the user's click operation on the synthetic video control (the third confirmation control), the storyboard, sequence information and related information selected by the user are transmitted to the background for video synthesis; at the same time, the waiting state and the generation failure situation are the same as the process when generating a short video; in addition, the generated video can also be liked, disliked, shared and downloaded.

[0298] 5) The information of the completed stage of the previous order can be viewed by returning (based on the backtracking control), but it cannot be edited.

[0299] It should be noted that if there is overlapping information in phase 2 and phase 3, if it is modified in phase 3 but regeneration is not triggered, it will be out of sync in phase 2. If regeneration is triggered, it will be synchronized to phase 2.

[0300] For interactive editing of videos, including but not limited to video support for style conversion and regeneration (style setting control), video canvas expansion (canvas expansion control), addition, deletion and modification of elements in the video (local editing control), such as adding specific subjects by circling specific areas and editing prompt words, motion brushes, camera parameter settings, video generation time extension, multi-image video, skeleton point driven video generation (action sequence generation control), video generation, such as the ability to generate new videos through 3 modes (i.e. video + picture, video + style selection, video + text), storyboard function (such as after generating a video, you can choose to add a new scene, re-fill in the prompt words, add a new plot to the entire video, and set each plot separately). settings, the newly added scene does not support uploading pictures again, and you can also delete a scene), video editing tools (background removal controls), object erasing tools (object erasing controls), color grading (this function only needs to enter descriptive text to easily process the video screen filter effects), super slow motion tools (frame rate adjustment controls, this function can reduce the frame rate of your video lens and convert it into a smooth slow motion video), video depth of field tools (depth of field setting controls. This function can automatically process the video into a picture with depth of field effect), scene detection (scene detection controls, this function can automatically split the lens of your uploaded video into multiple clips) and audio generation video (the third new video generation control).

[0301] In actual implementation, all videos that have been interactively edited are imported into a unified asset tab, and the function source of each video is identified, supporting quick jumps from each function tab; each function stores its own function assets; at the same time, for the successfully generated assets under the asset tab, a quick entry for advanced editing is added. Clicking the quick entry directly displays the interactive editing function interface (i.e., skipping the video upload step).

[0302] In actual implementation, the interactive editing process includes selecting / uploading videos, editing, generating, and storing assets. For video uploading, both pulling video assets and local uploading are supported, with single selection supported. Pulling existing assets displays all successfully generated videos under the Assets tab, allowing users to select. Local uploading, on the other hand, indicates that local video uploading is supported. After uploading, the video can be cropped to a specific duration, but the cropping duration is limited to integers of 2, 3, or 4 seconds. Custom cropping to non-integer durations such as 2.5 seconds is not supported.

[0303] In actual implementation, for the functions in interactive editing, the input source is all video. After uploading a video, the input video is interactively edited based on the interactive editing function interface; for the process of local video editing, as shown in Figure 26, during the video playback, the user is supported to select the required area. The video is automatically paused when the selection is made. After the selection is completed, a prompt word editing box (text input box) and a selector pop up, among which the editing box supports filling in prompt words, and the selector supports selection of forward effect, backward effect, all effect, etc. If no selection is made, the default selection is all effect; among which, forward effect indicates that the video from the start of the video to the current time includes the prompt word change effect, backward effect indicates that the video from the current time to the end of the video includes the prompt word change effect, and all effect indicates that the entire video includes the prompt word change effect; then perform general parameter settings, such as supporting the setting of special effects, video quality, sound effects, style, and setting the logical multiplexing video generation process.

[0304] As shown in Figure 27, after uploading the video, you can select an expansion ratio, such as 16:9, 9:16, or 1:1. After selecting the ratio, a canvas of the corresponding size is displayed, and the video is displayed in the middle of the canvas. If the original video size is larger than the canvas size, you need to scale the original video proportionally to keep the entire video within the canvas. You can then adjust the video size and position, which means you can adjust the video size and drag and drop to adjust its position within the canvas, but you must ensure that the entire video is within the canvas and cannot exceed the canvas. Then, perform regular parameter settings, such as supporting special effects, video quality, sound effects, and style, and setting the logic for multiplexing the video generation process.

[0305] In actual implementation, for the process of generating video during canvas expansion, in response to the click operation on the generate video control, the coordinate information of the video quality, ratio, and selected area is converted and transmitted to the server, and then based on the data returned by the server, the status of the video is determined, that is, the queue and generation process status. For the queue status, this state is the waiting queue state from clicking generate to the start of generation, and the corresponding state is waiting in queue. At the same time, the front end needs to display the estimated countdown task number; for the generating state, this state is the generation process state from the start of generation to the return of results.

[0306] In actual implementation, after the queue and generation process states, the generated video is received from the server and added to the Assets tab. This way, the newly added interactively generated video is included under the Assets tab, allowing the video to continue the non-interactive editing process, as described above.

[0307] It should be noted that after generating a video, you can also perform like operations, dislike operations, share operations, and download operations on the generated video; at the same time, when the generation fails, it also supports regeneration of the video.

[0308] Regarding the process of skeleton point driven video generation, the specific process of generating a video includes selecting an action template such as a dance action template, uploading an image, text description, and parameter setting. First, at least one dance video template (candidate action sequence) is displayed, and then a dance video template is selected from the at least one dance video template; then, the user performs input operations, such as uploading a local image and / or adding a text description, and then performs general parameter settings, such as supporting the setting of special effects, scale, video quality, sound effects, style, and setting the logic multiplexing video generation process;

[0309] For the process of generating videos in the process of skeleton point driven video generation, in response to the click operation on the generate video control, the dance template selected by the user and the information entered by the user are transmitted to the server, and then based on the data returned by the server, the status of the video is determined, that is, the queue and generation process status. For the queue status, this state is the waiting queue state from clicking generate to starting to enter the generation, and the corresponding state is waiting in the queue. At the same time, the front end needs to display the estimated countdown task number; for the generating state, this state is the generation process state from starting to return the result.

[0310] In actual implementation, after the queue and generation process, the generated video is received from the server and added to the Assets tab. This way, the Assets tab includes the video generated by the newly added interactive gameplay, allowing you to continue editing the video in non-interactive gameplay, as described in the video generation process above.

[0311] It should be noted that after generating a video, you can also perform like operations, dislike operations, share operations, and download operations on the generated video; at the same time, when the generation fails, it also supports regeneration of the video.

[0312] In actual implementation, videos in assets (asset videos) can be shared and downloaded. Only successfully generated videos can be shared, and videos in the generation process cannot be shared. At the same time, the shared content includes the generated video, the corresponding parameter configuration, text, and backing image information of the video;

[0313] In actual implementation, you can also continue to generate videos based on the videos in the assets, display a continue generation button (regenerate control), click it to jump to the editor interface (video generation interface), and at the same time fill all information back into the editor (content input area), supporting users to continue editing and generating.

[0314] In actual implementation, you can also perform secondary creation based on the video in the asset. For the generated video, a re-generate button (fifth editing control) is displayed. Clicking it will jump directly to the editor corresponding to the long / short video; backfill the corresponding text, pictures, and parameter configuration information, and then respond to the click operation on the generate control to display the video generated based on the backfilled content; the generated video at this time is a new video, and the generated video is stored in the asset module as a new asset video, without overwriting the original asset video. For the secondary creation of long videos, the information of stages one, two, and three is backfilled, and the process goes directly to stage three.

[0315] In actual implementation, you can also delete videos in assets. Both assets in the process of generation and completed assets support deletion. The deletion status is transmitted from the front end to the back end. If the back end determines that the asset is synthesized / generation has started, it will directly delete the data; if it is an asset video to be produced, the video generation request will be canceled.

[0316] Next, the technical solution of this application is explained from a technical perspective.

[0317] In actual implementation, referring to Figure 28, Figure 28 is a structural diagram of the video generation model provided by an embodiment of the present application. Based on Figure 28, the general video model adopted by the embodiment of the present application is a model for conditionally generating video based on a diffusion model. Unlike the classical Wensheng graph diffusion model, the embodiment of the present application is designed to be a 3D Unet network, which mainly includes a main network indicated by the dotted box 2802 in Figure 28 and a conditional control network indicated by the dotted box 2801. The conditional control network can encode and model a variety of input reference conditions (such as text, images, videos, skeletons, etc.), and send the encoded reference condition information to the main network to guide the main network learning. The main network mainly learns to predict the noise situation at each time step, so that the noise data can be subsequently predicted and denoised to obtain the final clear video frame. In the main network and the conditional control network, an additional timing modeling module is added to learn the timing-related motion information in video generation.

[0318] In actual implementation, the training process first inputs the video-text training data, and the training video is composed of B*T*C*H*W input videos (set as z0), where B is the training batch size, T is the number of frames of the generated video, and C, H, and W are the channels, height, and width of the video frames; among them, in some diffusion models based on latent variables, C, H, and W represent the channels, height, and width of the latent variables; then the noise is input, and the input video is denoised according to the commonly used denoising strategy and the number of denoising steps t, and then fed into the model, that is, z t =α t z0+δ t ε...Formula (1);

[0319] Where ∈ is Gaussian noise, α and δ are noise coefficients, and z0 is the input video.

[0320] Then perform video spatiotemporal modeling, with noisy input z t After entering the main network, the conv convolution in each 3D residual block and the transformer module in the Attention block will model and learn the input space and time sequence. Then, conditional modeling is performed, using common encoding modules such as CLIP to extract the embedding features of the text, and then the cross attention module in the Attention block performs attention calculation with the original input to guide the noise prediction of the 3D Unet; the image, video, and bone node information are conditionally encoded through the conditional control module, and then the output of different 3D residual modules is spliced ​​with the information in the main network through conditional control to guide the noise prediction of the 3D Unet. Finally, noise prediction is performed, and the noise input z is added. t After various conditional information is input into the 3D Unet of this application, the noise added at the noise addition step number t is predicted. The loss function is the MSE Loss of the real noise input. The network parameters are continuously updated by minimizing the loss.

[0321] In actual implementation, after the model training is completed, sampling can be performed. The sampling process is to use the existing conditional information (text, image, video, skeleton, etc.) to continuously predict the noise that needs to be removed in the random noise; then, denoising is performed to finally generate a video. t The images are input into the trained 3D Unet and the noise is removed after multiple rounds of noise prediction to obtain the final clear and noise-free video frames, which are then combined at a certain frame rate to obtain the final video.

[0322] In practical implementation, the local video editing model used in the interactive editing process, also known as the model for implementing local video editing, shares a similar structure to the general video model described above, with some differences in the inputs of the conditional control module. Local video editing requires the user to upload a video or directly use a generated video, and then specify the location of the local area to be edited. The local area typically specifies the top left and bottom right coordinates of a rectangular box.

[0323] In actual implementation, during the training and sampling process, refer to Figure 29, which is a schematic diagram of the sampling process provided by an embodiment of the present application. Based on Figure 29, the conditional control encoding module will mask the area specified by the user in the original video (filled with 0) to obtain a 0-1 binary image and Gaussian noise, and connect them according to the channel position and input them into the conditional control module to guide the learning of the main model. Then, the rest of the training and sampling processes can refer to the above description.

[0324] The skeleton-guided dance video model used in the interactive editing process, which generates videos driven by skeleton points, builds on the general video model by adding two types of conditional control modules: a skeleton sequence conditional control module and a given image conditional control module. The skeleton sequence is a set of dance skeleton templates provided by the platform, while the given image is a user-uploaded image of a human body.

[0325] See Figure 30, which is a schematic diagram of the processing process of the skeleton-guided dance video model provided in an embodiment of the present application. Based on Figure 30, the skeleton-guided dance video model supports two functions. The first function is to generate a video based on text and skeleton information. The text prompts here are mainly descriptions of the characters and background of the generated video, and the actions will be generated according to the skeleton sequence; the skeleton information is conditionally controlled using a lightweight convolutional network, and the text is directly encoded using the text encoder of the general model. The second function is to generate a video based on an image and skeleton information. The appearance of the characters in the generated video is consistent with the given image; the image information here is introduced in two forms: a control network and an image adapter, and then the rest of the training and sampling processes can refer to the above description.

[0326] Thus, this application proposes a relatively general video generation model based on various conditional controls. This model not only ensures that spatial and temporal motion modeling can be well learned during the video generation process (ensuring clear video quality and reasonable motion without distortion), but also uniformly utilizes various conditional information (text, images, videos, skeletons, etc.) to guide the final video generation process.

[0327] Applying the above-mentioned embodiment of the present application, at least one video generation control for generating videos of different video types is first displayed, and then based on the at least one video generation control, in response to the target video type being selected in at least two video types, a content input area corresponding to the target video type is displayed, and based on the content input area, the target content is input, thereby generating a first video of the target video type based on the input target content. In this way, not only is a video of the target video type corresponding to the target content generated based on the target content input by the user, so that the video content of the target video meets the user's needs, but also, based on the video generation control, a video of the target video type is generated, so that the type of video also meets the user's needs. Compared with the solution in which the user needs to edit the video to generate the final video, the present application reduces the editing operations that the user needs to perform on the video, which not only improves the efficiency of video generation, but also improves the efficiency of human-computer interaction and the hardware resource utilization of the electronic device.

[0328] The following further describes an exemplary structure of the video generating device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, as shown in FIG2 , the software modules stored in the video generating device 455 in the memory 450 may include:

[0329] A first display module 4551 is configured to display at least one video generation control in the video generation interface, wherein the at least one video generation control is used to generate videos of at least two video types;

[0330] A second display module 4552 is configured to generate a control based on the at least one video, and in response to a target video type being selected from the at least two video types, display a content input area corresponding to the target video type;

[0331] a third display module 4553 configured to display the input target content in response to a content input operation performed in the content input area;

[0332] The fourth display module 4554 is configured to display a first video of the target video type generated based on the target content in response to a video generation instruction.

[0333] In some embodiments, the number of the video generation controls is at least two, and different video generation controls correspond to different video types. The device also includes a determination module, which is configured to display a video generation interface including the at least two video generation controls in response to an opening instruction for the video generation interface; and determine that the target video type is in a selected state in response to the video generation control corresponding to the target video type in the video generation interface being in a selected state.

[0334] In some embodiments, the device also includes a parameter setting module, which is configured to display a parameter setting area in the video generation interface, and the parameter setting area is used to set the parameters of the first video; based on the parameter setting area, in response to the parameter setting operation, the video parameters set for the first video of the target video type are displayed; the fourth display module 4554 is also configured to respond to the video generation instruction and, based on the target content, display the generated first video of the target video type with the video parameters.

[0335] In some embodiments, the target video type is a long video type, and the first video of the long video type includes at least two storyboards; the device also includes a fifth display module, and the fifth display module is configured to display a storyboard disassembly interface in response to a determination instruction for the target content, and display at least two storyboard texts determined based on the target content in the storyboard disassembly interface; wherein each of the storyboard texts corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; the fourth display module 4554 is also configured to display the generated first video of the long video type including the storyboards corresponding to each of the storyboard texts in response to a video generation instruction triggered based on the at least two storyboard texts.

[0336] In some embodiments, the device further includes a sixth display module, which is configured to display a storyboard text quantity adjustment control, wherein the storyboard text quantity adjustment control is used to add or delete storyboard texts; in response to a storyboard text addition operation triggered based on the storyboard text quantity adjustment control, the added new storyboard text is displayed on the storyboard disassembly interface; in response to a determination operation triggered based on at least two storyboard texts including the new storyboard text, the video generation instruction is received.

[0337] In some embodiments, the target video type is a long video type, and the first video of the long video type includes multiple storyboards; the device also includes a seventh display module, and the seventh display module is configured to display a storyboard generation interface in response to a determination instruction for the target content, and display at least two storyboard texts and at least two storyboards in the storyboard generation interface, each of the storyboard texts corresponding to one storyboard, and the storyboards are generated based on the corresponding storyboard texts; the fourth display module 4554 is also configured to display the generated first video of the long video type including the at least two storyboards in response to a video generation instruction based on the at least two storyboard texts and the at least two storyboards.

[0338] In some embodiments, the device also includes a control module, which is configured to control the first target storyboard to be in a selected state in response to a selection operation on the first target storyboard among the at least two storyboards; the fourth display module 4554 is also configured to display the generated first video of the long video type including the first target storyboard based on the first target storyboard in the selected state and the storyboard text of the first target storyboard in response to the video generation instruction.

[0339] In some embodiments, the at least two storyboards have an arrangement order, and the first videos constituted by the at least two storyboards with different arrangement orders are different, and a sorting adjustment control is also displayed in the storyboard generation interface; the device also includes a sorting module, and the sorting module is configured to change the sorting of the at least two storyboards from the current arrangement order to the target arrangement order in response to a sorting adjustment operation for the at least two storyboards triggered based on the sorting adjustment control; the fourth display module 4554 is also configured to display the first video of the long video type including the at least two storyboards generated based on the target arrangement order in response to a video generation instruction based on the at least two storyboard texts and the at least two storyboards.

[0340] In some embodiments, the storyboard generation interface further displays a storyboard update control for updating the content of the storyboard; the device further includes a storyboard update module, which is configured to update the second target storyboard in the at least two storyboards to a first new storyboard based on the storyboard update control in response to a content update operation on the second target storyboard, wherein the content of the first new storyboard is associated with the content of the second target storyboard.

[0341] In some embodiments, the at least two storyboards are arranged in an order, and an upward association control is present for a third target storyboard among the at least two storyboards, wherein the upward association control is used to update the content of the third target storyboard based on the content of a previous storyboard of the third target storyboard, and the third target storyboard is any one of the at least two storyboards except the first storyboard; the device further includes an upward association module, wherein the upward association module is configured to update the third target storyboard to a second new storyboard in response to a triggering operation on the upward association control.

[0342] In some embodiments, the device also includes an eighth display module, and the eighth display module is configured to display function introduction information when the account that performs the triggering operation generates the first video of the long video type for the first time; wherein, the function introduction information is used to introduce the function of the upward-associated control.

[0343] In some embodiments, the seventh display module is further configured to display at least two storyboard texts in the storyboard generation interface, and display at least two candidate storyboards corresponding to each of the storyboard texts; for each of the storyboard texts, in response to a selection operation on a target candidate storyboard among the at least two candidate storyboards, control the target candidate storyboard to be in a selected state, and determine the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

[0344] In some embodiments, the device also includes a ninth display module, which is configured to display at least one video template in the video generation interface; in response to a selection operation on a target video template in the at least one video template, display a template content description of the target video template and a corresponding first editing control; wherein the first editing control is used to generate a first video of the target video type based on the template content description; in response to a trigger operation on the first editing control, determine the trigger operation as an input operation on the content input area; the third display module 4553 is also configured to display the template content description of the target video template in the content input area in response to a trigger operation on the first editing control, and determine the displayed template content description as the input target content.

[0345] In some embodiments, the device further includes a tenth display module, which is configured to display at least one video template in a video generation interface; in response to a trigger operation on a target video template in the at least one video template, display a detail interface including detail information of the target video template; wherein the detail interface further includes at least one of a second editing control and a first new video generation control, the second editing control being used to generate a first video of the target video type based on the template content description, the first new video generation control being used to re-enter the target content in the content input area, and generate the first video of the target video type based on the re-entered target content; in response to a trigger operation on the second editing control A trigger operation is performed to jump from the details interface to the video generation interface including the content input area of ​​the template content description carrying the target video template; in response to the trigger operation for the first new video generation control, the video generation interface is jumped from the details interface to the video generation interface including the blank content input area; the third display module 4553 is also configured to, when receiving a trigger operation for the second editing control, determine the trigger operation as the content input operation, and determine the template content description displayed in the content input area as the input target content; when receiving a trigger operation for the first new video generation control, in response to the content input operation performed in the blank content input area, display the input target content.

[0346] In some embodiments, the fourth display module 4554 is further configured to display a first video of the target video type generated based on the target content in a video display interface in response to the video generation instruction; wherein, the video display interface also includes at least one of the following: detailed information of the generated first video, a third editing control, and a second new video generation control; wherein, the third editing control is used to regenerate a second video of the target video type based on the target content, and the second new video generation control is used to re-enter the target content in the content input area, and generate a third video of the target video type based on the re-entered target content.

[0347] In some embodiments, the video display interface includes the third editing control, the target video type is a long video type, and the first video of the long video type includes at least two storyboards; the device also includes an eleventh display module, and the eleventh display module is configured to display a storyboard generation interface in response to a trigger operation on the third editing control; display at least two storyboard texts determined based on the target content on the storyboard generation interface, and display at least two storyboards determined based on the at least two storyboard texts; based on the storyboard generation interface, in response to an editing operation on the at least two storyboard texts and the at least two storyboards, display the edited at least two storyboard texts and the at least two storyboards; in response to a video generation instruction, display a second video of the long video type generated based on the edited at least two storyboard texts and the at least two storyboards.

[0348] In some embodiments, the video display interface includes the third editing control, and the target video type is a short video type; the device also includes a twelfth display module, and the twelfth display module is configured to display a video generation interface including a content input area in response to a trigger operation on the third editing control, in which the target content is displayed; in response to an editing operation on the target content, display new target content obtained by editing the target content; and in response to a video generation instruction, display a third video of a short video type generated based on the new target content.

[0349] In some embodiments, the video generation interface also includes a video display control, which is used to display at least one asset video, including at least one of a successful video of historical generation success, a failed video of historical generation failure, a generating video in the process of being generated, and a waiting video waiting to be generated; the device also includes a thirteenth display module, which is configured to display at least one asset video including a first video of the target video type in response to a trigger operation on the video display control.

[0350] In some embodiments, the at least one asset video includes a failure video, and the failure video is associated with a regeneration control, and the regeneration control is used to regenerate the failure video; the device also includes a regeneration module, and the regeneration module is configured to regenerate the target failure video in response to a trigger operation of the regeneration control associated with the target failure video in at least one failure video; when the regeneration is successful, the failure video in the at least one asset video is switched to a regenerated fourth video; when the regeneration fails, a regeneration failure prompt message is displayed.

[0351] In some embodiments, the at least one asset video includes a waiting video, and the device also includes a fourteenth display module, which is configured to display a first waiting prompt information in an associated area of ​​the waiting video, and the first waiting prompt information is used to indicate that the corresponding waiting video is waiting to be generated; in response to a trigger operation for the waiting video triggered based on the first waiting prompt information, a waiting cancel control is displayed, and the waiting cancel control is used to terminate the waiting process of the waiting video.

[0352] In some embodiments, the video generation interface also includes a style setting control, which is used to set the video style of the generated first video; the device also includes a style setting module, which is configured to display at least one video style in response to a trigger operation on the style setting control; in response to a selection operation on a target video style in the at least one video style, control the target video style to be in a selected state; the fourth display module 4554 is also configured to display a first video of the target video type with the target video style generated based on the target content in response to a video generation instruction.

[0353] In some embodiments, the style setting module is further configured to, in response to a determination instruction for the target video style in a selected state, display an identifier of the target video style on the style setting control, the identifier being associated with a style cancel control; wherein the identifier of the target video style is used to identify that the generated first video of the target video type has the target video style; in response to a triggering operation for the style cancel control, cancel the display of the identifier of the target video style on the style setting control; the fourth display module 4554 is further configured to, in response to a video generation instruction, display a first video of the target video type with a default video style generated based on the target content.

[0354] In some embodiments, the fourth display module 4554 is also configured to display a second waiting prompt information in response to a video generation instruction, wherein the second waiting prompt information is used to indicate the length of time required to wait before starting to generate the first video; when the waiting prompt information indicates that the length of time required to wait before starting to generate the first video is zero, the second waiting prompt information is canceled; and the first video of the target video type generated based on the target content is displayed.

[0355] In some embodiments, the device also includes an interaction module, which is configured to display at least one interaction control for the first video, and the interaction control is one of the following controls: a local editing control for local editing of the first video, a canvas expansion control for editing the size of the first video, a motion brush control for changing a static object in the first video into a dynamic object, a camera parameter setting control for setting the camera parameters of the first video, a background removal control for removing the background of the first video, an object erasing control for erasing a target object in the first video, a scene detection control for dividing the video into different video segments based on different scenes in the first video, a depth of field setting control for setting a depth of field effect for the first video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the first video, an action sequence generation control for setting an action sequence for an object in the first video, and a third new video generation control for generating a new video based on the audio data in the first video; in response to a trigger operation on a target interaction control among at least one of the interaction controls, the interaction operation indicated by the target interaction control is performed on the first video of the target video type.

[0356] The present invention provides a computer program product including computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the video generation method described in the present invention.

[0357] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the video generation method provided by an embodiment of the present application, for example, the video generation method shown in Figure 3.

[0358] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a CD-ROM; or various devices including one or any combination of the above memories.

[0359] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0360] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0361] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0362] In summary, the embodiments of the present application have the following beneficial effects:

[0363] (1) Not only is a target video type video generated based on the target content input by the user, so that the video content of the target video meets the user's needs, but also a target video type video is generated based on the video generation control, so that the type of the video also meets the user's needs. Compared with the solution in which the user needs to edit the video to generate the final video, the present application reduces the editing operations that the user needs to perform on the video, which not only improves the video generation efficiency, but also improves the human-computer interaction efficiency and the hardware resource utilization of the electronic device.

[0364] (2) It not only ensures that the spatial and temporal motion modeling can be well learned during the video generation process (ensuring clear video quality, reasonable motion and no distortion); at the same time, it can also uniformly use various conditional information (text, images, videos, skeletons, etc.) to guide the final video generation process.

[0365] (3) Improve the relevance of content and effects between storyboards based on upward association controls.

[0366] It should be noted that in the embodiments of the present application, when it comes to obtaining user operation data, target content input by the user and other related data, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0367] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A video generation method, which is executed by an electronic device, and the method includes: In a video generation interface, display at least one video generation control, where the at least one video generation control is used to generate videos of at least two video types; Based on the at least one video generation control, in response to a target video type among the at least two video types being in a selected state, display a content input area corresponding to the target video type; In response to a content input operation performed in the content input area, display the input target content; In response to a video generation instruction, display a first video of the target video type generated based on the target content.

2. The method according to claim 1, wherein, The number of the video generation controls is at least two, and different video generation controls correspond to different video types. Before displaying at least one video generation control in the video generation interface, the method further includes: In response to an open instruction for the video generation interface, display a video generation interface including the at least two video generation controls; Before in response to a target video type among the at least two video types being in a selected state and displaying a content input area corresponding to the target video type, the method further includes: In response to a video generation control corresponding to the target video type in the video generation interface being in a selected state, determine that the target video type is in a selected state.

3. The method according to claim 1 or 2, wherein The method further includes: Display a parameter setting area in the video generation interface, where the parameter setting area is used to set parameters of the first video; Based on the parameter setting area, in response to a parameter setting operation, display the video parameters set for the first video of the target video type; The step of in response to a video generation instruction and displaying a first video of the target video type generated based on the target content includes: In response to a video generation instruction, based on the target content, display a first video of the target video type with the video parameters.

4. The method according to any one of claims 1 to 3, wherein The target video type is a long video type, and the first video of the long video type includes at least two storyboards; After in response to a content input operation performed in the content input area and displaying the input target content, the method further includes: In response to a determination instruction for the target content, display a storyboard disassembly interface, and display at least two storyboard texts determined based on the target content in the storyboard disassembly interface; Wherein, each storyboard text corresponds to a storyboard, and the storyboard text is used to describe the corresponding storyboard; The step of in response to a video generation instruction and displaying a first video of the target video type generated based on the target content includes: In response to a video generation instruction triggered based on the at least two storyboard texts, display a first video of the long video type generated including storyboards corresponding to the respective storyboard texts.

5. The method according to claim 4, wherein, The method further includes: Display a storyboard text quantity adjustment control, where the storyboard text quantity adjustment control is used to increase or delete storyboard texts; In response to a storyboard text addition operation triggered based on the storyboard text quantity adjustment control, display an added new storyboard text in the storyboard disassembly interface; In response to a determination operation triggered based on at least two storyboard texts including the new storyboard text, the video generation instruction is received.

6. The method according to any one of claims 1 to 5, wherein The target video type is a long video type, and the first video of the long video type includes multiple storyboards; After the target content entered is displayed in response to an input operation performed on the content input area, the method further includes: In response to a determination instruction for the target content, a storyboard generation interface is displayed, and at least two storyboard texts and at least two storyboards are displayed in the storyboard generation interface, each storyboard text corresponding to a storyboard, and the storyboard is generated based on the corresponding storyboard text; In response to the video generation instruction, displaying a first video of the target video type generated based on the target content includes: Based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, displaying the generated first video of the long video type including the at least two storyboards.

7. The method according to claim 6, wherein, After the at least two storyboard texts and at least two storyboards are displayed in the storyboard generation interface, the method further includes: In response to a selection operation for a first target storyboard among the at least two storyboards, controlling the first target storyboard to be in a selected state; Based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, displaying the generated first video of the long video type including the at least two storyboards includes: Based on the first target storyboard in the selected state and the storyboard text of the first target storyboard, in response to the video generation instruction, displaying the generated first video of the long video type including the first target storyboard.

8. The method according to claim 6 or 7, wherein The at least two storyboards have an arrangement order, and the first videos formed by the at least two storyboards with different arrangement orders are different, and a sorting adjustment control is also displayed in the storyboard generation interface; After the at least two storyboard texts and at least two storyboards are displayed in the storyboard generation interface, the method further includes: In response to a sorting adjustment operation for the at least two storyboards triggered based on the sorting adjustment control, changing the sorting of the at least two storyboards from the current arrangement order to a target arrangement order; Based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, displaying the generated first video of the long video type including the at least two storyboards includes: Based on the at least two storyboard texts and at least two storyboards, in response to the video generation instruction, displaying the generated first video of the long video type including the at least two storyboards based on the target arrangement order.

9. The method according to any one of claims 6 to 8, wherein, A storyboard update control for updating the content of the storyboard is also displayed in the storyboard generation interface; After the at least two storyboard texts and at least two storyboards are displayed in the storyboard generation interface, the method further includes: Based on the storyboard update control, in response to a content update operation for a second target storyboard among the at least two storyboards, updating the second target storyboard to a first new storyboard, and the content of the first new storyboard is associated with the content of the second target storyboard.

10. The method according to any one of claims 6 to 9, wherein The at least two storyboards have an arrangement order, and there is an upward associated control in the third target storyboard among the at least two storyboards. The upward associated control is used to update the content of the third target storyboard based on the content of the previous storyboard of the third target storyboard. The third target storyboard is any storyboard except the first storyboard among the at least two storyboards; After displaying at least two storyboard texts and at least two storyboards in the storyboard generation interface, the method further includes: In response to a trigger operation on the upward associated control, updating the third target storyboard to a second new storyboard.

11. The method according to claim 10, wherein, Before updating the third target storyboard to a second new storyboard in response to a trigger operation on the upward associated control, the method further includes: If the account that executes the trigger operation generates the first video of the long video type for the first time, display function introduction information; Wherein, the function introduction information is used to introduce the function of the upward associated control.

12. The method according to any one of claims 6 to 11, wherein, The displaying at least two storyboard texts and at least two storyboards in the storyboard generation interface includes: Display at least two storyboard texts in the storyboard generation interface, and display at least two candidate storyboards corresponding to each of the storyboard texts; For each of the storyboard texts, in response to a selection operation on a target candidate storyboard among at least two of the candidate storyboards, control the target candidate storyboard to be in a selected state, and determine the target candidate storyboard in the selected state as the storyboard corresponding to the storyboard text.

13. The method according to any one of claims 1 to 12, wherein, The method further includes: In the video generation interface, display at least one video template; In response to a selection operation on a target video template among the at least one video template, display the template content description of the target video template and a corresponding first editing control; Wherein, the first editing control is used to generate the first video of the target video type based on the template content description; In response to a trigger operation on the first editing control, determine the trigger operation as an input operation for the content input area; The displaying the input target content in response to a content input operation performed in the content input area includes: In response to a trigger operation on the first editing control, in the content input area, display the template content description of the target video template, and determine the displayed template content description as the input target content.

14. The method according to any one of claims 1 to 12, wherein, The method further includes: In the video generation interface, display at least one video template; In response to a trigger operation on a target video template among the at least one video template, display a details interface including the details information of the target video template; Wherein, the details interface further includes at least one of a second editing control and a first new video generation control. The second editing control is used to generate the first video of the target video type based on the template content description. The first new video generation control is used to re-enter target content in the content input area and generate the first video of the target video type based on the re-entered target content; In response to a triggering operation on the second editing control, jump from the details interface to a video generation interface including a content input area carrying a template content description of a target video template; In response to a triggering operation on the first new video generation control, jump from the details interface to a video generation interface including a blank content input area; The response to a content input operation performed in the content input area, displaying the input target content, includes: When a triggering operation on the second editing control is received, determine the triggering operation as the content input operation, and determine the template content description displayed in the content input area as the input target content; When a triggering operation on the first new video generation control is received, in response to a content input operation performed in the blank content input area, display the input target content.

15. The method according to any one of claims 1 to 14, wherein, The response to a video generation instruction, displaying a first video of a target video type generated based on the target content, includes: In response to the video generation instruction, display a first video of a target video type generated based on the target content in a video display interface; Wherein, the video display interface further includes at least one of the following: details information of the generated first video, a third editing control, and a second new video generation control; Wherein, the third editing control is used to regenerate a second video of the target video type based on the target content, and the second new video generation control is used to re-enter target content in the content input area and generate a third video of the target video type based on the re-entered target content.

16. The method according to claim 15, wherein, The video display interface includes the third editing control, the target video type is a long video type, and the first video of the long video type includes at least two storyboards; After the response to the video generation instruction, displaying a first video of a target video type generated based on the target content in a video display interface, the method further includes: In response to a triggering operation on the third editing control, display a storyboard generation interface; In the storyboard generation interface, display at least two storyboard texts determined based on the target content, and display at least two storyboards determined based on the at least two storyboard texts; Based on the storyboard generation interface, in response to an editing operation on the at least two storyboard texts and the at least two storyboards, display the edited at least two storyboard texts and the at least two storyboards; In response to a video generation instruction, display a second video of a long video type generated based on the edited at least two storyboard texts and the at least two storyboards.

17. The method according to claim 15 or 16, wherein, The video display interface includes the third editing control, the target video type is a short video type; after the response to the video generation instruction, displaying a first video of a target video type generated based on the target content in a video display interface, the method further includes: In response to a triggering operation on the third editing control, display a video generation interface including a content input area, and the target content is displayed in the content input area; In response to an editing operation on the target content, display the new target content obtained by editing the target content; In response to a video generation instruction, display a third video of a short video type generated based on the new target content.

18. The method according to any one of claims 1 to 17, wherein The video generation interface further includes a video display control for displaying at least one asset video, where the at least one asset video includes at least one of a successfully generated video that was successfully generated in the past, a failed video that failed to be generated in the past, a video being generated that is in the process of being generated, and a waiting video that is waiting to be generated; After displaying the first video of the target video type generated based on the target content in response to the video generation instruction, the method further includes: In response to a trigger operation on the video display control, display at least one asset video including the first video of the target video type.

19. The method according to claim 18, wherein, The at least one asset video includes a failed video, and the failed video is associated with a regeneration control for regenerating the failed video; After displaying at least one asset video including the first video of the target video type in response to a trigger operation on the video display control, the method further includes: In response to a trigger operation on the regeneration control associated with a target failed video among at least one failed video, regenerate the target failed video; When the regeneration is successful, switch the failed video in the at least one asset video to the regenerated fourth video; When the regeneration fails, display a regeneration failure prompt message.

20. The method according to claim 18 or 19, wherein The at least one asset video includes a waiting video. After displaying at least one asset video including the first video of the target video type in response to a trigger operation on the video display control, the method further includes: In the associated area of the waiting video, display a first waiting prompt message for indicating that the corresponding waiting video is waiting to be generated; In response to a trigger operation on the waiting video triggered based on the first waiting prompt message, display a waiting cancellation control for terminating the waiting process of the waiting video.

21. The method according to any one of claims 1 to 19, wherein, The video generation interface further includes a style setting control for setting the video style of the generated first video; After displaying at least one video generation control in the video generation interface, the method further includes: In response to a trigger operation on the style setting control, display at least one video style; In response to a selection operation on a target video style among the at least one video style, control the target video style to be in a selected state; The displaying the first video of the target video type generated based on the target content in response to the video generation instruction includes: In response to the video generation instruction, display the first video of the target video type with the target video style generated based on the target content.

22. The method according to claim 21, wherein, After controlling the target video style to be in a selected state in response to a selection operation on the target video style among the at least one video style, the method further includes: In response to a determination instruction for the target video style in the selected state, on the style setting control, display an identifier of the target video style, and the identifier is associated with a style cancellation control; Among them, the identifier of the target video style is used to identify that the first video of the target video type generated has the target video style; In response to a trigger operation on the style cancellation control, cancel the display of the identifier of the target video style on the style setting control; The displaying, in response to a video generation instruction, of the first video of the target video type having the target video style based on the target content includes: In response to a video generation instruction, display the first video of the target video type having the default video style based on the target content.

23. The method according to any one of claims 1 to 22, wherein, The displaying, in response to a video generation instruction, of the first video of the target video type based on the target content includes: In response to a video generation instruction, display a second waiting prompt message, and the second waiting prompt message is used to indicate the duration to wait from the start of generating the first video; When the waiting prompt message indicates that the duration to wait from the start of generating the first video is zero, cancel the display of the second waiting prompt message; Display the first video of the target video type generated based on the target content.

24. The method according to any one of claims 1 to 23, wherein After the displaying, in response to a video generation instruction, of the first video of the target video type generated based on the target content, the method further includes: Display at least one interaction control for the first video, and the interaction control is one of the following controls: a local editing control for locally editing the first video, a canvas expansion control for editing the size of the first video, a motion brush control for changing a static object in the first video into a dynamic object, a shooting parameter setting control for setting shooting parameters of the first video, a background removal control for removing the background of the first video, an object erasing control for erasing a target object in the first video, a scene detection control for splitting the video into different video segments based on different scenes in the first video, a depth of field setting control for setting a depth of field effect for the first video, a frame rate adjustment control for adjusting the frame rate of the storyboard in the first video, an action sequence generation control for setting an action sequence for an object in the first video, and a third new video generation control for generating a new video based on the audio data in the first video; In response to a trigger operation on a target interaction control among at least one of the interaction controls, perform the interaction operation indicated by the target interaction control on the first video of the target video type.

25. A video generation device, the device includes: A first display module configured to display at least one video generation control in a video generation interface, and the at least one video generation control is used to generate videos of at least two video types; A second display module, configured to display a content input area corresponding to the target video type based on the at least one video generation control and in response to the target video type among the at least two video types being in a selected state; A third display module, configured to display the input target content in response to a content input operation performed in the content input area; A fourth display module, configured to display a first video of the target video type generated based on the target content in response to a video generation instruction.

26. An electronic device, comprising: A memory, configured to store computer-executable instructions or a computer program; A processor, configured to implement the video generation method according to any one of claims 1 to 24 when executing the computer-executable instructions or the computer program stored in the memory.

27. A computer-readable storage medium, storing computer-executable instructions or a computer program, where when the computer-executable instructions or the computer program are executed by a processor, the video generation method according to any one of claims 1 to 24 is implemented.

28. A computer program product, comprising computer-executable instructions or a computer program, where when the computer-executable instructions or the computer program are executed by a processor, the video generation method according to any one of claims 1 to 24 is implemented.

Citation Information

Patent Citations

  • Video editing processing method and device, electronic equipment and storage medium

    CN113709575A

  • Video content generation method and device, equipment, storage medium and program product

    CN114501105A

  • Video editing method and electronic equipment

    CN115052201A

  • Video processing method and device, electronic equipment and storage medium

    CN116916092A

Cited By

  • Mapping operator-based automatic generation method of sub-mirror rough sketch

    CN120997346A

  • Intelligent split creation method, electronic equipment, storage medium and program product

    CN121665083A