A video generation method, apparatus, device, medium, and program product

By displaying placeholder markers on the video track and responding to user actions, automatically filling in and adjusting the video duration and position, the lack of intelligence and flexibility in video generation methods is solved, achieving more intelligent and flexible video generation.

CN122293906APending Publication Date: 2026-06-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-03-31
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing video generation methods lack intelligence and flexibility, and are cumbersome for users.

Method used

By displaying the first video track in the track editing area and showing the first placeholder on the track, the length and position of the placeholder are associated with the interactive operation. The input information is displayed in response to the input operation, and the video is displayed under the generation operation, so as to realize the automatic filling and adjustment of the video duration and position.

Benefits of technology

The coordination between the track editing area and video generation has been optimized, improving the intelligence and flexibility of video generation. Users can intuitively determine the video duration and insertion position through interactive operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122293906A_ABST
    Figure CN122293906A_ABST
Patent Text Reader

Abstract

This disclosure provides a video generation method, apparatus, device, medium, and program product. The method includes: displaying a track editing area, the track editing area including a first video track; displaying a first placeholder identifier on the first video track, the length of the first placeholder identifier indicating the video duration, the length and position of the first placeholder identifier being associated with interactive operations within the track editing area; displaying first input information in a first area in response to an input operation; and displaying a first video on the first placeholder identifier in response to a generation operation, the first video conforming to the first input information. This method can solve the problem of insufficient intelligence and flexibility in related video generation methods, intuitively determining the duration and insertion position of the generated video through interactive operations within the track editing area, thus improving the intelligence and flexibility of the video generation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to multimedia technology, and more particularly to a video generation method, apparatus, device, medium, and program product. Background Technology

[0002] With the development of computer and multimedia technologies, video editing tools have become increasingly feature-rich, allowing users to create videos. For example, users can use video editing tools to edit multiple video clips into a video that meets their expectations. However, the related video generation methods lack intelligence and flexibility, and the operation is cumbersome for users. Summary of the Invention

[0003] This article provides a video generation method, apparatus, device, medium, and program product that can improve the intelligence and flexibility of video generation methods.

[0004] Firstly, this paper provides a video generation method, including: The video shows a track editing area, which includes a first video track. A first placeholder icon is displayed on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operations within the track editing area. In response to an input operation, the first input information is displayed in the first area; In response to the generation operation, a first video is displayed on the first placeholder identifier, wherein the first video conforms to the first input information.

[0005] Secondly, this paper also provides a video generation apparatus, including: The first display module is used to display the track editing area, wherein the track editing area includes a first video track; The second display module is used to display a first placeholder icon on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operations within the track editing area. The third display module is used to respond to input operations and display the first input information in the first area; The fourth display module is used to display the first video on the first placeholder in response to the triggering of the generation operation, wherein the first video conforms to the first input information.

[0006] Thirdly, this document also provides an electronic device, which includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method.

[0007] Fourthly, this document also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method.

[0008] Fifthly, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the video generation method.

[0009] This paper presents a video generation method. It involves displaying a track editing area, including a first video track; displaying a first placeholder on the first video track, the length of which indicates the video duration; and linking the length and position of the first placeholder to interactive operations within the track editing area. In response to an input operation, first input information is displayed in the first area; and in response to a generation operation, a first video is displayed on the first placeholder, conforming to the first input information. This method automatically fills the first video track with a first placeholder and adjusts its length and / or position through interactive operations within the track editing area, displaying the first video on the first placeholder. This allows control over the duration and / or position of the generated video by editing the duration and / or position of the first placeholder, optimizing the coordination between the track editing area and video generation. It addresses the lack of intelligence and flexibility in existing video generation methods, allowing for intuitive determination of the generated video's duration and insertion position through interactive operations within the track editing area, thus enhancing the intelligence and flexibility of the video generation method. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a diagram illustrating an application scenario of the video generation method presented in this paper. Figure 2 This is a flowchart illustrating a video generation method presented in this paper. Figure 3 This is a schematic diagram of an interactive page provided for this article; Figure 4 This is a flowchart illustrating another video generation method presented in this paper. Figure 5 This is a schematic diagram of yet another type of interactive page provided in this article; Figure 6 This is a schematic diagram of the structure of a video generation device provided in this paper; Figure 7 This is a schematic diagram of the structure of an electronic device provided in this article. Detailed Implementation

[0012] The examples herein will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that this document can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided to provide a more thorough and complete understanding of this document. It should be understood that the accompanying drawings and examples are for illustrative purposes only and are not intended to limit the scope of this document.

[0013] It should be understood that the steps described in the method embodiments herein may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0015] It should be noted that the concepts of "first" and "second" mentioned in this article are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0016] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of messages or information exchanged between multiple devices in the embodiments herein are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0018] It is understandable that before using the technical solutions disclosed in the examples in this article, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this article in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained.

[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.

[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0023] Figure 1 This diagram illustrates an application scenario of the video generation method presented in this paper. It should be noted that... Figure 1 This is merely an example of how this disclosure can be applied in one situation to help those skilled in the art understand the technical content of this disclosure in one situation, and does not mean that this disclosure cannot be used in other devices, systems, environments or scenarios.

[0024] like Figure 1 As shown, the system may include a server 110, a network 120, several terminal devices 130, and a database 140. The terminal devices include at least one of mobile terminals and PCs.

[0025] Server 110 can be a physical server containing a single host, or it can be a virtual server hosted in a host cluster. During operation, server 110 can run server-side programs for video editing applications, thus acting as a server for these applications. Examples of video editing applications include editing software.

[0026] exist Figure 1In the system configuration described, server 110 may include one or more components that implement the functions performed by server 110. These components may include at least one of software components and hardware components executable by one or more processors. A user operating terminal device 130 can interact with server 110 using a client to utilize the services provided by these components. It is understood that different system configurations may exist. Figure 1 This is a schematic diagram illustrating an application scenario for implementing a video generation method in one particular situation.

[0027] Terminal device 130 connects to server 110 via network 120. Typically, network 120 can be any type of network, and it can use any of the various available protocols (including TCP / IP (Transmission Control Protocol / Internet Protocol) and IPX (Internet Packet Exchange)) to support data communication. Network 120 can include various connection types, such as wired communication links, wireless communication links, or fiber optic cables.

[0028] Database 140 can be used to store historically generated videos and video generation information, etc. Database 140 can reside in various locations. For example, database 120 used by server 110 can be local to server 110, or it can be located remotely to server 110 and can communicate with server 110 via network 120 or a dedicated connection. Database 140 can be of different types. In some cases, database 140 used by server 110 can be a relational database. One or more of these databases 140 can store, update, and retrieve database 140 and data from database 140 in response to commands.

[0029] In some cases, one or more of the databases 140 may also be used by applications to store application data. The databases 140 used by applications may be of different types, such as key-value stores, object stores, or regular storage types supported by the file system.

[0030] It should be noted that, Figure 1 The system can be configured and operated in various ways to implement video generation methods and devices in various situations.

[0031] Figure 2This is a flowchart illustrating a video generation method provided in this paper. This paper applies to video creation scenarios, such as video editing. The method can be executed by a video generation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented through an electronic device, such as a mobile terminal, a PC, or a server.

[0032] like Figure 2 As shown, the method includes: S210. Display the track editing area, wherein the track editing area includes a first video track.

[0033] In some cases, the interactive page of a video editing tool displays the primary controls and a timeline editing area. The timeline editing area refers to the portion of the interactive page that contains the timeline track. Optionally, the interactive page is the page corresponding to the video generation tag.

[0034] In other cases, the first control is displayed in the first area of ​​the interactive page. Optionally, the first control is displayed via a parameter panel. The parameter panel is triggered by configuration items displayed in the first area. For example, in response to a click on a configuration item, the parameter panel is displayed on top of the interactive page. The first control is displayed within the parameter panel. Figure 3 This is a schematic diagram of an interactive page provided for this article. (For example...) Figure 3 As shown, the interactive page 300 displays configuration item 310 and track editing area 320. The track editing area includes a first video track 330. In response to a trigger operation on configuration item 310, parameter panel 340 is displayed. A first control 350 is displayed within parameter panel 340.

[0035] The first control is a switch control corresponding to the first placeholder identifier. The first placeholder identifier occupies a portion of a video segment of a first duration at a preset position on the video track, and the duration and position of the first placeholder identifier on the video track are adjustable. The first duration is a preset time. Optionally, the first placeholder identifier can also occupy a portion of a video segment of a second duration at a preset position on the video track, and the second duration is generated based on the duration of the video segments before and / or after the first placeholder identifier. For example, the length of the first placeholder identifier can be generated based on the duration of two video segments connected by the first placeholder identifier. Optionally, the length of the first placeholder identifier can be generated based on the duration of at least the preceding transition video. And / or, the length of the first placeholder identifier can be generated based on the duration of at least the following transition video.

[0036] The operating mode of the first control corresponds to the on / off state of the first placeholder icon. For example, the operating modes of the first control include a first mode and a second mode. The first mode corresponds to the on state of the first placeholder icon. The second mode corresponds to the off state of the first placeholder icon. If the first control is in the first mode, the first placeholder icon is on and displayed within the track. If the first control is in the second mode, the first placeholder icon is off and the first station control is not displayed within the track.

[0037] The track editing area includes at least one timeline track. The timeline track is used to hold video editing footage. The timeline track includes a main track and multiple sub-tracks. The main track is typically a video track. The multiple sub-tracks are typically used to hold at least one of the following: video, audio, effects, text, and stickers. Accordingly, the multiple sub-tracks include at least one of the following types: video track, audio track, effects track, text track, and sticker track.

[0038] S220. Display a first placeholder icon on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operation within the track editing area.

[0039] The video duration refers to the length of the generated video. Interactive operations within the timeline editing area include dragging operations on the first placeholder marker. Drag operations include at least one of the following: stretching, compressing, and adjusting the layer position of the first placeholder marker. The layer position corresponds to the layer of the timeline track where the first placeholder marker is located.

[0040] For example, in response to the first control being in a first mode, the first placeholder identifier is displayed in the first track segment, the first track segment being a blank segment on the first video track starting from the first identifier, wherein the first identifier represents the time information of the track editing area.

[0041] In this document, a first placeholder identifier is displayed on the first track segment, and preset information is displayed on the first placeholder identifier. The first identifier can be a time positioning identifier in the track editing area, used to indicate the time coordinate of the current editing operation. For example, the first identifier can be a time hand; dragging the time hand within the track editing area allows for quick positioning to a specific frame of the video. A blank segment refers to a track segment within the video track that is not occupied by video.

[0042] In some cases, in response to a click operation on the first control, the first control is converted to a first mode, and a first placeholder identifier is displayed on the first track segment corresponding to the first identifier on the first video track.

[0043] In other cases, in response to the operation of dragging the first control to the first video track, the first control is converted to a first mode, and a first placeholder is displayed on the first track segment corresponding to the first identifier on the first video track.

[0044] In this document, the initial position of the first placeholder is determined based on the first identifier, and the position of the first placeholder can be adjusted by dragging the first placeholder. The initial length of the first placeholder is a preset duration, and the length of the first placeholder can be adjusted by dragging the first placeholder within the first video track.

[0045] See Figure 3 As shown, in response to a click operation on the first control 350, a first placeholder identifier 370 is displayed on the first track segment. The first track segment is a blank segment on the first video track 330 corresponding to the first identifier 360.

[0046] In some cases, the relationship between the length and position of the first placeholder marker and interactive operations within the track editing area includes: the length of the first placeholder marker is linked to the stretching operation within the track editing area; the position of the first placeholder marker is linked to the dragging operation within the track editing area.

[0047] Optionally, in response to a drag operation on the first placeholder icon on the first video track, at least one of the duration and position of the first placeholder icon can be adjusted. For example, in response to a stretch operation on the first placeholder icon, the length of the first placeholder icon can be increased. In response to a drag operation on the first placeholder icon, the position of the first placeholder icon can be updated. Therefore, by dragging the first placeholder icon within the video track, the duration and insertion position of the generated video can be controlled, achieving a deep integration of interaction within the timeline track and video generation.

[0048] See Figure 3 As shown, a first placeholder identifier 370 is displayed between two video clips. In response to a stretching operation on the first placeholder identifier 370 (the arrow indicates the dragging direction), the length of the first placeholder identifier 370 is increased.

[0049] Optionally, in response to a drag operation on the first placeholder, the position of the first placeholder is adjusted on the first video track.

[0050] Optionally, in response to a dragging operation of the first placeholder icon outside the first video track, the first placeholder icon is displayed on a second video track. For example, the first placeholder icon is dragged from the first video track to a blank area within the track editing area, creating a second video track in the blank area, and the first placeholder icon is displayed within the second video track. Alternatively, the second video track is displayed in the track editing area. Dragging the first placeholder icon from the first video track to the second video track displays the first placeholder icon on the second video track.

[0051] It should be noted that if the duration of the first track segment is greater than or equal to the initial length of the first placeholder identifier, the first placeholder identifier will be displayed within the first track segment. If the duration of the first track segment is less than the initial length of the first placeholder identifier, the first placeholder identifier cannot be displayed through the first track segment. Alternatively, if the first identifier is located within a video segment within the first video track, there will be no blank segment starting from the first identifier within the first video track. Consequently, the duration of the first track segment will be zero, and the first placeholder identifier cannot be displayed through the first track segment.

[0052] In some cases, in response to the duration of the first track segment being less than the length of the first placeholder, the first placeholder is displayed in a second track segment, which is a blank segment on the second video track that starts from the first placeholder.

[0053] The second video track refers to a newly created video track or a video track within the track editing area that meets the display requirements of the first placeholder identifier. If the duration of a blank segment within a video track, starting from the first identifier, is greater than or equal to the length of the first placeholder identifier, then the video track is determined to be a second video track that meets the display requirements of the first placeholder identifier. Optionally, if multiple video tracks within the track editing area meet the display requirements of the first placeholder identifier, a second video track is determined from these multiple video tracks based on the track hierarchy.

[0054] Optionally, a second control is displayed, which is used to configure duration information, wherein the duration information is linked to the length of the first placeholder identifier.

[0055] The second control is used to set the video duration. Optionally, the second control is displayed in the parameter panel. See Figure 3, where the second control 390 is displayed in parameter panel 340.

[0056] Optionally, the second control may include a duration input box or a slider control, etc. Adjusting the video duration corresponding to the second control can synchronously adjust the length of the first placeholder indicator. Correspondingly, adjusting the length of the first placeholder indicator on the video track can synchronously update the video duration corresponding to the second control. Therefore, it can be seen that the parameter value of the second control and the length of the first placeholder indicator on the video track are linked. Here, the parameter value of the second control refers to the duration information.

[0057] S230, In response to an input operation, display the first input information in the first area.

[0058] The first area is a partial area within the interactive page. The interactive page includes a first tab and a second tab. When the first tab is selected, it displays task information within the first area. This task information represents historical videos and their generation information. Historical videos represent successfully generated videos. Generation information instructs the generative model to generate historical videos. Optionally, generation information includes at least one of the following: generation mode, prompt text, reference materials, model type, aspect ratio, resolution, duration, and camera movement. The generation mode represents the media content generation scenario. For example, generation modes include at least one of the following: image-to-video, first and last frame-to-video, video extension, video-to-video, and audio-to-video.

[0059] The second tab indicates the selected state and displays interaction information with the generative model in the first area. The generative model is used to perform a specified type of task based on user requests. For example, generative models include language models or multimodal models. Interaction information includes prompts, generation modes, and configuration parameters.

[0060] For example, in response to the second tab being selected, input controls, mode options, configuration items, and generation controls are displayed in the first area. See also Figure 3 As shown, the first area 3100 displays input control 3110, configuration item 310, and generation control 380, etc.

[0061] The input controls are used to input at least one of text, reference video, reference image, and reference audio. Optionally, the input controls include text controls and reference material controls, etc. Reference materials include at least one of reference video, reference image, and reference audio. The mode options include multiple generation modes, which represent the generation scenario of the second media content. Configuration items include model type, aspect ratio, resolution, and duration, etc.

[0062] In this document, the input operation is used to input at least one of the following: prompt text, reference materials, generation mode, and configuration parameters. Accordingly, the first input information includes at least one of the following: prompt text, reference materials, generation mode, and configuration parameters.

[0063] In some cases, in response to input within the first area, prompt text and reference materials are displayed on the input controls. The reference video is uploaded from the media library or selected from the task information. The generation mode is displayed in the mode options. Configuration parameters are displayed in the parameter panel corresponding to the configuration item.

[0064] In other scenarios, in response to a modification trigger operation on any historical video in the task information, the generation information of the triggered historical video is displayed in the first area, and the triggered historical video is also displayed in the timeline track. Based on this, the generation information of the historical video can be edited. Video generation is then performed based on the edited generation information.

[0065] S240. In response to the triggering of the generation operation, display the first video on the first placeholder identifier, wherein the first video conforms to the first input information.

[0066] The first video is generated by a generative model based on the first input information. For example, the first video can be used as a transition video. Optionally, if the first video is a transition video, two video segments adjacent to the first video can be used as reference videos for video generation.

[0067] For example, in response to a trigger operation on the generation control, a generative model generates a first video based on the first input information and the duration corresponding to the first placeholder, and displays the first video on the first placeholder. Optionally, if the duration of the first placeholder is not a positive integer, the generative model generates a video based on the positive integer corresponding to the duration of the first placeholder, and then obtains the first video by cropping or extracting frames from the generated video.

[0068] In some cases, in response to a trigger operation of the generation control, the first video is displayed on the first placeholder, and task information is displayed in the first area, wherein the task information includes the generation state of the first video. Specifically, when the second tab is selected, input controls, mode options, configuration items, and generation controls are displayed in the first area. A first placeholder is displayed on the first video track. In response to a click operation on the generation control, the first tab is switched to a selected state, and task information is displayed in the first area. The task information includes historical videos and the generation state of the first video. Optionally, the generation state of the first video is presented as a video cover with image blurring. If the first video is successfully generated, the generation state of the first video is changed to a generation success state. The generation success state of the first video is presented as the video cover of the first video. If the first video fails to generate, or if the first placeholder is deleted from the first video track during the generation process, the generation state of the first video is changed to a generation failure state. The generation failure state of the first video is presented as a failure indicator.

[0069] Optionally, the first control is in a second mode, and further includes: In response to an interaction with the second control, the duration information is updated, wherein the duration information indicates the length of the second placeholder identifier. In response to a triggering operation on the generate control, the second placeholder identifier is displayed on the first video track, and the first video is displayed on the second placeholder identifier.

[0070] Since the first control is in the second mode, meaning the first placeholder identifier is off, it will not be displayed on the first video track. In response to the interaction with the second control, the duration information corresponding to the second control is updated. The duration information corresponding to the second control indicates the length of the second placeholder. The second placeholder occupies a segment of the length corresponding to the duration information at a preset position on the video track. During the generation of the first video, preset information is displayed on the second placeholder. The video duration of the first video corresponds to the length of the second placeholder.

[0071] In response to a trigger operation on the generated control, a second placeholder is displayed on the first video track, and the length of the second placeholder is consistent with the duration information corresponding to the second control. Optionally, the length and position of the second placeholder can be adjusted, and the specific adjustment method is the same as that of the first placeholder, which will not be repeated here. A generative model is used to generate a first video based on the first input information and the duration corresponding to the second placeholder, and the first video is displayed on the second placeholder.

[0072] Optionally, in response to media content being dragged from the track editing area to the first position, the media content is displayed as reference material in the first area. The media content in the track editing area includes at least one of video, image, and audio. The first position is a pre-defined position. For example, the first position is the edge area of ​​the track editing area. In response to dragging a video clip from the video track to the edge area of ​​the track editing area, the video clip is displayed as a reference video in the first area. Optionally, a thumbnail of a video cover is displayed in the first area. Alternatively, in response to dragging an image from the sticker track to the edge area of ​​the track editing area, a sticker is displayed as a reference image in the first area. Optionally, a thumbnail of a sticker is displayed in the first area. Alternatively, in response to dragging an audio clip from the audio track to the edge area of ​​the track editing area, the audio clip is displayed as reference audio in the first area. Optionally, a thumbnail of an audio cover is displayed in the first area.

[0073] The technical solution presented in this paper utilizes a track editing area, which includes a first video track. A first placeholder is displayed on the first video track, its length indicating the video duration. The length and position of the first placeholder are related to interactive operations within the track editing area. In response to an input operation, first input information is displayed in the first area. In response to a generation operation, a first video is displayed on the first placeholder, conforming to the first input information. This solution automatically fills the first video track with a first placeholder and adjusts its length and / or position through interactive operations within the track editing area, displaying the first video on the first placeholder. This allows control over the duration and / or position of the generated video by editing the duration and / or position of the first placeholder, optimizing the coordination between the track editing area and video generation. It addresses the lack of intelligence and flexibility in existing video generation methods, enabling intuitive determination of the generated video's duration and insertion position through interactive operations within the track editing area, thus enhancing the intelligence and flexibility of the video generation method.

[0074] Figure 4 This is a flowchart illustrating another video generation method presented in this paper, which, based on the examples above, adds a limitation on the intelligent video extension method. For example... Figure 4 As shown, the method includes: S410. In response to an interactive operation on the second video, a third control and a third placeholder are displayed, wherein the interactive operation is used to extend the second video, and the length of the third placeholder is used to represent the extension duration.

[0075] The second video can include videos uploaded from a media library or videos generated through a generative model. Optionally, the second video can include a first video generated based on the above example. Interactive operations for the second video include dragging the first or last frame of the second video on the timeline. The third control is used to trigger the execution of a video extension task. The video extension task refers to the task of extending the duration of the second video. The third placeholder is used to display the extended video. The extended video refers to the portion of the video sequence obtained by extending the second video. During the generation of the extended video, preset information is displayed on the third placeholder. For example, the preset information is used to characterize the state during video generation.

[0076] For example, in response to a selection operation on the second video, a fourth control is displayed in the track editing area; in response to a trigger operation on the fourth control, a third identifier is displayed, wherein the third identifier is used to trigger a drag operation on the second video; in response to the triggering of the drag operation, a third control and a third placeholder identifier are displayed.

[0077] The fourth control represents the entry point for the intelligent video extension function. Optionally, the fourth control is displayed in the track editing area. Alternatively, it is displayed when the second video on the video track is selected. The third identifier includes a drag icon. By pressing and dragging the drag icon, a drag operation is performed on the second video, displaying the third control and the third placeholder icon.

[0078] In some cases, displaying the third control and the third placeholder in response to the triggering of the drag operation includes: displaying the third control and the third placeholder starting from the last frame of the second video in response to dragging the last frame of the second video in a direction away from the second video.

[0079] In other cases, displaying the third control and the third placeholder in response to the drag operation includes: displaying the third control and the third placeholder with the first frame as the endpoint in response to dragging the first frame of the second video in a direction away from the second video.

[0080] Figure 5 This is a schematic diagram of another type of interactive page provided in this article. For example... Figure 5As shown, the interactive page 500 displays a track editing area 510 and a first region 570. Task information is displayed within the first region 570. The track editing area 510 includes a timeline track 520 and a fourth control 530. A second video 540 is displayed on the timeline track 520. In response to dragging the last frame of the second video 540 away from the second video 540, a third control 550 (represented by a bold solid line) and a third placeholder marker 560 (represented by a bold solid line) are displayed. Alternatively, in response to dragging the first frame of the second video 540 away from the second video 540, a third control 550 (represented by a bold dashed line) and a third placeholder marker 560 (represented by a bold dashed line) are displayed.

[0081] Optionally, a third control, a third placeholder, and a drag duration are displayed at the top of the timeline track where the second video is located. The length of the third placeholder is the same as the drag duration. For example, during dragging the first or last frame of the second video, the length of the third placeholder is extended, and the drag duration is updated synchronously. Assuming the second video is dragged on the track for t seconds, the drag duration is t seconds.

[0082] S420. In response to a trigger operation on the third control, a third video is displayed on the third placeholder, wherein the third video is a video clip generated based on the second video.

[0083] In some cases, the third video is generated using the last frame as the first frame of the video. The first frame of the video refers to the first frame of the extended video. For example, in response to a trigger operation on the third control, a generative model is used to generate a third video by using the last frame of the second video as the first frame of the extended video. The preset information displayed on the third placeholder is then switched to the third video.

[0084] In other cases, the third video is generated using the first frame as the last frame of the video. Here, the last frame refers to the last frame of the extended video. For example, in response to a trigger operation on the third control, a generative model is used to generate a third video using the first frame of the second video as the last frame of the extended video. The preset information displayed on the third placeholder is then switched to the third video.

[0085] Optionally, the method further includes: displaying a fifth control; in response to a trigger operation on the fifth control, displaying the second video and a duration editing item in the first area, wherein the duration editing item is used to control the duration; in response to an input operation, displaying second input information in the first area, wherein the second input information includes modification description information for the second video; and in response to a generation operation, displaying the third video on the third placeholder, wherein the third video conforms to the second input information. The second video and draggable duration can be refilled into the first area via the fifth control, and the third video can be generated by combining the modification description information for the second video in the first area, thus establishing a link between editing operations on the timeline track and video generation.

[0086] The duration editing section displays time information, which is the extended duration of the video. The extended video duration can be modified using the duration control in the configuration settings. The second video serves as a reference for the third video. In response to any position adjustment operation between the second video and the duration editing section, their positions are swapped.

[0087] Optionally, a swap control is included between the second video and the duration editing item. In response to the swap control being triggered, the positions of the second video and the duration editing item are swapped, updating the position of the third placeholder on the timeline track. For example, if the third placeholder on the timeline track is after the second video, in response to the swap control being triggered, the third placeholder is displayed before the second video. Alternatively, if the third placeholder on the timeline track is before the second video, in response to the interactive control being triggered, the third placeholder is displayed after the second video.

[0088] In some cases, the second input information includes at least one of duration parameters, input text, and model type. The input text describes the video content to be extended. The duration parameter controls the extension duration. In response to the generation operation, based on the first positional relationship between the second video and the duration editing item, the second video is extended backward by the duration corresponding to the duration editing item. This backward extension uses the last frame of the second video as the first frame of the extended video, combined with the second input information, to generate the video. Alternatively, in response to the generation operation, based on the second positional relationship between the second video and the duration editing item, the second video is extended forward by the duration corresponding to the duration editing item. This forward extension uses the first frame of the second video as the last frame of the extended video, combined with the second input information, to generate the video.

[0089] See Figure 5In response to a drag operation on the second video 540, a fifth control 580 is displayed. In response to a trigger on the fifth control 580, an input control 590, a duration editing item 5100, a swapping control 5110, a configuration item 5120, and a generation control 5130 are displayed in the first area 570. The video cover of the second video is displayed on the input control 590. The dragged duration is displayed on the duration editing item 5100. In response to a click operation on the generation control 5130, a third video is displayed on the third placeholder marker 560.

[0090] The technical solution presented in this paper displays a third control and a third placeholder in response to interactive operations on the second video. The interactive operation extends the second video, and the length of the third placeholder represents the extension duration. In response to a trigger operation on the third control, a third video is displayed on the third placeholder; this third video is a video segment generated based on the second video. The video extension duration can be intuitively determined based on dragging operations on the third placeholder, and the third video can be generated based on the second video and the extended duration. The video extension operation can be directly performed on the timeline track, achieving direct coupling between video extension and interactive operations on the timeline track.

[0091] Figure 6 This is a schematic diagram of the structure of a video generation device provided in this paper. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0092] like Figure 6 As shown, the device includes: a first display module 610, a second display module 620, a third display module 630, and a fourth display module 640.

[0093] The first display module 610 is used for the track editing area, wherein the track editing area includes a first video track; The second display module 620 is used to display a first placeholder icon on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operations within the track editing area. The third display module 630 is used to respond to input operations and display the first input information in the first area; The fourth display module 640 is used to display a first video on the first placeholder in response to the triggering of the generation operation, wherein the first video conforms to the first input information.

[0094] Optionally, the second display module 620 is specifically used for: In response to the first control being in the first mode, the first placeholder identifier is displayed in the first track segment. The first track segment is a blank segment on the first video track that starts from the first identifier, wherein the first identifier represents the time information of the track editing area.

[0095] Optionally, it also includes: In response to the fact that the duration of the first track segment is less than the length of the first placeholder, the first placeholder is displayed in the second track segment, which is a blank segment on the second video track that starts from the first placeholder.

[0096] Optionally, it also includes: In response to a drag operation of the first placeholder icon on the first video track, adjust at least one of the duration and position of the first placeholder icon; In response to a drag operation of the first placeholder icon outside the first video track, the first placeholder icon is displayed on the second video track.

[0097] Optionally, it also includes: A second control is displayed, which is used to configure duration information, wherein the duration information is linked to the length of the first placeholder identifier.

[0098] Optionally, it also includes: In response to an interactive operation on the second control, the duration information is updated, wherein the duration information is used to indicate the length of the second placeholder identifier; In response to a trigger operation on the generated control, the second placeholder is displayed on the first video track, and the first video is displayed on the second placeholder.

[0099] Optionally, the fourth display module 640 is specifically used for: In response to the triggering operation of the generation control, the first video is displayed on the first placeholder identifier, and task information is displayed in the first area, wherein the task information includes the generation state of the first video.

[0100] Optionally, it also includes: In response to an interactive operation on the second video, a third control and a third placeholder are displayed, wherein the interactive operation is used to extend the second video, and the length of the third placeholder is used to represent the extension duration; In response to a trigger operation on the third control, a third video is displayed on the third placeholder, wherein the third video is a video clip generated based on the second video.

[0101] Optionally, the display of a third control and a third placeholder in response to an interactive operation on the second video includes: In response to a selection operation on the second video, a fourth control is displayed in the track editing area; In response to a trigger operation on the fourth control, a third identifier is displayed, wherein the third identifier is used to trigger a drag operation on the second video; In response to the drag operation, a third control and a third placeholder are displayed.

[0102] Optionally, the display of a third control and a third placeholder in response to the drag operation includes: In response to dragging the last frame of the second video away from the second video, the third control and a third placeholder starting from the last frame are displayed.

[0103] Optionally, the third video is generated using the last frame as the first frame of the video.

[0104] Optionally, the display of a third control and a third placeholder in response to the drag operation includes: In response to dragging the first frame of the second video away from the second video, the third control and a third placeholder marker ending at the first frame are displayed.

[0105] Optionally, the third video is generated using the first frame as the last frame of the video.

[0106] Optionally, it also includes: Display the fifth control; In response to a trigger operation on the fifth control, the second video and a duration editing item are displayed in the first area, wherein the duration editing item is used to control the duration; In response to an input operation, second input information is displayed in the first area, wherein the second input information includes modified description information for the second video; In response to the generation operation, the third video is displayed on the third placeholder, wherein the third video conforms to the second input information.

[0107] Optionally, it also includes: In response to media content being dragged to the first position within the track editing area, the media content is displayed as reference material in the first area.

[0108] The video generation apparatus provided in this paper can execute the video generation method provided in any example in this paper, and has the corresponding functional modules and beneficial effects of the execution method.

[0109] It is worth noting that the various units and modules included in the above-mentioned device are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this document.

[0110] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this article. See below for reference. Figure 7 It shows an electronic device suitable for implementing this paper (e.g. Figure 7 The diagram below shows the structure of the terminal device or server 700. The terminal device in this document may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not impose any limitations on the functionality and scope of this article.

[0111] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.

[0112] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0113] In particular, according to the examples herein, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the examples herein include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an example, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods herein.

[0114] The names of messages or information exchanged between multiple devices in the embodiments herein are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0115] The electronic device provided herein and the video generation method provided in the above examples belong to the same inventive concept. Technical details not described in detail in this example can be found in the above examples, and this example has the same beneficial effects as the above examples.

[0116] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the video generation method provided in the above example.

[0117] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0118] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0119] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0120] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: The video shows a track editing area, which includes a first video track. A first placeholder icon is displayed on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operations within the track editing area. In response to an input operation, the first input information is displayed in the first area; In response to the generation operation, a first video is displayed on the first placeholder identifier, wherein the first video conforms to the first input information.

[0121] Computer program code for performing the operations described herein may be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0122] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to the various examples herein. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0123] The units described herein can be implemented in software or hardware. The names of the units are not, in some cases, limiting to the unit itself.

[0124] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0125] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0126] The above description is merely a preferred example and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.

[0127] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain contexts. Similarly, while some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this paper. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.

[0128] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video generation method, comprising: The video shows a track editing area, which includes a first video track. A first placeholder icon is displayed on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operations within the track editing area. In response to an input operation, the first input information is displayed in the first area; In response to the generation operation, a first video is displayed on the first placeholder, wherein the first video conforms to the first input information.

2. The method according to claim 1, wherein displaying the first placeholder identifier on the first video track includes: In response to the first control being in the first mode, the first placeholder identifier is displayed in the first track segment. The first track segment is a blank segment on the first video track that starts from the first identifier, wherein the first identifier represents the time information of the track editing area.

3. The method according to claim 2, further comprising: In response to the fact that the duration of the first track segment is less than the length of the first placeholder, the first placeholder is displayed in the second track segment, which is a blank segment on the second video track that starts from the first placeholder.

4. The method according to claim 1, further comprising: In response to a drag operation of the first placeholder icon on the first video track, adjust at least one of the duration and position of the first placeholder icon; In response to a drag operation of the first placeholder icon outside the first video track, the first placeholder icon is displayed on the second video track.

5. The method according to claim 1, further comprising: A second control is displayed, which is used to configure duration information, wherein the duration information is linked to the length of the first placeholder identifier.

6. The method according to claim 5, further comprising: In response to an interactive operation on the second control, the duration information is updated, wherein the duration information is used to indicate the length of the second placeholder identifier; In response to a trigger operation on the generated control, the second placeholder is displayed on the first video track, and the first video is displayed on the second placeholder.

7. The method according to claim 1, wherein displaying the first video on the first placeholder in response to the triggering of the generation operation comprises: In response to the triggering operation of the generation control, the first video is displayed on the first placeholder identifier, and task information is displayed in the first area, wherein the task information includes the generation state of the first video.

8. The method according to claim 1, further comprising: In response to an interactive operation on the second video, a third control and a third placeholder are displayed, wherein the interactive operation is used to extend the second video, and the length of the third placeholder is used to represent the extension duration; In response to a trigger operation on the third control, a third video is displayed on the third placeholder, wherein the third video is a video clip generated based on the second video.

9. The method of claim 8, wherein displaying the third control and the third placeholder in response to an interactive operation on the second video comprises: In response to a selection operation on the second video, a fourth control is displayed in the track editing area; In response to a trigger operation on the fourth control, a third identifier is displayed, wherein the third identifier is used to trigger a drag operation on the second video; In response to the drag operation, a third control and a third placeholder are displayed.

10. The method of claim 9, wherein displaying the third control and the third placeholder in response to the triggering of the drag operation includes: In response to dragging the last frame of the second video away from the second video, the third control and a third placeholder starting from the last frame are displayed.

11. The method according to claim 10, wherein the third video is generated using the last frame as the first frame of the video.

12. The method of claim 9, wherein displaying the third control and the third placeholder in response to the triggering of the drag operation includes: In response to dragging the first frame of the second video away from the second video, the third control and a third placeholder marker ending at the first frame are displayed.

13. The method according to claim 12, wherein the third video is generated using the first frame as the video tail frame.

14. The method of claim 8, further comprising: Display the fifth control; In response to a trigger operation on the fifth control, the second video and a duration editing item are displayed in the first area, wherein the duration editing item is used to control the duration; In response to an input operation, second input information is displayed in the first area, wherein the second input information includes modified description information for the second video; In response to the generation operation, the third video is displayed on the third placeholder, wherein the third video conforms to the second input information.

15. The method according to claim 1, further comprising: In response to media content being dragged to the first position within the track editing area, the media content is displayed as reference material in the first area.

16. A video generation apparatus, comprising: The first display module is used to display the track editing area, wherein the track editing area includes a first video track; The second display module is used to display a first placeholder icon on the first video track. The length of the first placeholder icon is used to indicate the video duration. The length and position of the first placeholder icon are related to the interactive operations within the track editing area. The third display module is used to respond to input operations and display the first input information in the first area; The fourth display module is used to display the first video on the first placeholder in response to the triggering of the generation operation, wherein the first video conforms to the first input information.

17. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any one of claims 1-15.

18. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in any one of claims 1-15.

19. A computer program product comprising a computer program that, when executed by a processor, implements the video generation method as described in any one of claims 1-15.