Video generation method and device, equipment, medium and program product
By displaying a first track and a second display area during the video generation process, users can intuitively edit media materials and input descriptive text for screen changes, solving the problem of video effects not meeting expectations and improving the quality of video generation.
Patent Information
- Application Number
- CN202511575835.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-01-09
AI Technical Summary
In existing technologies, video generation methods suffer from issues such as video effects not meeting expectations and poor user experience.
By displaying a first track and a second display area, users can intuitively edit the first media material in the first display area and enter descriptive text in the second display area to represent the changes in the image, thereby generating a target video that meets the user's expectations.
It enables flexible and accurate generation of target videos based on multi-frame media footage and descriptive text, thereby improving the quality of video generation.
Smart Images

Figure CN121309934A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to multimedia technology, and more particularly to a video generation method, apparatus, device, medium, and program product. Background Technology
[0002] With the rapid development of multimedia technology and the increasing demand for content creation, video generation has become an important method of digital content production. Currently, users can generate videos by entering prompts or selecting preset templates, but this may result in video effects that do not meet expectations, leading to a poor user experience. Summary of the Invention
[0003] This disclosure provides a video generation method, apparatus, device, medium, and program product that can improve video quality and generate target videos that meet expectations.
[0004] In a first aspect, embodiments of this disclosure provide a video generation method, including:
[0005] The first track is displayed on the first page. The first track includes a first display area and a second display area. The first display area is used to display the first media material, and the second display area includes preset placeholder content for video clips.
[0006] In response to a trigger operation on the second display area, an editing area and a first control are displayed on the first page, wherein the editing area is used to input first descriptive text, and the first descriptive text represents at least part of the screen change information corresponding to the first media material;
[0007] In response to a trigger operation on the first control, a second page is displayed, on which a target video is displayed. The target video includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
[0008] Secondly, embodiments of this disclosure also provide a video generation apparatus, the apparatus comprising:
[0009] The first display module is used to display a first track on the first page. The first track includes a first display area and a second display area. The first display area is used to display first media material, and the second display area includes preset placeholder content for video clips.
[0010] The second display module is used to display an editing area and a first control on the first page in response to a trigger operation on the second display area. The editing area is used to input the first descriptive text, which represents at least part of the screen change information corresponding to the first media material.
[0011] The third display module is used to display a second page in response to a trigger operation on the first control. The second page displays a target video, which includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
[0012] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0013] One or more processors;
[0014] Storage device for storing one or more programs.
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any embodiment of this disclosure.
[0016] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in any embodiment of this disclosure.
[0017] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video generation method as described in any embodiment of this disclosure.
[0018] This disclosure provides a video generation method. By displaying a first track, including a first display area and a second display area, and the first display area including first media material, the first media material can be intuitively displayed, allowing for flexible editing based on creative needs. Furthermore, preset placeholder content for video clips is displayed in the second display area, enabling precise control over the position of the video clips to be generated. Further, in response to a trigger operation on the second display area, an editing area and a first control are displayed on a first page. The editing area is used to input first descriptive text, which represents at least a portion of the screen change information corresponding to the first media material. Since the video clips are generated based on the first media material and the first descriptive text, and both the first media material and the first descriptive text can be flexibly edited, it is easy to generate a target video that meets user expectations. In response to a trigger operation on the first control, a second page is displayed, showing the target video. The technical solution of this disclosure provides an optimized video generation method, enabling flexible and accurate generation of a target video based on multiple frames of first media material and first descriptive text. This solves the problem of unsatisfactory video effects in related technologies, improving the quality of generated videos. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 This is a flowchart illustrating a video generation method provided in an embodiment of the present disclosure;
[0021] Figure 2 This is a schematic diagram of a first page provided for an embodiment of the present disclosure;
[0022] Figure 3 This is a schematic flowchart of another video generation method provided in an embodiment of this disclosure;
[0023] Figure 4 This is a schematic diagram of another first page provided for an embodiment of this disclosure;
[0024] Figure 5 A flowchart illustrating yet another video generation method provided in this disclosure embodiment;
[0025] Figure 6 A schematic diagram of a third page provided for an embodiment of this disclosure;
[0026] Figure 7 A schematic diagram of a fourth page provided for an embodiment of this disclosure;
[0027] Figure 8 This is a schematic diagram of the structure of a video generation apparatus provided in an embodiment of the present disclosure;
[0028] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0034] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0035] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0036] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0037] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0038] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0039] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0040] Figure 1 This is a flowchart illustrating a video generation method provided in an embodiment of the present disclosure. This embodiment is applicable to video editing scenarios, such as adding a target video between at least two video segments. The method can be executed by a video generation device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.
[0041] like Figure 1 As shown, the method includes:
[0042] S110. Display a first track on the first page. The first track includes a first display area and a second display area. The first display area is used to display first media material, and the second display area includes preset placeholder content for video clips.
[0043] The first page represents the preview page of the first media material in the editing application. The editing application represents the application used to edit the media material. The editing application includes a preset page. For example, the preset page can be an editing page. The first page is a subpage of the editing page. The first page is triggered by interactive operations in the editing page. The first media material is an image or video used to generate the target video. The first media material can come from a preset media resource library. For example, the preset media resource library includes a local photo album or a cloud photo album. Optionally, the first media material can also be an image or video generated by a generation model. The first track represents the digital editing area of the first media material. In some embodiments, the first media material in the first track is arranged in chronological order. Optionally, a thumbnail of the first media material is displayed in the first track. A preview area is displayed above the first track on the first page, and a preview image of the first media material in the selected state is displayed in the preview area. A second display area is displayed between the thumbnails of two first media materials. The second display area displays preset placeholder content for video clips corresponding to two adjacent frames of first media materials. Optionally, the preset placeholder content for video clips includes attribute information such as the duration of the video clip.
[0044] Optionally, in response to a sliding operation within the first track, the first media material displayed on the first page is switched, allowing multiple first media materials to be displayed within the first track through sliding switching. In response to a trigger operation on the second display area, information on screen changes between the first media materials is obtained, and / or, information on screen changes for each video segment is obtained. For example, the first descriptive text can describe the screen change information between two adjacent first media materials. Optionally, the first descriptive text can also describe the screen change information for all video segments. For example, the same special effects can be added to all video segments through the first descriptive text.
[0045] Figure 2 This is a schematic diagram of a first page provided as an embodiment of this disclosure. (As shown...) Figure 2 As shown, the first page 210 displays a first track 220 and a preview area 230. The first track 220 includes a first display area 240 and a second display area 250. A thumbnail of the first media clip is displayed in the first display area 240. The attribute information of the video clip is displayed in the second display area 250. Optionally, the attribute information may include video duration, etc. A preview image of the selected first media clip is displayed in the preview area 230.
[0046] Optionally, in response to a trigger operation on a special effects control in the preset page, a special effects page is displayed, wherein the special effects page includes multiple candidate special effects templates. In response to a trigger operation on a target special effects template, a third page is displayed, wherein the third page includes a third control, the third control being used to trigger the acquisition of the first media material.
[0047] For example, in response to a trigger action on an effect control in the editing page, candidate controls corresponding to the effect control are displayed. In response to a trigger action on the target control among the candidate controls corresponding to the effect control, an effects page is displayed. The effects page includes multiple candidate effect templates. In response to a trigger action on the target effect template among the multiple candidate effect templates, a third page is displayed. The third page displays a reference frame for the third control and the target video. The reference frame for the target video can be the selected video frame of the second media material in the editing page. Alternatively, the reference frame for the target video can be the last video frame of the second media material in the editing page, etc. The third control is used to trigger the addition of the first media material after the reference frame.
[0048] S120. In response to a trigger operation on the second display area, an editing area and a first control are displayed on the first page, wherein the editing area is used to input first descriptive text, and the first descriptive text represents at least part of the screen change information corresponding to the first media material.
[0049] In this embodiment of the disclosure, the editing area is a control on the first page for inputting first descriptive text. The first descriptive text is descriptive text about screen change information corresponding to at least a portion of the first media materials, enabling user-customized video effects. Optionally, the first descriptive text may be descriptive text about global screen change information of the target video. And / or, the first descriptive text may be descriptive text about screen change information between portions of the first media materials. For example, the first descriptive text is descriptive text about screen change information of each video segment of the target video. Wherein, each video segment of the target video includes the video segment corresponding to the first media materials in frame 1 and frame 2, the video segment corresponding to the first media materials in frame 2 and frame 3, ..., and the video segment corresponding to the first media materials in frame (i-1) and frame 1. Alternatively, the first descriptive text is descriptive text about screen change information of video segments corresponding to first media materials in two adjacent frames. For example, the first descriptive text is descriptive text about screen change information between the first media materials in frame 1 and frame 2. Or, the first descriptive text is descriptive text about screen change information between the first media materials in frame 2 and frame 3. Alternatively, the first descriptive text may be a descriptive text about the changes in image quality between the first media material in frame i-1 and the first media material in frame i.
[0050] Optionally, in response to a text input operation within the editing area, a first descriptive text is displayed in the editing area. Alternatively, in response to a voice input operation within the editing area, voice data is converted into text content, a first descriptive text is determined based on the text content, and the first descriptive text is displayed in the editing area.
[0051] The first control is used to trigger the video generation operation. For example, the first control can be used to trigger the generation of the target video.
[0052] For example, in response to a trigger operation targeting the second display area, a first track, an editing area, and a first control are displayed on the first page, wherein the triggered second display area is in a selected state. For example, the selected state is to overlay a preset display effect onto the second display area. The preset display effect includes thickening the border, changing the border color, or making the border glow, etc. The editing area is in a state awaiting input. For example, the editing area in the state awaiting input includes preset prompt text, which is used to indicate the effective range of the screen change information described by the first descriptive text. Here, the effective range refers to the range of the second display area corresponding to the screen change information described by the first descriptive text.
[0053] In some embodiments, in response to an input operation on the editing area, first descriptive text is displayed in the editing area. In response to a switching operation on the second display area, the selected state of the second display area after switching is displayed on the first page, as well as the input-pending state of the editing area, allowing the input of the first descriptive text corresponding to the selected second display area to be entered in the editing area.
[0054] S130. In response to a trigger operation on the first control, a second page is displayed, on which a target video is displayed. The target video includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
[0055] The second page is a preview page for the target video. It includes a preview area, a regeneration control, and an application control. The preview area is used to play the target video. The regeneration control triggers the regeneration of the target video. The application control adds the target video to the second track of the preset page. Optionally, the target video is added to a position to be filled in the second track. This position could be the position corresponding to a selected video frame in the second track, or it could be a position after a segment of video footage on the preset timeline. Optionally, the reference frame for the target video is the selected video frame in the second track, or it could be the last frame of a segment of video footage.
[0056] For example, in response to a trigger operation on the first control, a video segment corresponding to the selected second display area is generated based on the first descriptive text corresponding to the selected second display area and the first media material displayed in the first display area adjacent to the selected second display area. Here, the first display area adjacent to the selected second display area refers to the first display area before and after the selected second display area. For example, first display area A, second display area B, and first display area C are three consecutive areas within the first track. If second display area B is the selected second display area, then first display area A and first display area C are the first display areas adjacent to the selected second display area. Using a pre-configured model generation service, a video segment between the first media material displayed in first display area A and first media material displayed in first display area C is generated based on the first media material displayed in first display area A, first media material displayed in first display area C, and the first descriptive text. Using the same method, video segments between the first media materials in every two first display areas within the first track are generated to obtain the target video.
[0057] The technical solution of this disclosure embodiment, by displaying a first track, including a first display area and a second display area, and the first display area including first media material, allows for flexible editing of the first media material based on creative needs through intuitive display of the first media material. Furthermore, preset placeholder content for video clips is displayed in the second display area, enabling precise control over the position of the video clips to be generated. Further, in response to a trigger operation on the second display area, an editing area and a first control are displayed on the first page. The editing area is used to input first descriptive text, which represents at least part of the screen change information corresponding to the first media material. Since the video clip is generated based on the first media material and the first descriptive text, and both the first media material and the first descriptive text can be flexibly edited, it is convenient to generate a target video that meets the user's expectations. In response to a trigger operation on the first control, a second page is displayed, showing the target video. The technical solution of this disclosure embodiment provides an optimized video generation method, enabling flexible and accurate generation of a target video based on multiple frames of first media material and first descriptive text, solving the problem of unsatisfactory video effects in related technologies, and improving the quality of video generation.
[0058] Figure 3 This is a flowchart illustrating another video generation method provided by an embodiment of the present disclosure. Based on the above embodiments, this embodiment specifically defines the display of the editing area on the first page.
[0059] like Figure 3 As shown, the method includes:
[0060] S310. Display a first track on the first page. The first track includes a first display area and a second display area. The first display area is used to display first media material, and the second display area includes preset placeholder content for video clips.
[0061] S320. In response to a trigger operation on the second display area, a first editing area, a second editing area, and a first control are displayed on the first page. The first descriptive text in the first editing area is used to describe the screen change information corresponding to the target video. The first descriptive text in the second editing area is used to describe the screen change information corresponding to a portion of the first media material. The display forms of the first editing area and the second editing area are linked.
[0062] In this embodiment, the first descriptive text within the first editing area ensures that all video clips of the first media materials within the first track exhibit the same special effects. Alternatively, a portion of the first media materials may refer to every two first media materials within the first track. Alternatively, a portion of the first media materials may refer to every three first media materials within the first track. The number of first media materials involved in a portion of the first media materials can be set according to actual needs. That is, video clips are generated using the first descriptive text within the second editing area and every N first media materials, where N is a positive integer set according to actual needs.
[0063] The interaction between the first and second editing areas can be described as follows: the display modes of the first and second editing areas are inverse processes of each other. The selection state of the second editing area is related to the display modes of the first and second editing areas. The display mode represents the visualization state of the first and second editing areas. For example, display modes include expanded and collapsed modes, or maximized and minimized modes, etc.
[0064] Optionally, in response to a trigger operation on the second display area, the first page displays an expanded first editing area and a collapsed second editing area. The position of the video clip to be generated can be switched by triggering an operation on the second display area within the first track.
[0065] S330, in response to the switching operation for the first editing area and the second editing area, switch the display mode of the first editing area and the second editing area, and switch the selection state of the second display area.
[0066] The switching operations include clicking, selecting, or boxing the first or second editing area. Optionally, preset switching controls are displayed in both the first and second editing areas. In response to a triggering operation on the preset switching control, the display state of the first and second editing areas is switched. For example, assuming the first editing area is in an expanded state, clicking the preset switching control within the first editing area displays the collapsed state of the first editing area on the first page, and the expanded state of the second editing area.
[0067] For example, in response to a trigger operation on the first editing area, the first editing area is switched from a first form to a second form, and the second editing area is switched from a second form to a first form. Similarly, in response to a trigger operation on the second editing area, the second editing area is switched from a first form to a second form, and the first editing area is switched from a second form to a first form.
[0068] The first and second forms are two display forms that are inverse processes of each other. For example, the first form is an expanded form, and the second form is a collapsed form. If the first editing area is in an expanded form, then the second editing area is in a collapsed form. If the first editing area is in a collapsed form, then the second editing area is in an expanded form. Optionally, if the editing area includes first descriptive text, then after the editing area is converted to a collapsed form, a preset identifier is displayed at the position corresponding to the editing area in the collapsed form. For example, the preset identifier can be a preset graphic.
[0069] In this embodiment, the selection state of the second display area is associated with the display forms of the first and second editing areas. In some embodiments, if the first editing area is in a first form, all second display areas within the first track are selected. If the second editing area is in a first form, the portion of the second display areas within the first track corresponding to the second editing area are selected. For example, the second display area corresponding to the second editing area is determined based on preset prompt text within the second editing area. Assuming that first display area A, second display area B, and first display area C are three consecutive areas within the first track, first display area A displays first media material A, first display area C displays first media material C, and the preset prompt text represents the generation of a video segment between first media material A and first media material C, the second display area corresponding to the second editing area can be determined as second display area B. This embodiment, by associating the selection state of the second display area with the first and second editing areas, achieves an intuitive display of the first media material corresponding to the video segment to be generated, facilitating the quick determination of the first media material corresponding to the video segment through the selection state of the second display area, and assisting the user in editing the first descriptive text related to the video segment based on the determined first media material.
[0070] Optionally, in response to the second editing area being in the first form, preset configuration items are displayed on the first page, wherein the preset configuration items represent the attribute information of the video segment. Optionally, the attribute information may include video duration, etc. In response to a trigger operation on the preset configuration items, multiple candidate configuration items are displayed, for example, multiple video duration options are displayed. In response to a trigger operation on a target configuration item among the multiple candidate configuration items, the preset configuration items are updated based on the target configuration item. In this embodiment of the disclosure, by displaying preset configuration items in response to the second editing area being in the first form, the attributes of video segments between some first media materials can be edited through the editing operation of the preset configuration items, which facilitates the individual editing of the relevant attributes of each video segment and realizes flexible adjustment of the attributes of different video segments.
[0071] Figure 4 This is a schematic diagram of another first page provided as an embodiment of this disclosure. (See diagram below.) Figure 4As shown, the first page 410 displays a first track 420, a first editing area 430, a second editing area 440, and a first control 450. The first track 420 includes a first display area 460 and a second display area 470. The first display area 460 displays first media material. The second display area 470 displays attribute information of a video clip. The first editing area 430 includes a first preset switching control 431. The second editing area 440 includes a second preset switching control 441. The first editing area 430 is in an expanded state, and the second editing area 440 is in a collapsed state. All second display areas are selected (indicated by a thick border). In response to a trigger operation on the first preset switching control 431, the first page 410 displays the expanded state of the second editing area 440, the collapsed state of the first editing area 430, and the portion of the second display area 470 corresponding to the second editing area 430 is converted to a selected state. For example, the first second display area 470 in the current first page is selected.
[0072] S340. In response to a trigger operation on the first control, a second page is displayed, on which a target video is displayed. The target video includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
[0073] For example, in response to a trigger operation on the first control, video clips corresponding to two adjacent first media materials are generated based on the first descriptive text in the first editing area corresponding to the selected second display area and / or the first descriptive text in the second editing area and the first media material.
[0074] It should be noted that if the first editing area includes the first descriptive text and the second editing area does not include the first descriptive text, then based on the first descriptive text and the first media material in the first track, the video content corresponding to the first descriptive text is added to the video clips corresponding to every two first media materials in the first track. Optionally, if the first editing area does not include the first descriptive text and the second editing area includes the first descriptive text, then based on the first descriptive text and the first media material in the first track corresponding to the selected second display area, a video clip corresponding to a portion of the first media material in the first track is generated. Optionally, if both the first and second editing areas include the first descriptive text, the video content in both cases is superimposed. By inputting the first descriptive text in the first and second editing areas respectively, it is possible to generate both the global video content between every two first media materials in the first track and the video clips between the portion of the first media material corresponding to the selected second display area, and then add the global video content to the generated video clips.
[0075] The technical solution of this disclosure embodiment displays a first track on a first page, the first track including a first display area and a second display area. In response to a trigger operation on the second display area, a first editing area, a second editing area, and a first control are displayed on the first page. Since the display forms of the first and second editing areas are linked, in response to a switching operation on the first and second editing areas, the display forms of the first and second editing areas are switched, and the selection state of the second display area is switched, increasing the interactive interest and improving the user experience. Furthermore, associating the selection state of the second display area with the display forms of the first and second editing areas allows for a more intuitive display of the first media material corresponding to the video clip to be generated, facilitating convenient and rapid input of the first descriptive text. Further, the corresponding video clip is quickly generated based on the first media material and the first descriptive text.
[0076] Figure 5 This is a flowchart illustrating another video generation method provided in this disclosure. Based on the above embodiments, this disclosure further defines the display method of the first page and the acquisition method of the first media material.
[0077] like Figure 5 As shown, the method includes:
[0078] S510. Display the second media material in the second track of the preset page.
[0079] The preset page is an interactive page for editing media data. The second track is the track within the preset page that carries media materials. Media materials include images or videos, etc. The second media material represents the original video or original image collection.
[0080] S520: In response to the interactive operation on the second media material meeting preset conditions, the second control is displayed on the second track.
[0081] In some embodiments, interactive operations on second media materials in a second track are acquired. Preset conditions are used to determine whether the interactive operations continuously add media materials to the second track. A pre-configured model service can be used to identify whether a user intends to continuously add media materials to the second track based on the interactive operations on the second media materials; if so, the interactive operations on the second media materials are determined to meet the preset conditions.
[0082] The second control serves as the entry point to the third page. The third page represents the media upload page. Users can select the method for obtaining the first media material through the third page. Optionally, the method includes obtaining the first media material from a preset media resource library or generating the first media material. The preset media resource library includes local photo albums or cloud photo albums, etc.
[0083] S530. In response to a trigger operation on the second control, a third page is displayed, wherein the third page includes a third control, which is used to trigger the acquisition of the first media material.
[0084] For example, in response to a triggering action on the second control, a third page is displayed. The third page displays a reference frame of the third control and the target video. The reference frame of the target video can be a selected video frame of the second media material on the preset page. Alternatively, the reference frame of the target video can be the last video frame of the second media material preceding the target video to be added on the preset page, etc. The third control is used to trigger the addition of the first media material after the reference frame.
[0085] Figure 6 This is a schematic diagram of a third page provided as an embodiment of this disclosure. (As shown...) Figure 6 As shown, a second media asset is displayed in the second track 620 of the preset page 610. In response to an interactive operation on the second media asset that meets preset conditions, a second control 630 is displayed at the position corresponding to the last frame of the second media asset. In response to a triggering operation of the second control 630, a third page 640 is displayed. A reference frame 650 and a third control 660 are displayed on the third page 640.
[0086] S540. In response to the trigger operation on the third control, display the material generation control.
[0087] See Figure 6 In response to a trigger operation on the third control 660, a control panel 670 is displayed at the top of the third page. The control panel 670 displays a media upload control 680 and a media generation control 690. The media upload control 680 is used to trigger the retrieval of the first media material from a preset media resource library. The media generation control 690 is used to trigger the generation of the first media material.
[0088] S550. In response to the trigger operation of the media generation control, a fourth page is displayed. The fourth page includes a media display area and a second descriptive text. The media display area includes the previous frame of the media to be generated and the preset placeholder content of the media to be generated.
[0089] The fourth page is the media generation page. Users obtain the necessary prompts for media generation through this page. The fourth page includes a media display area and an input box for entering a second descriptive text. In response to text input, the second descriptive text is displayed within the input box. Alternatively, in response to voice input, the second descriptive text is determined based on voice data and displayed within the input box. The second descriptive text describes the media material to be generated. Optionally, it can describe the differences between the media material to be generated and the previous frame, enabling fast and accurate media material generation. For example, the input box also displays model selection options and generation controls. The model selection options display candidate models for user selection. The generation controls trigger the media material generation operation.
[0090] Optionally, a media display area can be displayed above the input box. This area shows the previous frame of the media to be generated and the preset placeholder content for that frame. For example, if the media to be generated is frame k, the display area shows the image of frame (k-1) and the corresponding preset placeholder content. The preset placeholder content for frame k includes a preset graphic and the frame's identifier information. If the media to be generated becomes frame k+1, the display area shows the image of frame k and the corresponding preset placeholder content for frame k+1.
[0091] Figure 7 This is a schematic diagram of a fourth page provided as an embodiment of this disclosure. (See diagram below.) Figure 7 As shown, the fourth page 710 includes a media display area 720 and an input box 730. The media display area 720 includes the previous frame 740 of the media material to be generated and preset placeholder content 750 for the media material to be generated. The input box 730 displays model selection options 760 and generation controls 770. The input box 730 also displays a second descriptive text.
[0092] S560. In response to the confirmation operation for the second description text, display a plurality of candidate media materials, wherein the candidate media materials are media material generation results that conform to the second description text.
[0093] For example, in response to a triggering operation of the generation control within the input box, multiple candidate media assets are generated based on the second descriptive text and the previous frame of the media asset to be generated, using a pre-configured model service. These multiple candidate media assets are then displayed for selection.
[0094] S570. In response to the confirmation operation for the target media material among the plurality of candidate media materials, the first page is displayed.
[0095] For example, in response to a confirmation operation for a target media material among multiple candidate media materials, a first page is displayed. The first page displays the target media material and an add control. In response to a trigger operation for adding the control, a fourth page is displayed, and execution returns to step S560 to generate a first media material that meets the quantity requirements.
[0096] S580. Display a first track on the first page. The first track includes a first display area and a second display area. The first display area is used to display first media material, and the second display area includes preset placeholder content for video clips.
[0097] S590. In response to a trigger operation on the second display area, an editing area and a first control are displayed on the first page, wherein the editing area is used to input first descriptive text, and the first descriptive text represents at least part of the screen change information corresponding to the first media material.
[0098] S5100, In response to the trigger operation on the first control, a second page is displayed, and the target video is displayed on the second page, including preset placeholder content for the video clip, wherein the video clip is generated based on the first media material and the first descriptive text.
[0099] The technical solution of this disclosure embodiment, when the interactive operation of the second media material in the second track meets preset conditions, displays a second control in the second track. Since the second control is the entry point to the third page, the length of the access link to the third page can be shortened, enabling fast access to the third page. Furthermore, in response to the triggering operation of the third control on the third page, a material generation control is displayed. In response to the triggering operation of the material generation control, a fourth page is displayed. In response to the confirmation operation of the second description text on the fourth page, multiple candidate media materials are displayed. This enables the generation of a first media material that meets the user's expectations based on the second description text when no suitable material is available, enriching the sources of the first media material and reducing the difficulty of obtaining materials.
[0100] Figure 8 This is a schematic diagram of the structure of a video generation device provided in an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.
[0101] like Figure 8 As shown, the device includes: a first display module 810, a second display module 820, and a third display module 830.
[0102] The first display module 810 is used to display a first track on a first page. The first track includes a first display area and a second display area. The first display area is used to display first media material, and the second display area includes preset placeholder content for video clips.
[0103] The second display module 820 is used to display an editing area and a first control on the first page in response to a trigger operation on the second display area, wherein the editing area is used to input first descriptive text, and the first descriptive text represents at least part of the screen change information corresponding to the first media material;
[0104] The third display module 830 is used to display a second page in response to a trigger operation on the first control, and to display a target video on the second page. The target video includes multiple video clips, wherein the target video is generated based on the first media material and the first descriptive text.
[0105] Optionally, the second display module 820 is specifically used for:
[0106] The first page displays a first editing area and a second editing area. The first descriptive text in the first editing area is used to describe the screen change information corresponding to the target video, and the first descriptive text in the second editing area is used to describe the screen change information corresponding to a portion of the first media material. The display forms of the first editing area and the second editing area are linked.
[0107] Optionally, it also includes:
[0108] In response to a switching operation on the first editing area and the second editing area, the display mode of the first editing area and the second editing area are switched, and the selection state of the second display area is switched.
[0109] Optionally, switching the display mode of the first editing area and the second editing area in response to a switching operation on the first editing area and the second editing area includes:
[0110] In response to a trigger operation on the first editing area, the first editing area is switched from a first form to a second form, and the second editing area is switched from a second form to a first form;
[0111] In response to a trigger operation on the second editing area, the second editing area is switched from a first mode to a second mode, and the first editing area is switched from a second mode to a first mode.
[0112] Optionally, it also includes:
[0113] In response to the second editing area being in the first form, preset configuration items are displayed on the first page, wherein the preset configuration items represent the attribute information of the video segment.
[0114] Optionally, it also includes:
[0115] Display the second media material in the second track of the preset page;
[0116] In response to an interactive operation on the second media material that meets preset conditions, a second control is displayed on the second track;
[0117] In response to a trigger operation on the second control, a third page is displayed, wherein the third page includes a third control used to trigger the retrieval of the first media material.
[0118] Optionally, it also includes:
[0119] In response to a triggering operation on a special effects control in the preset page, a special effects page is displayed, wherein the special effects page includes multiple candidate special effects templates;
[0120] In response to a trigger operation on the target effect template, a third page is displayed, wherein the third page includes a third control, which is used to trigger the acquisition of the first media material.
[0121] Optionally, it also includes:
[0122] In response to a trigger operation on the third control, the material generation control is displayed;
[0123] In response to a trigger operation on the media generation control, a fourth page is displayed. The fourth page includes a media display area and a second descriptive text. The media display area includes the previous frame of the media to be generated and the preset placeholder content of the media to be generated.
[0124] In response to a confirmation operation on the second description text, multiple candidate media materials are displayed, wherein the candidate media materials are media material generation results that conform to the second description text;
[0125] In response to the confirmation operation for the target media material among the multiple candidate media materials, the first page is displayed.
[0126] The video generation apparatus provided in this disclosure can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0127] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0128] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 9 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 9 The diagram below shows the structure of the terminal device or server 900. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0129] like Figure 9 As shown, electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. An edit / output (I / O) interface 905 is also connected to bus 904.
[0130] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0131] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0132] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0133] The electronic device provided in this embodiment and the video generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0134] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video generation method provided in the above embodiments.
[0135] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0136] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0137] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0138] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0139] The first track is displayed on the first page. The first track includes a first display area and a second display area. The first display area is used to display the first media material, and the second display area includes preset placeholder content for video clips.
[0140] In response to a trigger operation on the second display area, an editing area and a first control are displayed on the first page, wherein the editing area is used to input first descriptive text, and the first descriptive text represents at least part of the screen change information corresponding to the first media material;
[0141] In response to a trigger operation on the first control, a second page is displayed, on which a target video is displayed. The target video includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
[0142] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0144] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0145] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0146] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0148] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0149] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video generation method, characterized in that, include: The first track is displayed on the first page. The first track includes a first display area and a second display area. The first display area is used to display the first media material, and the second display area includes preset placeholder content for video clips. In response to a trigger operation on the second display area, an editing area and a first control are displayed on the first page, wherein the editing area is used to input first descriptive text, and the first descriptive text represents at least part of the screen change information corresponding to the first media material; In response to a trigger operation on the first control, a second page is displayed, on which a target video is displayed. The target video includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
2. The method according to claim 1, characterized in that, The provision of an editing area on the first page includes: The first page displays a first editing area and a second editing area. The first descriptive text in the first editing area is used to describe the screen change information corresponding to the target video, and the first descriptive text in the second editing area is used to describe the screen change information corresponding to a portion of the first media material. The display forms of the first editing area and the second editing area are linked.
3. The method according to claim 2, characterized in that, Also includes: In response to a switching operation on the first editing area and the second editing area, the display mode of the first editing area and the second editing area are switched, and the selection state of the second display area is switched.
4. The method according to claim 3, characterized in that, The step of switching the display mode of the first editing area and the second editing area in response to a switching operation on the first editing area and the second editing area includes: In response to a trigger operation on the first editing area, the first editing area is switched from a first form to a second form, and the second editing area is switched from a second form to a first form; In response to a trigger operation on the second editing area, the second editing area is switched from a first mode to a second mode, and the first editing area is switched from a second mode to a first mode.
5. The method according to claim 4, characterized in that, Also includes: In response to the second editing area being in the first form, preset configuration items are displayed on the first page, wherein the preset configuration items represent the attribute information of the video segment.
6. The method according to claim 1, characterized in that, Also includes: Display the second media material in the second track of the preset page; In response to an interactive operation on the second media material that meets preset conditions, a second control is displayed on the second track; In response to a trigger operation on the second control, a third page is displayed, wherein the third page includes a third control used to trigger the retrieval of the first media material.
7. The method according to claim 1, characterized in that, Also includes: In response to a triggering operation on a special effects control in the preset page, a special effects page is displayed, wherein the special effects page includes multiple candidate special effects templates; In response to a trigger operation on the target effect template, a third page is displayed, wherein the third page includes a third control, which is used to trigger the acquisition of the first media material.
8. The method according to claim 6 or 7, characterized in that, Also includes: In response to a trigger operation on the third control, the material generation control is displayed; In response to a trigger operation on the media generation control, a fourth page is displayed. The fourth page includes a media display area and a second descriptive text. The media display area includes the previous frame of the media to be generated and the preset placeholder content of the media to be generated. In response to a confirmation operation on the second description text, multiple candidate media materials are displayed, wherein the candidate media materials are media material generation results that conform to the second description text; In response to the confirmation operation for the target media material among the multiple candidate media materials, the first page is displayed.
9. A video generation apparatus, characterized in that, include: The first display module is used to display a first track on the first page. The first track includes a first display area and a second display area. The first display area is used to display first media material, and the second display area includes preset placeholder content for video clips. The second display module is used to display an editing area and a first control on the first page in response to a trigger operation on the second display area. The editing area is used to input first descriptive text, which represents at least part of the screen change information corresponding to the first media material. The third display module is used to display a second page in response to a trigger operation on the first control. The second page displays a target video, which includes multiple video clips, wherein the video clips are generated based on the first media material and the first descriptive text.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any one of claims 1-8.
11. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the video generation method as described in any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video generation method as described in any one of claims 1-8.