Video generation method and apparatus, and device and storage medium

By acquiring media materials input by the user and determining the target video script from the list of video scripts, a video draft is generated, which solves the problem of the single video generation method in the existing technology, realizes diversified video generation methods, and improves the user experience.

WO2026001828A1PCT designated stage Publication Date: 2026-01-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102120
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-19
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet users' diverse video generation needs and lack a variety of video generation methods.

Method used

By acquiring media materials input by the user, the target video script is determined from the list of video scripts, and a video draft is generated based on the media materials input by the user and the target video script. The video draft and the target video script have a corresponding relationship, and there is a preset order relationship between the video clips and the script clips.

Benefits of technology

It enriches the video generation methods, meets users' diverse video generation needs, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102120_02012026_PF_FP_ABST
    Figure CN2025102120_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a video generation method and apparatus, and a device and a storage medium. The method comprises: first, acquiring a media material input by a user, and determining target video text from a video text list; and then, on the basis of the media material input by the user and the target video text, generating at least one video draft, wherein there is a correspondence between the video draft and the target video text, the video draft comprises a plurality of video segments, and there is a correspondence between the video segments and text segments.
Need to check novelty before this filing date? Find Prior Art

Description

A video generation method, device, apparatus and storage medium

[0001] Cross-reference to Related Applications

[0002] The present application claims priority to the Chinese patent application No. 202410866693.6, filed on June 28, 2024, and entitled "A video generation method, device, apparatus and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates to the field of data processing, and particularly relates to a video generation method, device, apparatus and storage medium. BACKGROUND

[0004] With the continuous development of video generation technology, video generation related functions are also more diversified. For example, generating a video by using a video template. SUMMARY

[0005] The present disclosure provides a video generation method, device, apparatus and storage medium.

[0006] In a first aspect, the present disclosure provides a video generation method, comprising:

[0007] obtaining media material input by a user;

[0008] and determining a target video script from a video script list, wherein the video script list displays a video script generated based on multimedia description content input by the user, and / or a video script obtained by editing the video script generated based on the multimedia description content input by the user;

[0009] generating at least one video draft based on the media material input by the user and the target video script;

[0010] The video draft and the target video script have a corresponding relationship, the video draft includes a plurality of video segments, the plurality of video segments and a plurality of script segments have a corresponding relationship, the plurality of script segments are obtained by text segmentation of the target video script corresponding to the video draft, the plurality of video segments are extracted from the media material input by the user for the plurality of script segments respectively, and the display order of the plurality of video segments in the video draft is determined based on a preset order relationship between the corresponding script segments.

[0011] In an optional implementation, before the determining of the target video script from the video script list, the method further comprises:

[0012] obtaining multimedia description content in the target text edit box; wherein the multimedia description content comprises first text content extracted from the target audio / video resource, and / or second text content input based on the target text edit box;

[0013] generating at least one video script based on the multimedia description content, and displaying the at least one video script in the video script list.

[0014] In an optional implementation, before the determining the target video script from the video script list, the method further comprises:

[0015] receiving multimedia description content input for at least one target attribute of the target object;

[0016] generating at least one video script based on the multimedia description content corresponding to the at least one target attribute respectively, and displaying the at least one video script in the video script list.

[0017] In an optional implementation, the receiving multimedia description content input for at least one target attribute of the target object comprises:

[0018] in response to multimedia description content input for a first target attribute of the target object, displaying at least one candidate recommended content corresponding to a second target attribute of the target object; wherein the at least one candidate recommended content is determined based on the multimedia description content;

[0019] receiving recommended content selected for the second target attribute from the at least one candidate recommended content;

[0020] Correspondingly, the generating at least one video script based on the multimedia description content corresponding to the at least one target attribute respectively comprises:

[0021] generating at least one video script based on the multimedia description content corresponding to the first target attribute and the recommended content selected for the second target attribute.

[0022] In an optional implementation, before the determining the target video script from the video script list, the method further comprises:

[0023] in response to an edit trigger operation for a first video script in the video script list, displaying a text edit box corresponding to the first video script;

[0024] receiving a text edit operation for the first video script based on the text edit box, to obtain an edited video script;

[0025] updating the first video script in the video script list with the edited video script.

[0026] In an optional implementation, before the generating the at least one video draft based on the media material input by the user and the target video script, the method further includes:

[0027] In response to a copy operation on a second video script in the video script list, adding a copy video script of the second video script in the video script list;

[0028] In response to a text adjustment operation on the copy video script, generating a third video script, and updating the copy video script in the video script list with the third video script.

[0029] In an optional implementation, before the generating the at least one video draft based on the media material input by the user and the target video script, the method further includes:

[0030] In response to a target interaction operation triggered on a fourth video script in the video script list, displaying an interaction window;

[0031] Receiving feedback information on the fourth video script based on the interaction window, and uploading the feedback information to a server.

[0032] In an optional implementation, before the generating the at least one video draft based on the media material input by the user and the target video script, the method further includes:

[0033] In response to a deletion operation triggered on a fifth video script in the video script list, deleting the fifth video script from the video script list.

[0034] In an optional implementation, after the generating the at least one video draft based on the media material input by the user and the target video script, the method further includes:

[0035] Displaying a video draft list on a video preview page, wherein the video draft list includes a display area corresponding to each of the at least one video draft, and the display area is used to display information of the corresponding video draft and a target video script corresponding to the video draft;

[0036] And playing a first video draft in the video draft list on the video preview page.

[0037] In an optional implementation, the generating the at least one video draft based on the media material input by the user and the target video script includes:

[0038] text segmentation is performed on the target video script to obtain a plurality of script segments, wherein the plurality of script segments have a preset sequence relationship;

[0039] corresponding video segments are respectively cropped from the media material input by the user for the plurality of script segments;

[0040] based on the preset sequence relationship, the video segments corresponding to the plurality of script segments are spliced to obtain a video draft corresponding to the target video script.

[0041] In a second aspect, the present disclosure provides a video generation device, the device comprising:

[0042] a first acquisition module configured to acquire media material input by a user;

[0043] a first determination module configured to determine a target video script from a video script list, wherein the video script list displays video scripts generated based on multimedia description content input by a user, and / or video scripts obtained after editing the video scripts generated based on the multimedia description content input by the user;

[0044] a first generation module configured to generate at least one video draft based on the media material input by the user and the target video script;

[0045] wherein the video draft and the target video script have a corresponding relationship, the video draft includes a plurality of video segments, the plurality of video segments and a plurality of script segments have a corresponding relationship, the plurality of script segments are obtained by performing text segmentation on a target video script corresponding to the video draft, the plurality of video segments are extracted from the media material input by the user for the plurality of script segments, and a display order of the plurality of video segments in the video draft is determined based on a preset sequence relationship between corresponding script segments.

[0046] In a third aspect, the present disclosure provides a computer-readable storage medium, the computer-readable storage medium storing instructions, when the instructions are run on a terminal device, the terminal device implements the above method.

[0047] In a fourth aspect, the present disclosure provides a video generation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, when the processor executes the computer program, the method described above is implemented.

[0048] In a fifth aspect, the present disclosure provides a computer program product, the computer program product comprising computer programs / instructions, when the computer programs / instructions are executed by a processor, the method described above is implemented.

[0049] Compared with the prior art, the technical scheme provided by the embodiments of the present disclosure has at least the following advantages:

[0050] The embodiment of the present disclosure provides a video generation method. First, media material input by a user is obtained, and a target video script is determined from a video script list, wherein the video script list displays a video script generated based on multimedia description content input by the user and / or a video script obtained by editing the video script generated based on the multimedia description content input by the user. Then, at least one video draft is generated based on the media material input by the user and the target video script, wherein the video draft has a corresponding relationship with the target video script, the video draft includes a plurality of video clips, the video clip has a corresponding relationship with a script clip, the script clip is obtained by text segmentation of the target video script corresponding to the video draft, the plurality of video clips are extracted from the media material input by the user for the plurality of script clips, and the display order of the plurality of video clips in the video draft is determined based on a preset order relationship between the corresponding script clips. BRIEF DESCRIPTION OF DRAWINGS

[0051] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0052] In order to more clearly illustrate the technical scheme in the embodiments of the present disclosure or the prior art, the accompanying drawings required to be used in the embodiments or the prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0053] FIG. 1 is a flowchart of a video generation method according to an embodiment of the present disclosure;

[0054] FIG. 2 is a schematic diagram of a video generation page according to an embodiment of the present disclosure;

[0055] FIG. 3 is a schematic diagram of another video generation page according to an embodiment of the present disclosure;

[0056] FIG. 4 is a schematic diagram of another video generation page according to an embodiment of the present disclosure;

[0057] FIG. 5 is a schematic diagram of a video preview page according to an embodiment of the present disclosure;

[0058] FIG. 6 is a schematic diagram of a video generation device according to an embodiment of the present disclosure;

[0059] FIG. 7 is a schematic diagram of a video generation device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0060] To meet the diversified video generation function needs of people, how to further enrich the video generation mode is a technical problem to be solved at present.

[0061] In order to enable the above-mentioned purposes, features and advantages of the present disclosure to be more clearly understood, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0062] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the description are only a part of the embodiments of the present disclosure, not all the embodiments.

[0063] With the continuous development of video generation technology, video generation related functions are also more diversified. For example, generating a video by using a video template, etc.

[0064] To meet the diversified video generation function needs of people, how to further enrich the video generation mode is a technical problem to be solved at present.

[0065] To this end, the embodiments of the present disclosure provide a video generation method, first, obtaining media material input by a user, and determining a target video script from a video script list, wherein the video script list displays a video script generated based on multimedia description content input by the user. Then, based on the media material input by the user and the target video script, at least one video draft is generated, wherein the video draft has a corresponding relationship with the target video script, the video draft includes a plurality of video segments, the video segment has a corresponding relationship with a script segment, the script segment is obtained by text segmentation of the target video script corresponding to the video draft, the plurality of video segments are extracted from the media material input by the user for the plurality of script segments, and the display order of the plurality of video segments in the video draft is determined based on the preset order relationship between the corresponding script segments.

[0066] The embodiments of the present disclosure support the function of generating a video draft based on a target video script generated by multimedia description content input by a user and media material input by the user, enrich the video generation mode, and thus meet the diversified video generation function needs of the user.

[0067] Based on this, the embodiments of the present disclosure provide a video generation method, referring to FIG. 1, a flowchart of a video generation method provided by the embodiments of the present disclosure, the method specifically includes:

[0068] S101: Obtain media material input by a user.

[0069] The video generation method provided by the embodiments of the present disclosure can be applied to a client, which can include a desktop computer, a notebook computer, a tablet computer, and the like.

[0070] The media material in the embodiments of the present disclosure is one or more media materials uploaded by a user, which can include media materials of media types such as audio and pictures, and can include media materials imported from a photo album of the user.

[0071] For the manner of obtaining the user input media material, the user input media material can be obtained in a manner that the user triggers a material upload control.

[0072] In actual application, in order to ensure the effect of video generation, the embodiments of the present disclosure can set the number of uploaded media materials, the time length of the media materials, and the like. For example, the number of uploaded media materials is between 5 and 30, and the time length of the media materials is within 1 second to 5 minutes.

[0073] As shown in FIG. 2, it is a schematic diagram of a video generation page provided by the embodiments of the present disclosure, which displays a material upload control 201.

[0074] S102: Determine a target video script from a video script list.

[0075] The video script list displays video scripts generated based on user input multimedia description content, and / or video scripts obtained by editing the video scripts generated based on the user input multimedia description content.

[0076] The video script list in the embodiments of the present disclosure is used to display video scripts, and the target video script is any one of the video scripts in the video script list. Specifically, when a selection operation on a video script is received, the video script is taken as the target video script.

[0077] In addition, the selection operation on the video script can also set a default selection manner to determine the target video script, and on the basis of the default selection, the user can manually change the selected video script.

[0078] In actual application, in order to facilitate the user to view the video scripts in the video script list, the video script list in the embodiments of the present disclosure further includes at least one script list unit, wherein the script list unit has a corresponding relationship with a video script, and one video script is displayed in each script list unit.

[0079] As shown in FIG. 2, the video generation page displays a video script list 202, and the video script list displays a script list unit, as shown in 203. The script list unit 203 displays a target video script 204, and the target video script is in a selected state.

[0080] The video script in the embodiment of the present disclosure is generated by the user input multimedia description content. In an optional implementation, the video script can be intelligently generated based on the user input multimedia description content. Specifically, when a script generation trigger operation is received, the multimedia description content is intelligently expanded to obtain expanded description content, and the expanded description content is used as the video script. For expanding the multimedia description content, a model can be used for expansion, and the specific implementation of the present disclosure is not limited in any way. The trigger operation for generating the script can include setting a script generation control on the video generation page. As shown in the script generation control 205 displayed in the video generation page shown in FIG. 2.

[0081] The multimedia description content can include multimedia description content in the form of voice, text, etc. The multimedia description content usually includes a description of the subject description object in the video draft to be generated. For example, the video draft to be generated is a marketing video for a certain object, and accordingly, the multimedia description content can include description content for the object. In fact, the specific description content included in the multimedia description content is not limited.

[0082] In an optional implementation, the multimedia content can include first text content extracted from a target audio / video resource. Obtaining the multimedia content in the target text editing box can include: first, obtaining the target audio / video resource in the target text editing box, then extracting voice data in the target audio / video resource based on the target audio / video resource, then performing voice recognition on the extracted voice data, recognizing the corresponding text content, and displaying the recognized text content as the first text content in the target text editing box. The target audio / video resource can include audio resources, video resources, etc. in the user album, and can also include audio resources, video resources, etc. obtained based on the shooting page.

[0083] The manner of obtaining the target audio resource can include: setting an extraction control in the target text edit box, obtaining the target audio and video resource by triggering the extraction control, and then extracting the text content corresponding to the voice data in the target audio and video resource. Specifically, when receiving a triggering operation on the extraction control in the target text edit box, the current page jumps to a user album selection page, and the user album selection page displays audio and video resources, and the user can select one of the audio and video resources as the target audio and video resource. As shown in the extraction control 208 displayed in the video generation page shown in FIG. 2.

[0084] In actual application, after the text content is extracted from the target audio resource, the embodiment of the present disclosure also supports editing operation on the text content, and after the text content is modified, the modified text content is displayed as the first text content in the target text edit box.

[0085] In another optional implementation, the multimedia description content can also include second text content manually input by the user in the target text edit box, and obtaining the multimedia content can include: when receiving an input operation of triggering the target text edit box, displaying the text content in the target text edit box, and obtaining the text content as the second text content. The second text content can include any text content, and the specific form can include: a short sentence, a long sentence, a paragraph, a keyword, etc.

[0086] As shown in the video generation page shown in FIG. 2, the target text edit box 206 is displayed, and when receiving an input operation of the user in the target text edit box, the text content is displayed in the target text edit box, and the text content is obtained as the second text content.

[0087] In another optional implementation, the multimedia description content can include both the first text content extracted from the target audio and video resource and the second text content input by the user in the target text edit box. The display order of the first text content and the second text content in the target text edit box can be determined based on requirements.

[0088] In one application scenario, first, the user triggers the target text edit box to obtain the second text content manually input by the user in the target text edit box, and then, based on the user triggering the extraction control, the target audio and video resource is determined, and the first text content is extracted from the target audio and video resource and displayed in the target text edit box.

[0089] In another application scenario, first, the user triggers the extraction control to determine the target audio and video resource, and extracts the first text content from the target audio and video resource, and then displays the first text content in the target text edit box. Based on the display of the first text content, the user manually inputs the text content in the target text edit box.

[0090] In actual applications, the display positions of the first text content and the second text content in the target text editing box can also be determined based on a current positioning position, for example, a position of a mouse cursor. For example, the second text content is displayed in the target text editing box, and if the first text content is inserted before the second text content, the mouse cursor can be inserted before a first character display position of the second text content.

[0091] Based on the above embodiments, the text description content in the target text editing box is obtained by extraction from the target audio / video or based on manual input by the user. To prompt the user about the manner of obtaining the multimedia description content, the disclosure embodiments can also display an input or extraction text content portal on the video generation page, so as to facilitate the user to understand the manner of obtaining the text description content.

[0092] As shown in the video generation page in FIG. 2, the input or extraction text content portal 207 is displayed. In an application scenario, when the user triggers the input or extraction text content portal 207, the target input box 206 is displayed for inputting the multimedia description content.

[0093] Based on the video script generated based on the multimedia description content input by the user, to provide more video scripts for the user, the disclosure embodiments also support continuing to generate video scripts based on the generated video scripts. In an optional implementation, based on the video script list displaying the video scripts, when a continue generation operation for the video script is received, the video script is continued to be generated based on the obtained multimedia content, and is displayed in the video script list in a time sequence, where the continued video script is similar to the first generated video script.

[0094] In addition, based on the video script generated based on the multimedia description content input by the user, the disclosure embodiments also support editing the multimedia description content and generating a video script based on the modified multimedia description content. In an optional implementation, based on the video script list displaying the video scripts, if an editing operation for the multimedia content is received, when a generation operation for the video script is received, the multimedia description content is re-obtained to generate a video script, and is displayed in the video script list in a time sequence, where the video script list displays the video script generated based on the multimedia description content before the editing and the video script generated based on the multimedia description content after the editing.

[0095] S103: generating at least one video draft based on the media material input by the user and the target video script.

[0096] The video draft has a corresponding relationship with the target video script, the video draft includes a plurality of video clips, the plurality of video clips has a corresponding relationship with a plurality of script clips, the plurality of script clips are obtained by text segmentation of the target video script corresponding to the video draft, and the plurality of video clips are extracted from the media material input by the user for the plurality of script clips respectively. The display order of the plurality of video clips in the video draft is determined based on the preset order relationship between the corresponding script clips.

[0097] In the embodiments of the present disclosure, the video draft can be a video preliminary draft to be edited, and the video draft has a one-to-one relationship with the target video script.

[0098] The video draft is composed of a plurality of video clips, and each video clip in the video draft has a corresponding relationship with a script clip. The script clip is obtained by text segmentation of the target video script based on a preset order relationship. Then, based on the obtained script clip, a corresponding video shot is matched from the media material, the video clip corresponding to the video shot is cut out from the media material, and then the video clip corresponding to the script clip is spliced based on the preset order relationship of the script clip to generate the video draft. The plurality of video clips can include a highlight clip in the media material.

[0099] The script clip is obtained by text segmentation of the target video script based on a preset order relationship. Specifically, the text segmentation can be based on the paragraphs and context relationship of the target video script.

[0100] Since a plurality of shot lenses can be matched from the media material based on the script clip, the video clip with high similarity that is matched can be preferentially cut out from the media material for splicing to generate the video draft.

[0101] In actual application, before the video clip corresponding to the script clip is spliced to obtain the video draft, a voice clip corresponding to the video clip of the plurality of script clips can also be generated. Specifically, the plurality of script clips can be used to generate a voice clip corresponding to the video clip by human voice reading. The voice clip has the same playing time information as the video clip, that is, the voice clip and the video clip are kept synchronous.

[0102] The method of adjusting the playing time of the voice clip to make the voice clip have the same playing time information as the video clip can include adjusting the playing time of the voice clip generated by human voice reading. Specifically, the speed of human voice reading is slowed down or accelerated according to the playing time of the video clip.

[0103] The video generation method provided in the embodiments of the present disclosure includes the following steps. First, media material input by a user is acquired, and a target video script is determined from a video script list. The video script list displays video scripts generated based on multimedia description content input by the user and / or video scripts obtained by editing the video scripts generated based on the multimedia description content input by the user. Then, at least one video draft is generated based on the media material input by the user and the target video script. The video draft has a corresponding relationship with the target video script. The video draft includes a plurality of video clips. The video clip has a corresponding relationship with a script clip. The script clip is obtained by text segmentation of the target video script corresponding to the video draft. The plurality of video clips are extracted from the media material input by the user for the plurality of script clips. The display order of the plurality of video clips in the video draft is determined based on a preset order relationship between the corresponding script clips.

[0104] The embodiments of the present disclosure support the function of generating a video draft based on a target video script generated based on multimedia description content input by a user and media material input by the user, enrich the video generation mode, and thus meet the diversified video generation function needs of the user.

[0105] In actual application, to improve the acquisition efficiency of the multimedia description content and further improve the user experience, the embodiments of the present disclosure can also support the acquisition of the multimedia description content in an intelligent recommendation manner.

[0106] In an optional implementation, first, an input box of a target attribute of a target object is displayed on a video generation page, and then multimedia description content input for the target attribute of the target object is received. At least one video script is generated based on the multimedia description content corresponding to the target attribute, and the video script is displayed in a video script list. The target object can be a central subject object in the generated video script and also a subject description object included in the media material uploaded by the user. The target attribute can be any attribute of the target object, for example, the name, characteristics, advantages, price, target population, preferential activities, video duration, video size, and the like of the target object.

[0107] For receiving the text content of the target attribute input of the target object, when receiving the multimedia description content of the first target attribute input of the target object, at least one candidate recommended content of the second target attribute of the target object is displayed at a preset position based on the multimedia description content, the candidate recommended content corresponds to the second target attribute of the target object, and then the recommended content selected for the second target attribute of the target object from the candidate recommended content is received, and the recommended content is displayed in the input box of the second target attribute. The preset position can be any display position, specifically, the preset position can be a position near the input box of the second target attribute of the target object.

[0108] That is, according to the multimedia description content of the first target attribute of the target object manually input by the user, the candidate recommended content of the second target attribute of the target object is automatically displayed, and the user can select one of the candidate recommended content as the recommended content of the second target attribute.

[0109] Among them, the first target attribute and the second target attribute are any two different attributes in the target attribute. The multimedia description content of the first target attribute of the target object can be any text content, specifically, for example, idiom, keyword, etc. The candidate recommended content of the second target attribute of the target object can be any recommended text content, specifically, for example, idiom, word, etc. For example, taking the target object as a moon cake as an example, the multimedia description content of the first target attribute (name) of the target object can be "moisturizing cream", and the candidate recommended content of the second target attribute (selling point) of the target object can include "moisturizing effect", "moisturizing and hydrating", etc. The user can select one of them as the recommended content of the second target attribute (selling point).

[0110] As shown in FIG. 3, a schematic diagram of a video generation page provided by an embodiment of the present disclosure is shown. The video generation page displays the first target attribute 301 and the second target attribute 302 of the target object. When the multimedia description content is input in the input box of the first target attribute 301, the candidate recommended content is displayed at a position below the input box of the second target attribute based on the second target attribute of the target attribute, such as the candidate recommended content 303. When the candidate recommended content is selected as the recommended content of the second target attribute, the recommended content is displayed in the input box of the second target attribute, and at this time the selected candidate recommended content is hidden at the position below the input box of the second target attribute.

[0111] After selecting the recommended content corresponding to the second target attribute of the target object, at least one video script is generated based on the multimedia description content corresponding to the first target attribute of the target object and the recommended content selected by the second target attribute, and is displayed in a video script list.

[0112] In actual applications, while the recommended content is selected for the second target attribute of the target object, the disclosed embodiments can also input text content for the second target attribute. Specifically, the text content input for the second target attribute of the target object is received and displayed in the input box of the second target attribute. Then, at least one video script is generated based on the multimedia description content of the first target attribute, the recommended content of the second target attribute, and the text content, and displayed in the video script list.

[0113] In actual applications, before the multimedia description content is obtained in an intelligent recommendation manner, the candidate recommended object content identifier can be displayed to prompt the user of the intelligent generation of the candidate recommended content. Specifically, the intelligent display of the candidate recommended object content identifier can be displayed in the input box of the second target attribute.

[0114] In an alternative embodiment, when the multimedia description content input for the first target attribute of the target object is received, the candidate recommended object content identifier is displayed near the input box of the second target attribute, and the candidate recommended object content identifier is hidden when at least one candidate recommended content is displayed. The candidate recommended object content identifier is used to prompt that the candidate recommended content of the second target attribute is currently being intelligently generated.

[0115] In actual applications, to prompt the user of the intelligent acquisition of the multimedia description content, an intelligent acquisition portal can also be provided, such as the intelligent acquisition portal 304 displayed in the video generation page shown in FIG. 3.

[0116] In the above generation of the video script based on the multimedia description content, the target video script is determined from the video script list.

[0117] In actual applications, before the target video script is determined from the video script list based on the multimedia description content input by the user, an editing operation can also be performed on the first video script in the video script list. Specifically, when at least one video script is displayed in the video script list, one of the video scripts in the video script list is selected as the first video script. When an editing trigger operation for the first video script is received, a text editing box corresponding to the first video script is displayed. Then, a text editing operation for the first video script is received based on the text editing box. Subsequently, the first video script in the video script list is updated based on the edited video script.

[0118] The first video script can be any video script in the video script list. The editing trigger operation for the first video script can include inputting text, modifying, or deleting the text in the first video script.

[0119] To simplify the interaction path, in one application scenario, when a video script is displayed in a video script list, if it is detected that a mouse hovers over one of the video scripts, the input box of the video script enters an editing state, at which time an editing operation can be triggered on the video script.

[0120] To enrich the related interactive operations of video processing, the embodiments of the present disclosure can also support a copy operation on a second video script in a video script list before generating a video draft based on media materials and a target video script, for text adjustment of a copy video script.

[0121] In an optional implementation, in the process of displaying the video scripts in the video script list, when a copy operation on the second video script is received, a copy video script of the second video script is additionally displayed in the video script list. Then, when a text adjustment operation on the copy video script is received, a third video script is generated, and the copy video script in the video script list is updated based on the third video script.

[0122] The second video script is any video script in the video script list. The copy operation on the second video script in the video script list can specifically include displaying a copy creation control in the input box of the second video script. In one application scenario, if it is detected that a mouse hovers over the second video script in the video script list, the copy creation control is displayed in the input box of the second video script. When a user triggers the copy creation control, a copy video script of the second video script is additionally displayed in the video script list.

[0123] As shown in FIG. 4, it is a schematic diagram of a video generation page provided by the embodiments of the present disclosure. The video generation page displays a video script list 401, and the video script list displays video scripts, including a video script 402. When a mouse hovers over the display position of the video script 402 in the video script list, a copy creation control is displayed in the input box of the video script 402. When a user triggers the copy creation control 403, a copy video script 404 of the video script is additionally displayed in the video script list.

[0124] In actual applications, the embodiments of the present disclosure can also support a target interaction operation on a fourth video script in a video script list.

[0125] In an optional implementation, in the process of displaying the video scripts in the video script list, when a target interaction operation triggered on the fourth video script in the video script list is received, an interaction window is displayed on the video script list. Based on the interaction window, feedback information on the fourth video script is received, and the feedback information is uploaded to a server.

[0126] The fourth video script is any video script in the video script list. Triggering the target interaction operation for the fourth video script specifically can include, if it is detected that the mouse hovers over the fourth video script in the video script list, displaying an interaction control in the input box of the fourth video script, when the interaction control is triggered, triggering the execution of the corresponding target interaction operation, and displaying an interaction window. The interaction control is used to trigger the target interaction operation corresponding to the interaction control. The interaction control can include a like control, a dislike control, etc., and the target interaction operation can include a like operation, a dislike operation, etc.

[0127] As shown in FIG. 4, it can be seen from the figure that when the mouse hovers at the display position of the video script 402, an interaction control is displayed in the input box of the video script 402, which includes a like control 405, when the user triggers the like control 405, an interaction window 406 is displayed in the video script list, and the interaction window is used to receive feedback information for the fourth video script.

[0128] Based on the above embodiment, the embodiments of the present disclosure also support a delete operation for a fifth video script in the video script list. In an optional implementation, during the process of displaying the video script in the video script list, when a delete operation for a fifth video script in the video script list is received, the fifth video script is deleted from the video script list.

[0129] The fifth video script is any video script in the video script list. The delete operation for the fifth video script can include, if it is detected that the mouse hovers over the fifth video script in the video script list, a delete control is displayed in the input box of the fifth video script, when the delete control is triggered, the fifth video script is removed from the video script list.

[0130] By supporting the editing, copying, target interaction operation, and delete operation of the video script in the video script list, the video processing related interaction operation is further enriched, and the user experience is improved.

[0131] After triggering the editing, copying, target interaction operation, and delete operation of the video script in the video script list, a video draft can be generated based on the video script in the video script list and the media material. After obtaining the video draft, in order to enrich the related interaction mode of video processing, the embodiments of the present disclosure also support previewing the video draft.

[0132] In an optional implementation, when a triggering operation for the video draft control is received, at least one video draft is generated based on the media material and the target video script, and a video preview page is displayed, the video preview page displays a video draft list, information of a corresponding video draft in a display area of the video draft list, and a target video script corresponding to the video draft, and a first video draft in the video draft list is automatically played in a preview window of the video preview page. The first video draft can be any video draft in the video draft list. That is, while the video preview page is displayed, the video drafts in the video draft list are automatically cycled in the preview window of the video preview page.

[0133] The display area included in the video preview page has a corresponding relationship with the video drafts. The displayed information of the video drafts in the display area can include cover information and duration information of the video drafts, and the cover information of the video drafts can be a screenshot of a highlight segment in the media material input by the user.

[0134] As shown in FIG. 5, a schematic diagram of a video preview page provided by an embodiment of the present disclosure is shown. When a triggering operation for a video draft control 501 is received, a video preview page is displayed, the video preview page displays a video draft list 502, cover information, duration information of a corresponding video draft, and a target video script corresponding to the video draft in a display area 503 of the video draft list, and the video drafts displayed in the video draft list have a corresponding relationship with the target video script. In addition, a preview window 504 is also displayed in the video preview page, and a video draft is being played in the preview window.

[0135] In addition, the embodiment of the present disclosure can also support switching to play a video draft. In an optional implementation, when a switching-to-play triggering operation for a second video draft in the video draft list is received, the first video draft displayed in the preview window is switched to display the second video draft. The second video draft is any video draft in the video draft list except the first video draft. The switching-to-play triggering operation for the second video draft can include a click, double-click, or other triggering operation for the second video draft.

[0136] In actual application, based on the media material input by the user and the target video script, a video draft is generated, and the video draft is displayed on the video preview page. Based on the target video script and the media material, more video drafts are triggered to be generated to provide more video drafts to the user. That is, the video drafts are continuously generated based on the same target video script. The number of video clips included in each video draft corresponding to the target video script can be the same or different, and each video clip can correspond to the same or different script segment. The video duration of each generated video draft can be different.

[0137] In an optional implementation, a generate more drafts control is arranged on the video preview page. When the generate more drafts control is triggered, the relevance between the video script and the video clip in the video draft is scored and sorted based on the video script segmented from the target video script and the video clip cropped from the media material, and the video draft is generated based on the video clip and the video script with high ranking in the order, and displayed in the video draft list.

[0138] In actual application, after the video draft in the video draft list is displayed on the video preview page, the editing operation and the export operation of at least one video draft are supported to facilitate the user to edit and publish. Specifically, when the editing trigger operation of at least one video draft is received, the video draft is displayed on the video editing page, and the video editing information of the video draft is received. When the export operation of the video draft is received, the result video corresponding to the video draft is generated based on the video editing information.

[0139] The editing trigger operation of the video draft can include displaying an editing control on the video preview page. Specifically, when it is detected that the mouse hovers over one video draft in the video draft list, the editing control corresponding to the video draft is displayed. When the user triggers the editing control, the video editing page is displayed, and the video editing operation is displayed on the video editing page. When the video editing information of the video draft is received, the editing operation is performed on the video draft. Specifically, the editing operation can be performed on each video clip in the video draft respectively.

[0140] The video editing operation can include a music editing operation and other lightweight editing operations. The video editing page specifically, for example, a lightweight editing page.

[0141] In another optional implementation, when the user triggers the editing control on the video preview page, the video editing page is displayed, and the video track, the audio track, etc. are displayed on the video editing page. The video editing page specifically, for example, a multi-track video editing page. Through the multi-track video editing page, more editing operations can be performed on the video draft to meet the user's needs.

[0142] In actual applications, when the user triggers the editing control on the video draft page, an editing page with light editing operations can be first displayed to facilitate the user to perform preliminary editing operations on the video draft. Then, the user can trigger the editing control on the editing page with light editing operations to display a video editing page with multiple tracks to further edit the video draft, thereby meeting the user's needs.

[0143] For the export operation on the video draft, specifically, when the mouse hovers over one of the video drafts, the export control corresponding to the video draft is displayed on the video preview page. In an optional embodiment, when a trigger operation on the export control is received, a corresponding result video is generated based on the video draft.

[0144] In another optional embodiment, the export operation can also be triggered for the edited video draft. Specifically, the export control is displayed on the video editing page. When a trigger operation on the export control on the video editing page is received, the edited video draft displayed on the video editing page is used to generate a corresponding result video.

[0145] Based on the above method embodiments, the present disclosure further provides a video generation device. Referring to FIG. 6, which is a structural schematic diagram of a video generation device according to an embodiment of the present disclosure, the device comprises:

[0146] The first acquisition module 601 is configured to acquire the media material input by the user.

[0147] The first determination module 602 is configured to determine a target video script from a video script list. The video script list displays a video script generated based on the multimedia description content input by the user and / or a video script obtained by editing the video script generated based on the multimedia description content input by the user.

[0148] The first generation module 603 is configured to generate at least one video draft based on the media material input by the user and the target video script.

[0149] The video draft and the target video script have a corresponding relationship. The video draft includes a plurality of video segments. The plurality of video segments and a plurality of script segments have a corresponding relationship. The plurality of script segments are obtained by text segmentation of the target video script corresponding to the video draft. The plurality of video segments are extracted from the media material input by the user for the plurality of script segments. The display order of the plurality of video segments in the video draft is determined based on the preset order relationship between the corresponding script segments.

[0150] In an alternative implementation, the apparatus further includes:

[0151] a second obtaining module configured to obtain multimedia description content in a target text edit box, wherein the multimedia description content comprises first text content extracted from a target audio / video resource and / or second text content input based on the target text edit box;

[0152] a second generating module configured to generate at least one video script based on the multimedia description content and display the at least one video script in the video script list.

[0153] In an alternative implementation, the apparatus further includes:

[0154] a first receiving module configured to receive multimedia description content input for at least one target attribute of a target object;

[0155] a third generating module configured to generate at least one video script based on the multimedia description content corresponding to the at least one target attribute of the target object respectively and display the at least one video script in the video script list.

[0156] In an alternative implementation, the first receiving module includes:

[0157] a first display sub-module configured to display at least one candidate recommended content corresponding to a second target attribute of the target object in response to the multimedia description content input for a first target attribute of the target object, wherein the at least one candidate recommended content is determined based on the multimedia description content;

[0158] a first receiving sub-module configured to receive recommended content selected for the second target attribute from the at least one candidate recommended content;

[0159] Correspondingly, the third generating module is specifically configured to:

[0160] generate at least one video script based on the multimedia description content corresponding to the first target attribute and the recommended content selected for the second target attribute.

[0161] In an alternative implementation, the apparatus further includes:

[0162] an edit module configured to display a text edit box corresponding to a first video script in the video script list in response to an edit triggering operation for the first video script;

[0163] a receiving edit module configured to receive a text edit operation for the first video script based on the text edit box to obtain an edited video script;

[0164] The first updating module is configured to update the first video script in the video script list by using the edited video script.

[0165] In an optional implementation, the apparatus further includes:

[0166] The copying module is configured to, in response to a copying operation on a second video script in the video script list, add a duplicate video script displaying the second video script in the video script list.

[0167] The second updating module is configured to, in response to a text adjustment operation on the duplicate video script, generate a third video script and update the duplicate video script in the video script list by using the third video script.

[0168] In an optional implementation, the apparatus further includes:

[0169] The interaction module is configured to, in response to a target interaction operation triggered on a fourth video script in the video script list, display an interaction window.

[0170] The feedback module is configured to receive feedback information on the fourth video script based on the interaction window, and upload the feedback information to a server.

[0171] In an optional implementation, the apparatus further includes:

[0172] The deleting module is configured to, in response to a deleting operation triggered on a fifth video script in the video script list, delete the fifth video script from the video script list.

[0173] In an optional implementation, the apparatus further includes:

[0174] The previewing module is configured to display a video draft list on a video preview page, wherein the video draft list includes a display area corresponding to each of the at least one video draft, and the display area is configured to display information of the corresponding video draft and a target video script corresponding to the video draft.

[0175] The playing module is configured to play a first video draft in the video draft list on the video preview page.

[0176] In an optional implementation, the first generating module includes:

[0177] The text dividing module is configured to divide the target video script into a plurality of script segments, wherein the plurality of script segments have a preset order relationship.

[0178] a clipping module configured to clip a corresponding video segment for each of the plurality of script segments from the media material input by the user;

[0179] a splicing module configured to splice the video segments corresponding to the plurality of script segments based on the preset sequence relationship to obtain a video draft corresponding to the target video script.

[0180] In the video generation method provided by the embodiments of the present disclosure, first, the media material input by the user is obtained, and a target video script is determined from a video script list, wherein the video script list displays a video script generated based on multimedia description content input by the user and / or a video script obtained after editing the video script generated based on the multimedia description content input by the user. Then, at least one video draft is generated based on the media material input by the user and the target video script, wherein the video draft has a corresponding relationship with the target video script, the video draft includes a plurality of video segments, the video segment has a corresponding relationship with a script segment, the script segment is obtained by text segmentation of the target video script corresponding to the video draft, the plurality of video segments are extracted from the media material input by the user for the plurality of script segments, and the display order of the plurality of video segments in the video draft is determined based on a preset sequence relationship between the corresponding script segments.

[0181] The embodiments of the present disclosure support the function of generating a video draft based on a target video script generated based on multimedia description content input by a user and media material input by the user, enrich the video generation method, and thus meet the diversified video generation function needs of the user.

[0182] In addition to the above method and device, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores instructions, and when the instructions are run on a terminal device, the terminal device implements the video generation method provided by the embodiments of the present disclosure.

[0183] The embodiments of the present disclosure further provide a computer program product, which includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the video generation method provided by the embodiments of the present disclosure is implemented.

[0184] In addition, the embodiments of the present disclosure further provide a video generation device, as shown in FIG. 7, which can include:

[0185] The processor 701, the memory 702, the input device 703, and the output device 704. The number of processors 701 in the video generation device can be one or more, and one processor is taken as an example in FIG. 7. In some embodiments of the present disclosure, the processor 701, the memory 702, the input device 703, and the output device 704 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 7.

[0186] The memory 702 can be used to store software programs and modules, and the processor 701 can execute various functional applications and data processing of the video generation device by running the software programs and modules stored in the memory 702. The memory 702 can mainly include a program storage area and a data storage area, and the program storage area can store an operating system, at least one application program required by a function, and the like. In addition, the memory 702 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. The input device 703 can be used to receive input digital or character information, and generate signal input related to user settings and function control of the video generation device.

[0187] Specifically in the present embodiment, the processor 701 loads the executable file corresponding to the process of one or more application programs into the memory 702 according to the following instructions, and runs the application program stored in the memory 702 by the processor 701, thereby realizing the various functions of the video generation device described above.

[0188] It should be noted that in this paper, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0189] The foregoing is merely illustrative of the various implementations of the present disclosure and the general principles thereof. Numerous modifications can be made to these illustrations, and equivalents can be substituted therefor, without departing from the scope of the present disclosure. The specific embodiments commensurate with the specific application are intended to be illustrative only and not limiting of the scope of the application as set forth in the following claims.

Claims

1. A method for generating a video, wherein the method comprises: obtaining a media material inputted by a user; and determining a target video script from a video script list, wherein the video script list displays video scripts generated based on multimedia description content inputted by the user, and / or video scripts obtained by editing the video scripts generated based on the multimedia description content inputted by the user; generating at least one video draft based on the media material inputted by the user and the target video script; wherein the video draft has a corresponding relationship with the target video script, the video draft comprises a plurality of video clips, the plurality of video clips have a corresponding relationship with a plurality of script clips, the plurality of script clips are obtained by text segmentation of the target video script corresponding to the video draft, the plurality of video clips are extracted from the media material inputted by the user for the plurality of script clips respectively, and a display order of the plurality of video clips in the video draft is determined based on a preset order relationship between the corresponding script clips.

2. The method of claim 1, wherein before the determining the target video script from the video script list, the method further comprises: obtaining multimedia description content in a target text editing box, wherein the multimedia description content comprises first text content extracted from a target audio / video resource, and / or second text content inputted based on the target text editing box; generating at least one video script based on the multimedia description content, and displaying the at least one video script in the video script list.

3. The method of claim 1, wherein before the determining the target video script from the video script list, the method further comprises: receiving multimedia description content inputted for at least one target attribute of a target object; generating at least one video script based on multimedia description content corresponding to the at least one target attribute respectively, and displaying the at least one video script in the video script list.

4. The method of claim 3, wherein the receiving multimedia description content inputted for at least one target attribute of a target object comprises: in response to multimedia description content inputted for a first target attribute of a target object, displaying at least one candidate recommended content corresponding to a second target attribute of the target object, wherein the at least one candidate recommended content is determined based on the multimedia description content; receiving recommended content selected for the second target attribute from the at least one candidate recommended content; correspondingly, the generating at least one video script based on multimedia description content corresponding to the at least one target attribute respectively comprises: generating at least one video script based on multimedia description content corresponding to the first target attribute and the recommended content selected for the second target attribute.

5. The method of claim 1, wherein before the determining the target video script from the video script list, the method further comprises: in response to an editing trigger operation for a first video script in the video script list, displaying a text editing box corresponding to the first video script. receive a text editing operation for the first video script based on the text editing box, to obtain an edited video script; update the first video script in the video script list by using the edited video script.

6. The method of claim 1, wherein before the generating at least one video draft based on the user input media material and the target video script, the method further comprises: in response to a copy operation for a second video script in the video script list, adding a copy video script displaying the second video script in the video script list; in response to a text adjustment operation for the copy video script, generating a third video script, and updating the copy video script in the video script list by using the third video script.

7. The method of claim 1, wherein before the generating at least one video draft based on the user input media material and the target video script, the method further comprises: in response to a target interaction operation triggered for a fourth video script in the video script list, displaying an interaction window; receive feedback information for the fourth video script based on the interaction window, and upload the feedback information to a server.

8. The method of claim 1, wherein before the generating at least one video draft based on the user input media material and the target video script, the method further comprises: in response to a deletion operation triggered for a fifth video script in the video script list, deleting the fifth video script from the video script list.

9. The method of claim 1, wherein after the generating at least one video draft based on the user input media material and the target video script, the method further comprises: displaying a video draft list on a video preview page; wherein the video draft list includes a display area corresponding to each of the at least one video draft, and the display area is used to display information of the corresponding video draft and a target video script corresponding to the video draft; and playing a first video draft in the video draft list on the video preview page.

10. The method of claim 1, wherein the generating at least one video draft based on the user input media material and the target video script comprises: performing text segmentation on the target video script to obtain a plurality of script segments; wherein the plurality of script segments have a preset order relationship; cutting out a corresponding video segment for each of the plurality of script segments from the user input media material; based on the preset order relationship, performing splicing processing on the video segments corresponding to the plurality of script segments to obtain a video draft corresponding to the target video script.

11. A video generation apparatus, wherein the apparatus comprises: a first obtaining module configured to obtain user input media material; The first determining module is configured to determine a target video script from a video script list, wherein the video script list displays video scripts generated based on multimedia description content input by a user and / or video scripts obtained by editing the video scripts generated based on the multimedia description content input by the user; The first generating module is configured to generate at least one video draft based on the media material input by the user and the target video script; The video draft has a corresponding relationship with the target video script, the video draft includes a plurality of video clips, the plurality of video clips have a corresponding relationship with a plurality of script clips, the plurality of script clips are obtained by performing text segmentation on the target video script corresponding to the video draft, the plurality of video clips are extracted from the media material input by the user for the plurality of script clips respectively, and a display order of the plurality of video clips in the video draft is determined based on a preset order relationship between the corresponding script clips.

12. A computer readable storage medium, wherein the computer readable storage medium stores instructions, and the instructions, when executed on a terminal device, cause the terminal device to implement the method according to any one of claims 1-10.

13. A video generating apparatus comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1-10 when executing the computer program.

14. A computer program product, wherein the computer program product comprises computer programs / instructions, and the computer programs / instructions, when executed on a processor, implement the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Video editing method and device, electronic equipment and medium

    CN116489466A

  • Video generation method and apparatus, and electronic device

    CN116527994A

  • Video automatic generation method and device based on character originality

    CN117082293A

  • Video generation method and device, electronic equipment and readable storage medium

    CN118214921A

  • Contextualized Video Segment Selection For Video-Filled Text

    US20210110164A1