Video generation method, apparatus, device and storage medium

By obtaining multimedia description content and user input materials, a video draft is generated based on a preset sequence relationship, which solves the problem of a single video generation method in the existing technology, realizes diversified video generation methods, and improves user experience.

WO2025201137A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/083442
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-03-19
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In the existing technology, the video generation method is relatively simple and it is difficult to meet the diverse video generation function requirements of users.

Method used

By acquiring multimedia description content and multimedia materials input by the user, determining description content segments based on the multimedia description content, extracting corresponding video segments from the multimedia materials input by the user, and generating a video draft according to a preset sequence relationship.

Benefits of technology

It enriches the video generation methods, meets users' diverse video generation function needs, and improves the flexibility of video generation and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025083442_02102025_PF_FP_ABST
    Figure CN2025083442_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a video generation method, an apparatus, a device and a storage medium. The method comprises: acquiring multimedia description content; acquiring a multimedia material input by a user; and generating a video draft on the basis of the multimedia description content and the multimedia material which is input by the user, the video draft comprising a plurality of video clips, the plurality of video clips having a correspondence relationship with a plurality of description content segments, and the plurality of description content segments being determined on the basis of the multimedia description content. That is, by means of the description content segments determined on the basis of the multimedia description content, the video clips are extracted from the multimedia material input by the user, and then, on the basis of a preset sequence relationship between the description content segments, the video draft formed by the extracted video clips is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method, device, equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application number 202410370301.7, filed on March 28, 2024, entitled “A video generation method, device, equipment and storage medium”. The entire contents of that application are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of data processing, and in particular to a video generation method, apparatus, device, and storage medium. Background Art

[0004] With the continuous development of video generation technology, video generation related functions have become more diverse, for example, using video templates to generate videos. Summary of the Invention

[0005] In order to solve the above technical problems, the embodiments of the present disclosure provide a video generation method, apparatus, device and storage medium.

[0006] In a first aspect, the present disclosure provides a video generation method, the method comprising:

[0007] Acquiring multimedia description content; wherein the multimedia description content includes text description content and / or voice description content;

[0008] and, obtaining multimedia materials input by the user;

[0009] Generate a video draft based on the multimedia description content and the multimedia material input by the user;

[0010] Among them, the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0011] In an optional implementation, generating a video draft based on the multimedia description content and the multimedia material input by the user includes:

[0012] Determining a plurality of description content segments based on the multimedia description content; wherein the plurality of description content segments have a preset sequence relationship;

[0013] Cutting corresponding video segments for each of the plurality of description content segments from the multimedia material input by the user;

[0014] Based on the preset sequence relationship, the video segments corresponding to the multiple description content segments are spliced ​​together to obtain a video draft.

[0015] The determining of a plurality of description content segments based on the multimedia description content includes:

[0016] generating a video copy based on the multimedia content;

[0017] The video copy is split into multiple description content segments; wherein the multiple description content segments have a preset sequence relationship.

[0018] In an optional implementation, obtaining multimedia description content includes:

[0019] The text description content in the target text editing box is obtained; wherein the text description content includes a first text content extracted from the target audio and video resource, and / or a second text content input based on the target text editing box.

[0020] In an optional implementation, obtaining multimedia description content includes:

[0021] receiving text content input for at least one target attribute of a target object;

[0022] The multimedia description content is acquired based on the text content corresponding to the at least one target attribute.

[0023] In an optional implementation, the receiving text content input for at least one target attribute of the target object includes:

[0024] In response to a third text content input for a first target attribute of a target object, displaying at least one candidate recommendation content corresponding to a second target attribute of the target object; wherein the at least one candidate recommendation content is determined based on the third text content;

[0025] receiving a fourth text content selected from the at least one candidate recommended content, and / or a fifth text content input for the second target attribute;

[0026] Accordingly, the acquiring of multimedia description content based on the text content corresponding to the at least one target attribute includes:

[0027] Multimedia description content is acquired based on at least one of the third text content, the fourth text content, and the fifth text content.

[0028] In an optional embodiment, before splicing the video segments corresponding to the plurality of description content segments based on the preset sequence relationship to obtain a video draft, the method further includes:

[0029] Generating voice segments for corresponding video segments based on the multiple description content segments; wherein the voice segments and the video segments have the same play time information;

[0030] Accordingly, based on the preset sequence relationship, the video segments corresponding to the plurality of description content segments are spliced ​​together to obtain a video draft, including:

[0031] Based on the preset sequence relationship, the video segments with the voice segments are spliced ​​together to obtain a video draft.

[0032] In an optional implementation manner, after generating a video draft based on the multimedia description content and the multimedia material input by the user, the method further includes:

[0033] displaying the draft video on a video preview page;

[0034] In response to the video generation operation performed on the video preview page, at least one candidate video draft is generated based on the multimedia description content and the multimedia material input by the user, and the at least one candidate video draft is displayed on the video preview page.

[0035] In an optional implementation manner, after generating a video draft based on the multimedia description content and the multimedia material input by the user, the method further includes:

[0036] In response to an editing trigger operation on the video draft, displaying the video draft on a video editing page;

[0037] receiving video editing information for the video draft;

[0038] In response to the export operation on the video draft, a result video corresponding to the video draft is generated based on the video editing information.

[0039] In a second aspect, the present disclosure provides a video generation device, the device comprising:

[0040] A first acquisition module is configured to acquire multimedia description content; wherein the multimedia description content includes text description content and / or voice description content;

[0041] The second acquisition module is used to acquire multimedia materials input by the user;

[0042] A first generating module, configured to generate a video draft based on the multimedia description content and the multimedia material input by the user;

[0043] Among them, the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0044] In a third aspect, the present disclosure provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium. When the instructions are executed on a terminal device, the terminal device implements the above method.

[0045] In a fourth aspect, the present disclosure provides a video generation device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.

[0046] In a fifth aspect, the present disclosure provides a computer program product, which includes a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0048] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0049] FIG1 is a flow chart of a video generation method provided by an embodiment of the present disclosure;

[0050] FIG2 is a schematic diagram of a video generation page provided by an embodiment of the present disclosure;

[0051] FIG3 is a schematic diagram of a user album page provided by an embodiment of the present disclosure;

[0052] FIG4 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure;

[0053] FIG5 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure;

[0054] FIG6 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure;

[0055] FIG7 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure;

[0056] FIG8 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure;

[0057] FIG9 is a schematic diagram of a video preview page provided by an embodiment of the present disclosure;

[0058] FIG10 is a schematic diagram of a video editing page provided by an embodiment of the present disclosure;

[0059] FIG11 is a schematic diagram of another video editing page provided by an embodiment of the present disclosure;

[0060] FIG12 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure;

[0061] FIG13 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0062] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0063] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0064] With the continuous development of video generation technology, video generation related functions have become more diverse, for example, using video templates to generate videos.

[0065] In order to meet people's diverse needs for video generation functions, how to further enrich video generation methods is a technical problem that needs to be solved urgently.

[0066] To this end, an embodiment of the present disclosure provides a video generation method, which obtains multimedia description content, wherein the multimedia description content includes text description content and / or voice description content, and obtains multimedia materials input by a user, and generates a video draft based on the multimedia description content and the multimedia materials input by the user, wherein the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and the multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0067] The disclosed embodiments extract corresponding video segments from multimedia materials input by the user by determining description content segments based on the multimedia description content, and then generate a video draft composed of the extracted video segments based on a preset sequence relationship between the description content segments. As can be seen, the disclosed embodiments support generating a video draft by first determining description content segments and then extracting video segments from multimedia materials input by the user based on the description content segments, enriching the video generation method and thus meeting the diverse video generation function requirements of users.

[0068] Based on this, an embodiment of the present disclosure provides a video generation method. Referring to FIG1 , which is a flow chart of a video generation method provided by an embodiment of the present disclosure, the method specifically includes:

[0069] S101: Acquire multimedia description content.

[0070] The multimedia description content includes text description content and / or voice description content.

[0071] The video generation method provided by the embodiments of the present disclosure may be applied to a client. For example, the client may include a client deployed on a smart phone, a client deployed on a tablet computer, and the like.

[0072] In the embodiments of the present disclosure, the multimedia description content may include received text description content and may also include received voice description content. In other words, the multimedia description content is not limited to any specific format and may be in voice, text, or other formats. The multimedia description content typically includes a description of the main object in the draft video to be generated. For example, if the draft video to be generated is a marketing video for a specific target, the multimedia description content may include description content for that target. In fact, the specific description content included in the multimedia description content is not limited.

[0073] In actual applications, there are many ways to obtain multimedia description content. In an optional implementation, the multimedia description content can be obtained based on the target text editing box. Specifically, the text description content in the target text editing box is obtained. The text description content may include the first text content extracted based on the target audio and video resources. The target audio and video resources can be audio resources, video resources, etc. selected from the user's album, or audio resources, video resources, etc. captured based on the shooting page.

[0074] Among them, the first text content extracted from the target audio and video resources can specifically include: first extracting the voice data extracted from the target audio and video resources, and then performing voice recognition on the extracted voice data to identify the corresponding text content, and displaying the identified text content in the target text editing box. The method for obtaining the target audio resource can include: the user clicks the extraction control set on the page where the target text editing box is located to trigger the extraction of the text content corresponding to the voice data in the target audio and video resource. Specifically, the text content can be obtained by performing voice recognition on the target audio and video resource. The embodiment of the present disclosure does not impose any restrictions on the specific implementation method of extracting the first text content from the target audio and video resource.

[0075] As shown in Figure 2, a schematic diagram of a video generation page provided in an embodiment of the present disclosure is shown. An extraction control 201 is displayed on the video generation page for obtaining target audio and video resources. As shown in Figure 3, a schematic diagram of a user album page provided in an embodiment of the present disclosure is shown. When the user triggers the extraction control 201, the user album page is displayed, and a plurality of optional audio and video resources are displayed on the user album page. The user determines the target audio and video resource by triggering a selection operation for any one or more audio and video resources. As shown in Figure 4, another schematic diagram of a user album page provided in an embodiment of the present disclosure is shown. The user album page displays audio and video resources 401 in a selected state, i.e., the target audio and video resource. By triggering the extraction control 402 displayed on the user album page, the first text content is triggered to be extracted from the target audio resource.

[0076] In actual application, the extracted first text content can be displayed in the target text edit box, as shown in Figure 5, which is a schematic diagram of another video generation page provided by this public embodiment. The first text content is displayed in the target text edit box 501 displayed on the video generation page.

[0077] In another optional embodiment, the acquired multimedia description content may include second text content entered by the user in the target text edit box. Specifically, the second text content manually entered by the user in the target text edit box is used to form the multimedia description content. The second text content may include any text content, including short sentences, long sentences, paragraphs, keywords, and the like.

[0078] As shown in FIG2 , the target text editing box 202 is displayed in the video generation page. When the user triggers the target text editing box, the video generation page shown in FIG5 is displayed, and the user can manually enter the second text content in the target text editing box 501 .

[0079] In another optional embodiment, the obtained multimedia description content may include not only the first text content extracted from the target audio and video resources, but also the second text content entered by the user in the target text edit box. The order in which the first and second text contents are displayed in the target text edit box can be determined based on demand.

[0080] In one application scenario, first, based on the user triggering the target text editing box, the second text content manually entered by the user in the target text editing box is obtained. Then, based on the user triggering the extraction control, the target audio and video resources are determined, and the first text content is extracted from the target audio and video resources and displayed in the target text editing box.

[0081] In another application scenario, the user can first trigger the extraction control to determine the target audio and video resources, extract the first text content from the target audio and video resources, and then obtain the second text content manually entered by the user in the target text editing box based on the display of the first text content in the target text editing box.

[0082] In actual applications, the display positions of the first text content and the second text content in the target text editing box may be based on the current positioning position, for example, the position of the mouse cursor.

[0083] Based on the above embodiment, obtaining the text description content in the target text editing box includes extracting it from the target audio and video, or based on manual input by the user. For this reason, the embodiment of the present disclosure can also display the entry for inputting or extracting text content on the video generation page to facilitate users to understand how to obtain the text description content.

[0084] As shown in FIG. 2 , the video generation page displays an entry 203 for inputting or extracting text content.

[0085] In practical applications, in order to improve the efficiency of obtaining multimedia description content and further enhance user experience, the embodiments of the present disclosure also support obtaining multimedia description content in an intelligent recommendation manner.

[0086] Specifically, the system first receives textual content input for at least one target attribute of a target object, and then obtains multimedia description content based on the textual content corresponding to the at least one target attribute. The target object can be the central subject in the generated video, and the target attribute can be any attribute of the central subject, such as its name, features, advantages, price, applicable audience, promotional offers, video length, etc.

[0087] For receiving text content input for at least one target attribute of a target object, it can include when receiving third text content input for a first target attribute of the target object, determining at least one candidate recommended content based on the third text content, and the at least one candidate recommended content corresponds to the second target attribute of the above-mentioned target object.

[0088] That is, upon receiving the third text content of the first target attribute of the target object manually input by the user, candidate recommendation content for the second attribute of the target object is automatically displayed.

[0089] The first target attribute and the second target attribute are any two different attributes in the target attribute. The third text content can be any text content, such as idioms, keywords, etc. The candidate recommended content can be any recommended text content, such as idioms, words, etc.

[0090] In actual applications, before obtaining multimedia description content in an intelligent recommendation manner, a candidate recommendation object content identifier can also be displayed to prompt the user of the intelligent generation of candidate recommendation content function. Specifically, the intelligent display candidate recommendation object content identifier can be displayed in the input box corresponding to the second attribute.

[0091] Figure 6 shows a schematic diagram of a video generation page provided by an embodiment of the present disclosure. The video generation page displays a first target attribute 601. Upon receiving a third text input for the first target attribute, a candidate recommendation object content identifier 603 is displayed within the input box corresponding to the second attribute 602, indicating that candidate recommendation content for the second attribute is currently being intelligently generated. While at least one candidate recommendation content is being displayed, the candidate recommendation object content identifier is hidden.

[0092] FIG7 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure. The video generation page displays the attributes of a target object, including a first target attribute 701. When a user enters a third text content of the first attribute in an input box 702, at least one candidate recommendation content for a second attribute 703 is displayed, including candidate recommendation content 704.

[0093] On the basis of displaying at least one candidate recommended content based on the third text content, the fourth text content and / or the fifth text content can be received as the text content of the second attribute, and then multimedia content is obtained based on at least one of the third text content, the fourth text content and the fifth text content.

[0094] Among them, the fourth text content can be at least one candidate recommended content selected by the user from at least one candidate recommended content, and the fifth text content can be the text content of the second attribute manually input by the user. As shown in Figure 8, a schematic diagram of another video generation page provided in an embodiment of the present disclosure is shown. When a selection operation is received for the candidate recommended content 704 in the video generation page shown in Figure 7 as the fourth text content, the fourth text content is displayed in the input box corresponding to the second target attribute. If the fifth text content corresponding to the second target attribute is received, the fifth text content is displayed in the input box corresponding to the second attribute.

[0095] In actual applications, in order to prompt the user of the smart acquisition function, a smart acquisition entrance may also be displayed, such as the smart acquisition entrance 604 displayed in the video generation page shown in FIG6 .

[0096] S102: Acquire multimedia materials input by the user.

[0097] In the embodiment of the present disclosure, the multimedia material is one or more multimedia materials uploaded by the user, and the multimedia material may include audio media material and picture media material.

[0098] The method of obtaining multimedia materials input by the user may include obtaining the multimedia materials input by the user by triggering a material upload control.

[0099] As shown in FIG. 2 above, the material upload control 204 in the video generation page.

[0100] S103: Generate a video draft based on the multimedia description content and the multimedia material input by the user.

[0101] Among them, the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0102] In the embodiment of the present disclosure, the video draft may be an unedited video draft. In order to generate a video draft that meets the needs of the user, multiple video drafts may be generated for the user to choose from, for example, 5 video drafts may be generated.

[0103] Since a video draft is composed of multiple video clips, among the multiple video drafts generated, the video clips in each video draft may be the same or different, the number of video clips contained in each video draft may be the same or different, and the video length of each video draft may be different.

[0104] The triggering operation for generating a video draft may include a triggering operation for generating a video control. Specifically, a video generation control may be displayed, such as the video generation control 205 displayed on the video generation page shown in FIG2 .

[0105] In an optional implementation, multiple description content segments are determined based on the multimedia description content, wherein the multiple description content segments have a preset sequential relationship. According to the multimedia material input by the user, corresponding video segments are respectively cut out from the multimedia material for the multiple description content segments. Based on the preset sequential relationship, the video segments corresponding to the multiple description content segments are spliced ​​to obtain a video draft.

[0106] Among them, the description content segment can refer to a description text segment. Determining multiple description content segments based on the multimedia description content can include directly determining the acquired multimedia description content as a description content segment. The preset sequence relationship between multiple description content segments can be a preset sequence relationship between multimedia description contents. The preset sequence relationship can be specifically, for example, a paragraph relationship, a context relationship, etc.

[0107] Determining multiple description content segments based on the multimedia description content may also include intelligently generating a video text based on the acquired multimedia content, and then splitting the video text into multiple description content segments with a preset sequence relationship. Specifically, the multimedia description content is subjected to content expansion processing to obtain the expanded description content, and the expanded description content is used as the video text, and then the video text is split into multiple description content segments with a preset sequence relationship to improve the quality of the video draft. For the expansion processing of the multimedia description content, a model can be used for expansion, and the specific implementation method is not limited in any way in the embodiments disclosed herein.

[0108] Regarding cropping video segments corresponding to the description content segments from the multimedia material input by the user, wherein the multiple video segments may be highlight segments in the multimedia material, specifically, the multimedia material may be compressed, and then the compressed multimedia material may be subjected to image recognition by the server to determine highlight segments with higher picture quality, and then the highlight segments in the compressed multimedia material may be replaced with corresponding segments in the multimedia material input by the user, and the highlight segments may be cropped out.

[0109] Then, based on a preset sequence relationship between the multiple description content segments, the video segments corresponding to the multiple description content segments are spliced ​​together to obtain a video draft.

[0110] In practical applications, before the video clips corresponding to the multiple description content segments are spliced ​​together to generate a video draft, it is necessary to generate audio clips for the video clips corresponding to the multiple description content segments. Specifically, the multiple description content segments can be read aloud by a human voice to generate audio clips corresponding to the multiple video segments. The audio clips and the video segments have the same playback time information, that is, the audio clips are synchronized with the video segments.

[0111] The method of having the voice segment and the video segment have the same play time information may include adjusting the play time of the voice segment generated by human voice reading, specifically, slowing down or speeding up the speed of human voice reading according to the play time of the video segment.

[0112] In the video generation method provided by the embodiment of the present disclosure, multimedia description content is obtained, wherein the multimedia description content includes text description content and / or voice description content, and multimedia materials input by a user are obtained, and a video draft is generated based on the multimedia description content and the multimedia materials input by the user, wherein the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and the multiple description content clips, the multiple description content clips are determined based on the multimedia description content, the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips, and the display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0113] The disclosed embodiments extract corresponding video segments from multimedia materials input by the user by determining description content segments based on the multimedia description content, and then generate a video draft composed of the extracted video segments based on a preset sequence relationship between the description content segments. As can be seen, the disclosed embodiments support generating a video draft by first determining description content segments and then extracting video segments from multimedia materials input by the user based on the description content segments, enriching the video generation method and thus meeting the diverse video generation function requirements of users.

[0114] In actual applications, after generating a video draft, it also supports previewing the video draft, and supports generating more candidate video drafts based on the multimedia material description content and multimedia materials input by the user on the video preview page to meet user needs and further enhance user experience.

[0115] Specifically, at least one video draft is displayed on the video preview page. When a video generation operation is received on the video preview page, at least one candidate video draft is generated based on the acquired multimedia description content and the multimedia material input by the user, and the candidate video draft is displayed on the video preview page.

[0116] Among them, at least one video draft displayed on the video preview page can be displayed on the video preview page according to a sorting strategy such as picture relevance and compliance with current hot topics. The video generation operation on the video preview page can include setting a generation control on the video preview page to generate candidate video drafts and displaying the candidate video drafts on the video preview page.

[0117] Regarding displaying candidate video drafts on a video preset page, the candidate video drafts may be displayed through a preset sliding operation on the video preset page.

[0118] Figure 9 is a schematic diagram of a video preview page provided by an embodiment of the present disclosure. The video preview page displays at least one video draft, including a video draft 901 and a generation control 902. When the generation control 902 is triggered, at least one candidate video draft is generated based on the acquired multimedia description content and the multimedia material input by the user, and the candidate video draft is displayed on the video preview page.

[0119] In actual applications, while displaying a video draft on the video preview page, editing and exporting operations for at least one video draft are also supported to facilitate user editing and publishing. Specifically, when an edit trigger operation is received for at least one video draft, the video draft is displayed on the video editing page, and video editing information for the video draft is received. When an export operation is received for the video draft, a result video corresponding to the video draft is generated based on the video editing information.

[0120] The editing trigger operation for the video draft may include displaying an editing control on the video preview page. Specifically, first, a selected operation for at least one video draft is received. When the user triggers the editing control, a video editing page is displayed, and each video clip in the video draft is displayed on the video editing page. The video editing operation and multiple video clips in the selected video draft are displayed on the video editing page. When video editing information for the video draft is received, the video draft is edited. Specifically, the editing operation can be performed on each video clip in the video draft.

[0121] As shown in the video preview page in FIG9 , when the video draft 901 is selected and the editing control 903 is triggered, the video editing page is displayed, on which the various video clips in the video draft and various video editing operations are displayed.

[0122] In one optional embodiment, when a user selects a video draft and triggers an edit control on the video preview page, a video editing page is displayed. The video editing page displays multiple video clips of the video draft. The user can select one of the video clips as the video clip to be edited. When video editing information for the video clip to be edited is received, video editing operations are performed on the video clip to be edited. The video editing operations may include lightweight editing operations such as text editing, script editing, and music editing. The video editing page is specifically a lightweight editing page.

[0123] As shown in Figure 10, a schematic diagram of a video editing page provided by an embodiment of the present disclosure is shown. The video editing page displays multiple video clips of a video draft, as well as lightweight editing operations such as text editing operations, script editing operations, and music editing operations.

[0124] In another optional embodiment, when a user triggers an edit control on a video preview page, a video editing page is displayed, which displays video tracks, audio tracks, etc., and specifically, a multi-track video editing page. This multi-track video editing page facilitates more editing operations on the video draft to meet user needs.

[0125] As shown in Figure 11, it is a schematic diagram of another video editing page provided by an embodiment of the present disclosure. The video editing page displays video tracks, audio tracks, etc.

[0126] In actual applications, when the user selects a video draft to trigger the editing controls on the video preview page, an editing page with lightweight editing operations will be displayed first, allowing the user to perform preliminary editing operations on the video draft. Then the user can trigger the editing controls on the video editing page to display a video editing page with multiple tracks to further edit the video draft to meet user needs.

[0127] For the export operation of the video draft, the export control can be displayed. Specifically, the export control can be displayed on the video editing page. In an optional implementation, when a trigger operation for the export control is received, a result video corresponding to the video draft is generated based on the video editing information.

[0128] In practical applications, an export control can also be displayed on the video preview page. In an optional implementation, while displaying at least one video draft on the video preview page, a target video draft can be selected. When an export control is triggered for the target video, a result video corresponding to the target video draft is generated. This is shown in the export control 904 displayed on the video preview page shown in Figure 9 above.

[0129] Based on the above method embodiment, the present disclosure further provides a video generation device. Referring to FIG12 , which is a schematic structural diagram of a video generation device provided in an embodiment of the present disclosure, the device includes:

[0130] A first acquisition module 1201 is configured to acquire multimedia description content, wherein the multimedia description content includes text description content and / or voice description content;

[0131] The second acquisition module 1202 is used to acquire multimedia materials input by the user;

[0132] A first generating module 1203 is configured to generate a video draft based on the multimedia description content and the multimedia material input by the user;

[0133] Among them, the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0134] In an optional implementation manner, the first generating module includes:

[0135] A first determining submodule is configured to determine a plurality of description content segments based on the multimedia description content; wherein the plurality of description content segments have a preset sequence relationship;

[0136] A first cropping submodule is configured to crop corresponding video segments for each of the plurality of description content segments from the multimedia material input by the user;

[0137] The first splicing submodule is used to splice the video segments corresponding to the multiple description content segments based on the preset sequence relationship to obtain a video draft.

[0138] In an optional implementation, the first determining submodule includes:

[0139] A first generating submodule, configured to generate a video text based on the multimedia content;

[0140] The first splitting submodule is used to split the video text into multiple description content segments; wherein the multiple description content segments have a preset sequence relationship.

[0141] In an optional implementation manner, the first acquisition module is specifically configured to:

[0142] The text description content in the target text editing box is obtained; wherein the text description content includes a first text content extracted from the target audio and video resource, and / or a second text content input based on the target text editing box.

[0143] In an optional implementation, the first acquisition module includes:

[0144] A first receiving submodule is configured to receive text content inputted for at least one target attribute of a target object;

[0145] The first acquisition submodule is configured to acquire multimedia description content based on text content corresponding to the at least one target attribute.

[0146] In an optional implementation, the first receiving submodule includes:

[0147] A first display submodule is configured to, in response to a third text content inputted for a first target attribute of a target object, display at least one candidate recommendation content corresponding to a second target attribute of the target object; wherein the at least one candidate recommendation content is determined based on the third text content;

[0148] a second receiving submodule, configured to receive a fourth text content selected from the at least one candidate recommended content, and / or a fifth text content input for the second target attribute;

[0149] Accordingly, the first acquisition submodule is specifically configured to:

[0150] Multimedia description content is acquired based on at least one of the third text content, the fourth text content, and the fifth text content.

[0151] In an optional embodiment, the device further includes:

[0152] A second generating module is configured to generate voice segments for corresponding video segments based on the plurality of description content segments; wherein the voice segments and the video segments have the same play time information;

[0153] Accordingly, the first splicing submodule is specifically used to:

[0154] Based on the preset sequence relationship, the video segments with the voice segments are spliced ​​together to obtain a video draft.

[0155] In an optional embodiment, the device further includes:

[0156] A first display module, configured to display the video draft on a video preview page;

[0157] The third generation module is used to generate at least one candidate video draft in response to the video generation operation performed on the video preview page, based on the multimedia description content and the multimedia material input by the user, and display the at least one candidate video draft on the video preview page.

[0158] In an optional embodiment, the device further includes:

[0159] A second display module, configured to display the video draft on a video editing page in response to an editing trigger operation on the video draft;

[0160] A first receiving module is configured to receive video editing information for the video draft;

[0161] The fourth generating module is used to generate a result video corresponding to the video draft based on the video editing information in response to the export operation on the video draft.

[0162] In the video generation method provided by the embodiment of the present disclosure, multimedia description content is obtained, wherein the multimedia description content includes text description content and / or voice description content, and multimedia materials input by a user are obtained, and a video draft is generated based on the multimedia description content and the multimedia materials input by the user, wherein the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and the multiple description content clips, the multiple description content clips are determined based on the multimedia description content, the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips, and the display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

[0163] The disclosed embodiments extract corresponding video segments from multimedia materials input by the user by determining description content segments based on the multimedia description content, and then generate a video draft composed of the extracted video segments based on a preset sequence relationship between the description content segments. As can be seen, the disclosed embodiments support generating a video draft by first determining description content segments and then extracting video segments from multimedia materials input by the user based on the description content segments, enriching the video generation method and thus meeting the diverse video generation function requirements of users.

[0164] In addition to the above-mentioned method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal device, the terminal device implements the video generation method described in the embodiments of the present disclosure.

[0165] The embodiments of the present disclosure further provide a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the video generation method described in the embodiments of the present disclosure is implemented.

[0166] In addition, the embodiment of the present disclosure further provides a video generation device, as shown in FIG13 , which may include:

[0167] Processor 1301, memory 1302, input device 1303, and output device 1304. The video generation device may include one or more processors 1301, with one processor being used as an example in FIG13 . In some embodiments of the present disclosure, processor 1301, memory 1302, input device 1303, and output device 1304 may be connected via a bus or other means, with FIG13 using a bus as an example.

[0168] Memory 1302 can be used to store software programs and modules. Processor 1301 executes the various functional applications and data processing of the video generation device by running the software programs and modules stored in memory 1302. Memory 1302 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, and the like. Furthermore, memory 1302 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Input device 1303 may be used to receive input digital or character information and generate signal input related to user settings and function control of the video generation device.

[0169] Specifically in this embodiment, the processor 1301 will load the executable files corresponding to the processes of one or more applications into the memory 1302 according to the following instructions, and the processor 1301 will run the applications stored in the memory 1302, thereby realizing the various functions of the above-mentioned video generation device.

[0170] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0171] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, comprising: Acquiring multimedia description content; wherein the multimedia description content includes text description content and / or voice description content; and, obtaining multimedia materials input by the user; Generate a video draft based on the multimedia description content and the multimedia material input by the user; Among them, the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

2. The method according to claim 1, wherein The step of generating a video draft based on the multimedia description content and the multimedia material input by the user includes: Determining a plurality of description content segments based on the multimedia description content; wherein the plurality of description content segments have a preset sequence relationship; Cutting corresponding video segments for each of the plurality of description content segments from the multimedia material input by the user; Based on the preset sequence relationship, the video segments corresponding to the multiple description content segments are spliced ​​together to obtain a video draft.

3. The method according to claim 2, wherein: The determining of a plurality of description content segments based on the multimedia description content includes: generating a video copy based on the multimedia content; The video copy is split into multiple description content segments; wherein the multiple description content segments have a preset sequence relationship.

4. The method according to claim 1, wherein The obtaining of multimedia description content includes: The text description content in the target text editing box is obtained; wherein the text description content includes a first text content extracted from the target audio and video resource, and / or a second text content input based on the target text editing box.

5. The method according to claim 1, wherein The obtaining of multimedia description content includes: receiving text content input for at least one target attribute of a target object; The multimedia description content is acquired based on the text content corresponding to the at least one target attribute.

6. The method according to claim 5, wherein: The receiving text content input for at least one target attribute of the target object includes: In response to a third text content input for a first target attribute of a target object, displaying at least one candidate recommendation content corresponding to a second target attribute of the target object; wherein the at least one candidate recommendation content is determined based on the third text content; receiving a fourth text content selected from the at least one candidate recommended content, and / or a fifth text content input for the second target attribute; Accordingly, the acquiring of multimedia description content based on the text content corresponding to the at least one target attribute includes: Multimedia description content is acquired based on at least one of the third text content, the fourth text content, and the fifth text content.

7. The method according to claim 2 or 3, wherein: Before the video clips corresponding to the plurality of description content clips are spliced ​​together based on the preset sequence relationship to obtain a video draft, the method further includes: Generating voice segments for corresponding video segments based on the multiple description content segments; wherein the voice segments and the video segments have the same play time information; Accordingly, based on the preset sequence relationship, the video segments corresponding to the plurality of description content segments are spliced ​​together to obtain a video draft, including: Based on the preset sequence relationship, the video segments with the voice segments are spliced ​​together to obtain a video draft.

8. The method according to claim 1, wherein After generating a video draft based on the multimedia description content and the multimedia material input by the user, the method further includes: displaying the draft video on a video preview page; In response to the video generation operation performed on the video preview page, at least one candidate video draft is generated based on the multimedia description content and the multimedia material input by the user, and the at least one candidate video draft is displayed on the video preview page.

9. The method according to claim 1, wherein After generating a video draft based on the multimedia description content and the multimedia material input by the user, the method further includes: In response to an editing trigger operation on the video draft, displaying the video draft on a video editing page; receiving video editing information for the video draft; In response to the export operation on the video draft, a result video corresponding to the video draft is generated based on the video editing information.

10. A video generation device, comprising: A first acquisition module is configured to acquire multimedia description content; wherein the multimedia description content includes text description content and / or voice description content; The second acquisition module is used to acquire multimedia materials input by the user; A first generating module, configured to generate a video draft based on the multimedia description content and the multimedia material input by the user; Among them, the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips. The multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively. The display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.

11. A computer-readable storage medium storing instructions, which, when executed on a terminal device, enable the terminal device to implement the method according to any one of claims 1 to 9.

12. A video generating device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Multimedia data processing method and device, electronic equipment and storage medium

    CN109756751A

  • Video generation method and device, equipment, medium and product

    CN114501064A

  • Video data generation method and device, equipment and storage medium

    CN116320605A

  • Video generation method and device, computer equipment and storage medium

    CN117082304A

  • System and method for generating video

    WO2021259322A1