Video generation method and device, equipment and storage medium

By obtaining the media materials and video subject details input by the user, a video draft is generated, which solves the problem of the single video generation method in the existing technology, realizes diversified video generation methods, and improves the user experience.

CN121603737APending Publication Date: 2026-03-03BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411132781.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet users' diverse video generation needs and lack a variety of video generation methods.

Method used

By obtaining the object details of the media materials and video theme objects input by the user, a video draft is generated. The video draft contains multiple video clips and text clips. The text clips are extracted from the media materials based on the object details of the video theme objects, and the video clips are determined from the media materials input by the user.

Benefits of technology

It has enriched the ways to generate videos, met the diverse video generation needs of users, and improved the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121603737A_ABST
    Figure CN121603737A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, equipment and a storage medium, and the method comprises the steps: obtaining a media material inputted by a user, receiving a video subject object inputted by the user, obtaining the object detail information of the video subject object from a target database, and generating a video draft based on the media material input by the user and the object detail information of the video subject object. According to the embodiment of the invention, after the video subject object input by the user is received, the text fragment is generated based on the object detail information of the video subject object acquired from the target data and the information extracted from the media material, and the video draft is generated based on the text fragment and the media material input by the user. According to the invention, video generation modes are enriched, so that diversified video generation function requirements of users are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and more particularly to a video generation method, apparatus, device, and storage medium. Background Technology

[0002] With the continuous development of video generation technology, video generation-related functions have become more diversified. For example, generating videos using video templates.

[0003] To meet people's diverse needs for video generation functions, how to further enrich video generation methods is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a video generation method, apparatus, device, and storage medium.

[0005] In a first aspect, this disclosure provides a video generation method, the method comprising:

[0006] Obtain media materials input by the user;

[0007] The system receives a video theme object input by the user and retrieves the object details of the video theme object from the target database; wherein, the video theme object is the object used to describe the video draft.

[0008] Based on the media materials input by the user and the object details information of the video theme object, the video draft is generated;

[0009] The video draft includes multiple video clips, which correspond to multiple text clips. The multiple text clips are generated based on at least one of the object details information of the video subject and information extracted from the media material. The multiple video clips are determined from the media material input by the user for the multiple text clips.

[0010] In one optional implementation, before generating the video draft based on the media material input by the user and the object details information of the video theme object, the method further includes:

[0011] Obtain the media description content input for the video subject object;

[0012] Accordingly, generating the video draft based on the media material input by the user and the object details information of the video theme object includes:

[0013] The video draft is generated based on the media material input by the user, the object details information of the video theme object, and the media description content; wherein, the text fragment is generated based on at least one of the object details information of the video theme object, information extracted from the media material, and the media description content.

[0014] In one optional implementation, the method for receiving the video topic object input by the user includes:

[0015] In response to the input of a target search term, display the candidate objects corresponding to the target search term;

[0016] In response to the selection operation for a target candidate object, the target candidate object is determined as a video subject object.

[0017] In one optional implementation, determining the target candidate object as a video topic object in response to a selection operation for a target candidate object includes:

[0018] In response to the selection operation for a target candidate object, a list of resource objects corresponding to the target candidate object is displayed; wherein, the list of resource objects includes the resource objects bound to the target candidate object;

[0019] In response to the selection operation for a first target resource object in the list of resource objects, the first target resource object is determined as a video theme object.

[0020] In an optional implementation, after determining the first target resource object as a video theme object in response to the selection operation for the first target resource object in the resource object list, the method further includes:

[0021] In response to an edit-triggered operation on the video subject object, a list of resource objects corresponding to the target candidate object is displayed;

[0022] In response to the selection operation of a second target resource object in the list of resource objects, the second target resource object is determined as a video theme object.

[0023] In one optional implementation, the method further includes:

[0024] In response to the absence of a candidate object corresponding to the target search term, a content addition control is displayed;

[0025] In response to a triggering operation that adds a control to the content, the target search term is determined as the object description information of a video theme object; wherein, the object description information of the video theme object is used to generate the text fragment.

[0026] In one optional implementation, obtaining the media description content input for the video subject object includes:

[0027] Display candidate descriptions determined based on the video topic object;

[0028] In response to the selection operation for the target candidate description content, the target candidate description content is determined as the media description content input for the video subject object.

[0029] In one optional implementation, before obtaining the media material input by the user, the method further includes:

[0030] Receives selection operations for the target script template;

[0031] Accordingly, generating the video draft based on the media material input by the user and the object details information of the video theme object includes:

[0032] The video draft is generated based on the media material input by the user, the target script template, and the object details information of the video theme object; wherein the text fragment is generated based on at least one of the object details information of the video theme object, the structural features of the target script template, and information extracted from the media material.

[0033] In one optional implementation, generating the video draft based on the media material input by the user and the object details information of the video theme object includes:

[0034] The video script is generated based on at least one of the object details information of the video subject and the information extracted from the media material;

[0035] The video script is segmented into multiple text fragments; wherein the multiple text fragments have a preset order relationship.

[0036] From the media materials input by the user, corresponding video segments are cut out from the multiple text fragments respectively;

[0037] Based on the preset order relationship, the video segments corresponding to the multiple text segments are spliced ​​together to obtain the video draft.

[0038] In one optional implementation, the step of trimming corresponding video segments from the multiple text fragments from the media material input by the user includes:

[0039] Perform highlight recognition on the media material input by the user to identify highlight segments in the media material;

[0040] The highlight segment is cropped from the media material, and the text segment corresponding to the highlight segment is determined from the plurality of text segments.

[0041] In an optional implementation, after cropping corresponding video segments from the multiple text fragments from the user-input media material, the method further includes:

[0042] An audio segment is generated based on the text segment for the corresponding video segment; wherein the audio segment and the video segment have the same playback time information;

[0043] Accordingly, the step of splicing the video segments corresponding to the multiple text segments based on the preset order relationship to obtain the video draft includes:

[0044] Based on the preset order relationship corresponding to the multiple text segments, the video segments containing the audio segments are spliced ​​together to obtain the video draft.

[0045] In one optional implementation, the target database is used to store the object details information of the video subject object in a structured manner.

[0046] In one optional implementation, the object details information of the video subject object includes information for displaying on the object details page of the video subject object.

[0047] Secondly, this disclosure provides a video generation apparatus, the apparatus comprising:

[0048] The first acquisition module is used to acquire media materials input by the user;

[0049] The first receiving module is used to receive a video theme object input by the user and obtain object details information of the video theme object from the target database; wherein, the video theme object is the object used to describe the video draft;

[0050] The first generation module is used to generate the video draft based on the media material input by the user and the object details information of the video theme object;

[0051] The video draft includes multiple video clips, which correspond to multiple text clips. The multiple text clips are generated based on at least one of the object details information of the video subject and information extracted from the media material. The multiple video clips are determined from the media material input by the user for the multiple text clips.

[0052] Thirdly, this disclosure provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to implement the above-described method.

[0053] Fourthly, this disclosure provides a video generation device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0054] Fifthly, this disclosure provides a computer program product comprising a computer program / instruction that, when executed by a processor, implements the method described above.

[0055] The technical solution provided in this disclosure has at least the following advantages compared with the prior art:

[0056] This disclosure provides a video generation method, which involves acquiring media material input by a user, receiving a video theme object input by the user, and obtaining object details information of the video theme object from a target database. The video theme object is the object described in the video draft. Based on the media material input by the user and the object details information of the video theme object, a video draft is generated. The video draft includes multiple video segments, which correspond to multiple text segments. The multiple text segments are generated based on at least one of the object details information of the video theme object and information extracted from the media material. The multiple video segments are determined separately from the multiple text segments in the media material input by the user.

[0057] After receiving a video theme object input by the user, the embodiments of this disclosure support generating text fragments based on object details of the video theme object obtained from target data and information extracted from media materials, as well as generating video drafts based on text fragments and media materials input by the user. It can be seen that this disclosure enriches the video generation methods, thereby meeting the diverse video generation function needs of users. Attached Figure Description

[0058] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0059] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 A flowchart of a video generation method provided in this disclosure embodiment;

[0061] Figure 2 This is a schematic diagram of a video generation page provided in an embodiment of this disclosure;

[0062] Figure 3 A schematic diagram of a search page provided in an embodiment of this disclosure;

[0063] Figure 4 A schematic diagram of another video generation page provided in an embodiment of this disclosure;

[0064] Figure 5 A schematic diagram of another video generation page provided in an embodiment of this disclosure;

[0065] Figure 6 A schematic diagram of another video generation page provided in an embodiment of this disclosure;

[0066] Figure 7 A schematic diagram of another video generation page provided in an embodiment of this disclosure;

[0067] Figure 8 This is a schematic diagram of the structure of a video generation device provided in an embodiment of the present disclosure;

[0068] Figure 9 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this disclosure. Detailed Implementation

[0069] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0070] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0071] With the continuous development of video generation technology, video generation-related functions have become more diversified. For example, generating videos using video templates.

[0072] To meet people's diverse needs for video generation functions, how to further enrich video generation methods is a technical problem that urgently needs to be solved.

[0073] To this end, this disclosure provides a video generation method, which involves acquiring media material input by a user, receiving a video theme object input by the user, and obtaining object details information of the video theme object from a target database. The video theme object is the object described in the video draft. Based on the media material input by the user and the object details information of the video theme object, a video draft is generated. The video draft includes multiple video segments, which correspond to multiple text segments. The multiple text segments are generated based on at least one of the object details information of the video theme object and information extracted from the media material. The multiple video segments are determined separately from the multiple text segments in the media material input by the user.

[0074] After receiving a video theme object input by the user, the embodiments of this disclosure support generating text fragments based on object details of the video theme object obtained from target data and information extracted from media materials, as well as generating video drafts based on text fragments and media materials input by the user. It can be seen that this disclosure enriches the video generation methods, thereby meeting the diverse video generation function needs of users.

[0075] Based on this, the present disclosure provides a video generation method, referring to... Figure 1 The flowchart below illustrates a video generation method provided in this embodiment of the present disclosure. The method specifically includes:

[0076] S101: Obtain media materials input by the user.

[0077] The video generation method provided in this disclosure can be applied to a client, such as a client deployed on a smartphone, a client deployed on a tablet, etc.

[0078] In this embodiment of the disclosure, the media material is one or more media materials uploaded by the user. The media material may include media materials imported from the user's photo album. Specifically, the media material may include audio media materials, image media materials, etc.

[0079] Methods for obtaining media materials input by the user can include obtaining media materials input by the user by triggering a material upload control.

[0080] S102: Receive the video theme object input by the user, and obtain the object details information of the video theme object from the target database.

[0081] The video subject object is the object described in the video draft.

[0082] In other words, the video theme object is the main descriptive object in the video draft. Specifically, the video theme object can be POI information, or more specifically, it can be a store object or a product object under a store object. For example, the video theme object could be a product object of that store.

[0083] In one optional implementation, the video's subject matter can be determined using a target search term. Specifically, when input for a target search term is received, candidate objects corresponding to that target search term are displayed. When a selection operation is received for a candidate object, that candidate object is determined as the video's subject matter. The target search term may include keywords related to the video's subject matter; specifically, the target search term may be store information. The candidate objects corresponding to the target search term may include search results based on the target search term.

[0084] The input operation for the target search term may include a trigger operation on the search box, displaying a search page with a search box for entering the target search term.

[0085] like Figure 2 The diagram shown is a schematic of a video generation page provided in an embodiment of this disclosure. The video generation page displays a search box 201. When the search box 201 is triggered, a search page is displayed. When input of a target search term 202 is received on the search page, candidate objects corresponding to the target search term, such as candidate object 203, are displayed.

[0086] The target candidate object is one of the candidate objects corresponding to the target search term; specifically, the target candidate object is the store object.

[0087] As mentioned above Figure 2 In the video generation page shown, when candidate object 203 is selected, the candidate object is taken as the target candidate object and the target candidate object is determined as the video theme object, as shown in 204 on the video generation page.

[0088] In practical applications, based on the above-mentioned selection of target candidate objects, if the target candidate object is bound to a resource object, then the resource object bound to the target candidate object can be determined as the video theme object. The resource object can be a product object under the store object.

[0089] In one optional implementation, when a selection operation for a target candidate object is received, if it is detected that the target candidate object is bound to a resource object, a list of resource objects corresponding to the target candidate object is displayed. When a selection operation for the first target resource object in the list of resource objects is received, the first target resource object is determined as the video theme object.

[0090] The resource object list displays the resource objects bound to the target candidate object. The first target resource object is one of the resource objects.

[0091] like Figure 3 The diagram shown is a schematic of a search page provided in an embodiment of this disclosure. The search page displays candidate objects corresponding to a target search term. When a selection operation is received for a target candidate object 301, a list of resource objects corresponding to the target candidate object is displayed, showing the resource objects bound to the target candidate object. When a selection operation is received for a first target resource object 302, the first target resource object is identified as a video topic object.

[0092] The operation of identifying the first target resource object as the video theme object may include setting a determination control. Specifically, when a trigger operation is received for the determination control, the first target resource object is identified as the video theme object.

[0093] After identifying the first target resource object as the video theme object, this embodiment further enhances the user experience by supporting edit-triggered operations on the video theme object. These edit-triggered operations are used to change the first target resource object under the target candidate object, i.e., to change the video theme object.

[0094] In one optional implementation, when an edit trigger operation is received for a video theme object, a list of resource objects corresponding to the target candidate object is displayed. When a selection operation is received for a second target resource object in the resource object list, the second target resource object is determined as the video theme object. The second target resource object is any resource object other than the first target resource object among the resource objects bound to the target candidate object.

[0095] The editing trigger operation for the video theme object can include setting a drop-down control. Specifically, when a click operation is received on the drop-down control for the video theme object, a list of resource objects corresponding to the target candidate object is displayed.

[0096] like Figure 4 The diagram shown is a schematic of a video generation page provided in an embodiment of this disclosure. The video generation page currently displays a video theme object 401 and a dropdown control 402. When the dropdown control 402 is triggered, a list of resource objects corresponding to the target candidate object is displayed. When the second target resource object 403 in the resource object list is selected and the confirmation control is triggered, the second target resource object is determined as the video theme object, as shown in diagram 404.

[0097] In practical applications, embodiments of this disclosure also support modifications to the target candidate object. In one application scenario, if the target candidate object is a video theme object, when an edit trigger operation is received for the video theme object (target candidate object), the search page for the target candidate object is displayed. When a modification operation is received for the target search term, the corresponding candidate object is displayed based on the modified target search term. When an operation is received to select one of the candidate objects from the above candidate objects, that candidate object is used as the video theme object.

[0098] In another application scenario, after determining the target candidate object and designating the first target resource object among the resource objects bound to the target candidate object as the video theme object (i.e., the first target resource object under the target candidate object is the video theme object), when an edit trigger operation is received for the video theme object (the first target resource object), a list of resource objects corresponding to the target candidate object is displayed. When a sliding operation is triggered for the list of resource objects, the list of resource objects slides to display the resource objects, and a search input box for the target search term is displayed during the sliding display. When a modification input operation for the target search term is received, a search is performed based on the modified target search term, and the candidate objects corresponding to the modified target search term are displayed to determine the target candidate object. Then, a new resource object is determined from the list of resource objects corresponding to the target candidate object to be used as the video theme object.

[0099] In this embodiment of the disclosure, based on the received video subject object, the object details information of the video subject object is obtained from the target database. The target database is used to store the object details information of the video subject object in a structured manner, that is, the target database is used to classify, store and manage the object details information of the video subject object in a structured manner.

[0100] This object details information is used to generate the video draft. This information can include the information displayed on the object details page of the video subject object. Specifically, this object details information can include the name, address, etc., from the object details page.

[0101] S103: Generate the video draft based on the media material input by the user and the object details information of the video theme object;

[0102] The video draft includes multiple video clips, which correspond to multiple text clips. The multiple text clips are generated based on at least one of the object details information of the video subject and information extracted from the media material. The multiple video clips are determined from the media material input by the user for the multiple text clips.

[0103] In this embodiment of the disclosure, the video draft can be an unedited initial video draft. To generate a video draft that meets the user's needs, multiple video drafts can be generated based on the media materials input by the user and the object details information of the video theme object for the user to choose from. For example, five video drafts can be generated.

[0104] Since a video draft is composed of multiple video clips, the video clips in each of the generated video drafts may be the same or different, the number of video clips in each video draft may be the same or different, and the video length of each video draft may be different.

[0105] The triggering operation for generating a video draft can include a triggering operation for the video generation control. Specifically, it can display the video generation control.

[0106] like Figure 5 The diagram shown is a schematic of a video generation page provided in an embodiment of this disclosure. The video generation page displays user-uploaded media materials, as shown in 501, a video theme object 502, and a video generation control 503. When a trigger operation is received on the video generation control 503, a video draft is generated, as shown in 504.

[0107] Furthermore, the video draft in this embodiment is composed of multiple video clips spliced ​​together, and the display order of the multiple video clips in a video draft is determined based on a preset order relationship between the corresponding multiple text clips. These multiple video clips are obtained by trimming media materials input by the user.

[0108] In one optional implementation, the acquisition of text fragments can be based on object details information of a video theme object. The object details information of the video theme object may include information displayed on the object details page of the video theme object. The object details information can be used to characterize the relevant features of the video theme object. Specifically, the video theme object may include the name information, attribute resource value information, evaluation information, etc. of the video theme object.

[0109] In one application scenario, if the video subject is a store object, the object details information of the video subject object may include the store name, store address, store environment, store reviews, etc.

[0110] In another application scenario, if the video subject is a product object under a store object, the object details information of the video subject can include the details information of the product object, specifically the product name information, product review information, product price information, product description information, product raw material information, etc.

[0111] In another alternative implementation, the text fragment can also be generated based on user-uploaded media materials. Specifically, the visual content of the media materials can be analyzed to generate the text fragment.

[0112] Optionally, the text snippet can be generated based on the object details of the video subject object, or it can be generated based on media materials uploaded by the user.

[0113] In one optional implementation, a video script is generated based on at least one of the information extracted from the media material, according to the object details information of the video subject object. Then, the generated video script is segmented into multiple text fragments. Subsequently, corresponding video fragments are cut from the multiple text fragments from the media material input by the user. Based on the preset order relationship between the multiple text fragments, the video fragments corresponding to the multiple text fragments are spliced ​​together to obtain a video draft.

[0114] Specifically, the video script can be segmented into multiple text segments based on a preset order relationship, which may include paragraph relationships and contextual relationships.

[0115] From the media material input by the user, the corresponding video segments are cut out from multiple text fragments. Specifically, highlight recognition can be performed on the media material input by the user to identify the highlight segments in the media material. Then, the highlight segments are cut out from the media material and the text segments corresponding to the highlight segments are identified from multiple text fragments.

[0116] In practical applications, since multiple highlight segments can be matched from media footage based on text fragments, the highlight segments with higher similarity can be prioritized and cropped from the media footage for splicing to generate a video draft.

[0117] In practical applications, after generating a video draft, this embodiment of the disclosure also supports previewing the video draft to meet user needs and further improve the user experience.

[0118] As mentioned above Figure 5 The video generation page shown above displays at least one video draft, including video draft 504, and a video clip corresponding to video draft 504 displayed in the video preview area 505 of the video generation page.

[0119] In the video generation method provided in this embodiment, media materials input by the user are obtained, a video theme object input by the user is received, and object details information of the video theme object is obtained from a target database. The video theme object is the object used to describe the video draft. Based on the media materials input by the user and the object details information of the video theme object, a video draft is generated. The video draft includes multiple video segments, and there is a correspondence between the multiple video segments and multiple text segments. The multiple text segments are generated based on at least one of the object details information of the video theme object and information extracted from the media materials. The multiple video segments are determined separately from the multiple text segments in the media materials input by the user.

[0120] After receiving a video theme object input by the user, the embodiments of this disclosure support generating text fragments based on object details of the video theme object obtained from target data and information extracted from media materials, as well as generating video drafts based on text fragments and media materials input by the user. It can be seen that this disclosure enriches the video generation methods, thereby meeting the diverse video generation function needs of users.

[0121] In practical applications, to enrich the ways videos are generated and further enhance the user experience, this embodiment also supports obtaining media description content input for a video subject object. This media description content is used to determine text fragments.

[0122] The media description can include recommended audience information for the video's subject. Specifically, if the video's subject is a store, the media description can include recommended audience information for that store, such as its selling points. If the video's subject is a product within a store, the media description can include recommended audience information for that product, such as its selling points.

[0123] Specifically, the process involves obtaining media description content input for a video subject object, generating a video draft based on the user-input media material, the object details of the video subject object, and the media description content of the video subject object. The text fragments contained in the video draft can be generated based on at least one of the object details of the video subject object, information extracted from the media material, and the media description content.

[0124] The media description content input for the video subject object can include intelligent generation based on the object information on the object details page of the video subject object, or it can include media description content manually input by the user.

[0125] In one optional implementation, after determining the video subject object, based on the object information on the object details page of the video subject object, the model intelligently analyzes and displays the candidate description content of the video subject object, and when a selection operation for the target candidate description content is received, the target candidate content is determined as the media description content of the video subject object.

[0126] In this embodiment of the disclosure, before displaying the candidate description content of the video subject object, a recommendation prompt message may also be displayed, which is used to indicate that the candidate description content is intelligently generated.

[0127] like Figure 6 The diagram illustrates a video generation page provided in an embodiment of this disclosure. When a video theme object 601 is determined, recommended prompt information 602 is displayed on the video generation page. When candidate description information for the video theme object is analyzed, the recommended prompt information is hidden, and candidate description content is displayed, such as 603 "delicious and flavorful", 604 "high-quality ingredients", etc. When a selection operation is received for the target candidate description content "high-quality ingredients", the target candidate content is determined as the media description content input for the video theme object.

[0128] In practical applications, when determining video subject objects based on target search terms, there may be cases where no corresponding candidate object exists for the target search term. Therefore, this embodiment of the disclosure also supports using the target search term as object description information for the video subject object.

[0129] In one optional implementation, when an input operation for a target search term is received, if no corresponding candidate object is found for the target search term, a content addition control is displayed. When a trigger operation for the content addition control is received, the target search term is determined as the object description information of the video theme object, and then a video draft is generated based on the object description information of the video theme object and the media material input by the user.

[0130] like Figure 7 The diagram shown is a schematic of a video generation page provided in an embodiment of this disclosure. The video generation page displays a search box 701. When the search box 701 is triggered, a search page is displayed. When an input of a target search term is received on the search page, if no corresponding candidate object is found for the target search term, a content addition control 702 is displayed. When a trigger operation is received for the content addition control, the target search term is determined as object description information for the video theme object, as shown in 703.

[0131] By using target search terms as object description information for video subject objects, we can enrich the ways of video generation and improve the user experience.

[0132] In practical applications, to further enhance the user experience, this embodiment of the disclosure also supports generating video drafts based on script templates corresponding to videos that the user is interested in.

[0133] In one optional implementation, a selection operation for a target script template is received, and a video draft is generated based on user-input media materials, the target script template, and object details information of a video theme object. The video draft contains video segments and text segments that correspond to each other; the text segments are generated based on at least one of the following: object details information of the video theme object, structural features of the target script template, and information extracted from the media materials.

[0134] The target script template is a script template that carries a video theme object. A target script template carrying a video theme object can refer to a script template used to promote that video theme object. The structural characteristics of the target script template include its layout and other features.

[0135] In practical applications, to further enhance the user experience, human voice readings can also be added to video drafts.

[0136] In one optional implementation, audio segments are generated for corresponding video segments based on multiple text segments. These audio segments and video segments share the same playback time information. Specifically, multiple text segments can be read aloud to generate audio segments corresponding to multiple video segments. These audio segments and video segments share the same playback time information, meaning the audio segments and video segments are synchronized.

[0137] For methods where audio and video segments have the same playback time information, this can include adjusting the playback time of the audio segment generated by human voice reading; specifically, slowing down or speeding up the human voice reading based on the playback time of the video segment.

[0138] Then, based on the preset order relationship corresponding to the multiple text fragments, the video fragments with audio fragments are spliced ​​together to obtain a video draft.

[0139] In practical applications, after generating a video draft, this embodiment of the disclosure also supports generating more video drafts for the video subject and user-input media materials to meet user needs and further improve the user experience.

[0140] Specifically, at least one video draft is displayed on the video generation page. When a video generation operation is received on the video generation page, at least one candidate video draft is generated based on the object details of the video subject object and the media material input by the user, and the candidate video draft is displayed on the video preview page.

[0141] The video drafts displayed on the video generation page can be sorted according to factors such as image relevance. The video generation operation on the video generation page can include setting a generation control; when this control is triggered, candidate video drafts are generated and displayed on the video generation page.

[0142] To display candidate video drafts on the video generation page, you can use a preset swipe gesture on the video generation page to display the candidate video drafts.

[0143] In practical applications, while displaying video drafts on the video preview page, editing and export operations on the video drafts are also supported, facilitating user editing and publishing. Specifically, when an edit trigger operation is received for at least one video draft, the video draft is displayed on the video editing page, and video editing information for the draft is received. When an export operation is received for the video draft, the resulting video corresponding to the draft is generated based on the video editing information.

[0144] The editing trigger operation for the video draft can include displaying editing controls on the video preview page. Specifically, it first receives a selection operation for at least one video draft. When the user triggers the editing control, a video editing page is displayed, showing each video segment in the video draft. The video editing page displays the video editing operation and the selected video segments from the video draft. When video editing information for the video draft is received, the editing operation is performed on the video draft. Specifically, the editing operation can be performed on each video segment in the video draft separately.

[0145] In one optional implementation, when a user selects a video draft and triggers the editing controls on the video preview page, a video editing page is displayed. The video editing page displays multiple video clips containing the video draft. The user can select one of these video clips as the video clip to be edited. When video editing information for the video clip to be edited is received, video editing operations are performed on that video clip. These video editing operations can include lightweight editing operations such as text editing, script editing, and music editing. Specifically, the video editing page is, for example, a lightweight editing page.

[0146] In another optional implementation, when the user triggers the editing controls on the video preview page, a video editing page is displayed. This video editing page displays video tracks, audio tracks, etc., and specifically, for example, a multi-track video editing page. This multi-track video editing page facilitates further editing of the video draft, meeting user needs.

[0147] In practical applications, when a user selects a video draft and triggers the editing controls on the video preview page, an editing page with lightweight editing operations can be displayed first, allowing the user to perform preliminary editing operations on the video draft. Then, the user can trigger the editing controls on the video editing page to display a multi-track video editing page for further editing of the video draft, thus meeting the user's needs.

[0148] For exporting video drafts, an export control can be displayed. Specifically, the export control can be displayed on the video editing page. In one optional implementation, when a trigger operation for the export control is received, the resulting video corresponding to the video draft is generated based on the video editing information.

[0149] In practical applications, an export control can also be displayed on the video preview page. In one optional implementation, while at least one video draft is displayed on the video preview page, a target video draft can be selected. When an operation on the target video triggers the export control, the resulting video corresponding to the target video draft is generated.

[0150] Based on the above method embodiments, this disclosure also provides a video generation apparatus, with reference to... Figure 8 This is a schematic diagram of a video generation device provided in an embodiment of the present disclosure. The device includes:

[0151] The first acquisition module 801 is used to acquire media materials input by the user;

[0152] The first receiving module 802 is used to receive a video theme object input by the user and obtain object details information of the video theme object from the target database; wherein, the video theme object is an object used to describe the video draft;

[0153] The first generation module 803 is used to generate the video draft based on the media material input by the user and the object details information of the video theme object;

[0154] The video draft includes multiple video clips, which correspond to multiple text clips. The multiple text clips are generated based on at least one of the object details information of the video subject and information extracted from the media material. The multiple video clips are determined from the media material input by the user for the multiple text clips.

[0155] In one optional embodiment, the apparatus further includes:

[0156] The second acquisition module is used to acquire the media description content input for the video theme object;

[0157] Accordingly, the first generation module is specifically used for:

[0158] The video draft is generated based on the media material input by the user, the object details information of the video theme object, and the media description content; wherein, the text fragment is generated based on at least one of the object details information of the video theme object, information extracted from the media material, and the media description content.

[0159] In one optional implementation, the first receiving module includes:

[0160] The first display submodule is used to display candidate objects corresponding to the target search term in response to the input of the target search term;

[0161] The first determination submodule is used to determine the target candidate object as a video theme object in response to the selection operation for the target candidate object.

[0162] In one optional implementation, the first determining submodule includes:

[0163] The second display submodule is used to display a list of resource objects corresponding to the target candidate object in response to the selection operation of the target candidate object; wherein the list of resource objects includes the resource objects bound to the target candidate object;

[0164] The second determination submodule is used to determine the first target resource object as a video theme object in response to the selection operation of the first target resource object in the resource object list.

[0165] In one optional embodiment, the apparatus further includes:

[0166] The first display module is used to display a list of resource objects corresponding to the target candidate object in response to an edit trigger operation on the video theme object;

[0167] The first determining module is configured to determine the second target resource object as a video theme object in response to a selection operation of a second target resource object in the resource object list.

[0168] In one optional embodiment, the apparatus further includes:

[0169] The second display module is used to add content controls in response to the failure to find a candidate object corresponding to the target search term;

[0170] The second determining module is used to determine the target search term as the object description information of a video theme object in response to a trigger operation that adds a control to the content; wherein the object description information of the video theme object is used to generate the text fragment.

[0171] In one optional implementation, the second acquisition module includes:

[0172] Display candidate descriptions determined based on the video topic object;

[0173] In response to the selection operation for the target candidate description content, the target candidate description content is determined as the media description content input for the video subject object.

[0174] In one optional embodiment, the apparatus further includes:

[0175] The second receiving module is used to receive the selection operation for the target script template;

[0176] Accordingly, the first generation module is specifically used for:

[0177] The video draft is generated based on the media material input by the user, the target script template, and the object details information of the video theme object; wherein the text fragment is generated based on at least one of the object details information of the video theme object, the structural features of the target script template, and information extracted from the media material.

[0178] In one optional implementation, the first generation module includes:

[0179] The first generation submodule is used to generate video script based on at least one of the object details information of the video theme object and the information extracted from the media material;

[0180] The segmentation module is used to segment the video text into multiple text segments; wherein the multiple text segments have a preset order relationship.

[0181] The cropping module is used to crop out corresponding video segments from the multiple text fragments from the media materials input by the user.

[0182] The splicing module is used to splice the video segments corresponding to the multiple text segments based on the preset order relationship to obtain the video draft.

[0183] In one optional implementation, the cropping module includes:

[0184] The highlight recognition module is used to perform highlight recognition on the media material input by the user and determine the highlight segments in the media material;

[0185] The third determining module is used to cut out the highlight segment from the media material and determine the text segment corresponding to the highlight segment from the plurality of text segments.

[0186] In one optional embodiment, the apparatus further includes:

[0187] The second generation module is used to generate an audio segment for the corresponding video segment based on the text segment; wherein the audio segment and the video segment have the same playback time information;

[0188] Accordingly, the splicing module is specifically used for:

[0189] The third determining module is used to splice video segments containing audio segments based on a preset order relationship corresponding to the multiple text segments to obtain the video draft.

[0190] In one optional implementation, the target database is used to store the object details information of the video subject object in a structured manner.

[0191] In one optional implementation, the object details information of the video subject object includes information for displaying on the object details page of the video subject object.

[0192] In the video generation apparatus provided in this embodiment, a video theme object input by a user is received, and object details information of the video theme object is obtained from a target database. The video theme object is the object described in the video draft. Based on the media material input by the user and the object details information of the video theme object, a video draft is generated. The video draft includes multiple video segments, and there is a correspondence between the multiple video segments and multiple text segments. The multiple text segments are generated based on at least one of the object details information of the video theme object and information extracted from the media material. The multiple video segments are determined separately from the multiple text segments in the media material input by the user.

[0193] After receiving a video theme object input by the user, the embodiments of this disclosure support generating text fragments based on object details of the video theme object obtained from target data and information extracted from media materials, as well as generating video drafts based on text fragments and media materials input by the user. It can be seen that this disclosure enriches the video generation methods, thereby meeting the diverse video generation function needs of users.

[0194] In addition to the methods and apparatus described above, this disclosure also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to implement the video generation method described in this disclosure.

[0195] This disclosure also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, implement the video generation method described in this disclosure.

[0196] In addition, this disclosure also provides a video generation device, see [link to relevant documentation]. Figure 9 As shown, it may include:

[0197] The video generation device includes a processor 901, a memory 902, an input device 903, and an output device 904. The number of processors 901 in the video generation device can be one or more. Figure 9 Taking a processor as an example. In some embodiments of this disclosure, the processor 901, memory 902, input device 903, and output device 904 can be connected via a bus or other means, wherein, Figure 9 Taking the example of a connection between China and Israel via a bus.

[0198] The memory 902 can be used to store software programs and modules. The processor 901 executes various functional applications and data processing of the video generation device by running the software programs and modules stored in the memory 902. The memory 902 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc. In addition, the memory 902 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The input device 903 can be used to receive input digital or character information, and to generate signal inputs related to user settings and function control of the video generation device.

[0199] Specifically in this embodiment, the processor 901 loads the executable files corresponding to the processes of one or more applications into the memory 902 according to the following instructions, and the processor 901 runs the applications stored in the memory 902, thereby realizing the various functions of the video generation device described above.

[0200] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0201] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, characterized in that, The method includes: Obtain media materials input by the user; The system receives a video theme object input by the user and retrieves the object details of the video theme object from the target database; wherein, the video theme object is the object used to describe the video draft. Based on the media materials input by the user and the object details information of the video theme object, the video draft is generated; The video draft includes multiple video clips, which correspond to multiple text clips. The multiple text clips are generated based on at least one of the object details information of the video subject and information extracted from the media material. The multiple video clips are determined from the media material input by the user for the multiple text clips.

2. The method according to claim 1, characterized in that, Before generating the video draft based on the media material input by the user and the object details information of the video theme object, the process also includes: Obtain the media description content input for the video subject object; Accordingly, generating the video draft based on the media material input by the user and the object details information of the video theme object includes: The video draft is generated based on the media material input by the user, the object details information of the video theme object, and the media description content; wherein, the text fragment is generated based on at least one of the object details information of the video theme object, information extracted from the media material, and the media description content.

3. The method according to claim 1, characterized in that, The video topic object received from user input includes: In response to the input of a target search term, display the candidate objects corresponding to the target search term; In response to the selection operation for a target candidate object, the target candidate object is determined as a video subject object.

4. The method according to claim 3, characterized in that, The step of determining the target candidate object as a video topic object in response to the selection operation for the target candidate object includes: In response to the selection operation for a target candidate object, a list of resource objects corresponding to the target candidate object is displayed; wherein, the list of resource objects includes the resource objects bound to the target candidate object; In response to the selection operation for a first target resource object in the list of resource objects, the first target resource object is determined as a video theme object.

5. The method according to claim 4, characterized in that, After determining the first target resource object as a video theme object in response to the selection operation for the first target resource object in the resource object list, the method further includes: In response to an edit-triggered operation on the video subject object, a list of resource objects corresponding to the target candidate object is displayed; In response to the selection operation of a second target resource object in the list of resource objects, the second target resource object is determined as a video theme object.

6. The method according to claim 3, characterized in that, The method further includes: In response to the absence of a candidate object corresponding to the target search term, a content addition control is displayed; In response to a triggering operation that adds a control to the content, the target search term is determined as the object description information of a video theme object; wherein, the object description information of the video theme object is used to generate the text fragment.

7. The method according to claim 2, characterized in that, The step of obtaining the media description content input for the video subject object includes: Display candidate descriptions determined based on the video topic object; In response to the selection operation for the target candidate description content, the target candidate description content is determined as the media description content input for the video subject object.

8. The method according to claim 1, characterized in that, Before obtaining the media material input by the user, the method further includes: Receives selection operations for the target script template; Accordingly, generating the video draft based on the media material input by the user and the object details information of the video theme object includes: The video draft is generated based on the media material input by the user, the target script template, and the object details information of the video theme object; wherein the text fragment is generated based on at least one of the object details information of the video theme object, the structural features of the target script template, and information extracted from the media material.

9. The method according to claim 1, characterized in that, The process of generating the video draft based on the media material input by the user and the object details information of the video theme object includes: The video script is generated based on at least one of the object details information of the video subject and the information extracted from the media material; The video script is segmented into multiple text fragments; wherein the multiple text fragments have a preset order relationship. From the media materials input by the user, corresponding video segments are cut out from the multiple text fragments respectively; Based on the preset order relationship, the video segments corresponding to the multiple text segments are spliced ​​together to obtain the video draft.

10. The method according to claim 9, characterized in that, The step of trimming corresponding video segments from the multiple text fragments from the media materials input by the user includes: Perform highlight recognition on the media material input by the user to identify highlight segments in the media material; The highlight segment is cropped from the media material, and the text segment corresponding to the highlight segment is determined from the plurality of text segments.

11. The method according to claim 9, characterized in that, After cropping corresponding video segments from the multiple text fragments from the media materials input by the user, the process further includes: An audio segment is generated based on the text segment for the corresponding video segment; wherein the audio segment and the video segment have the same playback time information; Accordingly, the step of splicing the video segments corresponding to the multiple text segments based on the preset order relationship to obtain the video draft includes: Based on the preset order relationship corresponding to the multiple text segments, the video segments containing the audio segments are spliced ​​together to obtain the video draft.

12. The method according to claim 1, characterized in that, The target database is used to store the object details information of the video subject object in a structured manner.

13. The method according to claim 12, characterized in that, The object details information of the video subject object includes information used to display on the object details page of the video subject object.

14. A video generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire media materials input by the user; The first receiving module is used to receive a video theme object input by the user and obtain object details information of the video theme object from the target database; wherein, the video theme object is the object used to describe the video draft; The first generation module is used to generate the video draft based on the media material input by the user and the object details information of the video theme object; The video draft includes multiple video clips, which correspond to multiple text clips. The multiple text clips are generated based on at least one of the object details information of the video subject and information extracted from the media material. The multiple video clips are determined from the media material input by the user for the multiple text clips.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method as described in any one of claims 1-13.

16. A video generation device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1-13.

17. A computer program product, characterized in that, The computer program product includes a computer program / instruction that, when executed by a processor, implements the method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Video generation method and device, equipment and storage medium

    CN117979088A

  • Video generation method and device, equipment and storage medium

    CN117998159A

  • Object searching method and device, computer equipment and storage medium

    CN118093792A