Video generation method and apparatus, device, and storage medium
By obtaining the media materials and video subject details input by the user, a video draft is generated, which solves the problem of the single video generation method in the existing technology, realizes diversified video generation methods, and improves the user experience.
Patent Information
- Application Number
- PCT/CN2025/108449
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-16
- Filing Date
- 2025-07-14
- Publication Date
- 2026-02-19
AI Technical Summary
Existing video generation technologies are insufficient to meet users' diverse video generation needs and lack a variety of generation methods.
By obtaining the object details of the media materials and video theme objects input by the user, a video draft is generated. The video draft includes multiple video clips and text clips. The text clips are extracted from the media materials based on the object details of the video theme objects, and there is a correspondence between the video clips and the text clips.
It has enriched the ways to generate videos, met the diverse video generation needs of users, and improved the user experience.
Smart Images

Figure CN2025108449_19022026_PF_FP_ABST
Abstract
Description
Video generation method, device, apparatus and storage medium
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202411132781.X, filed on August 16, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0003] The present disclosure relates to a video generation method, device, apparatus and storage medium. BACKGROUND
[0004] With the continuous development of video generation technology, video generation related functions are also more diversified. For example, generating a video by using a video template.
[0005] In order to meet people's diversified video generation function needs, how to further enrich the video generation method is a technical problem to be solved at present. SUMMARY
[0006] In order to solve the above technical problems, the embodiments of the present disclosure provide a video generation method, device, apparatus and storage medium.
[0007] In a first aspect, the present disclosure provides a video generation method, comprising:
[0008] obtaining media material input by a user;
[0009] receiving a video theme object input by a user, and obtaining object detail information of the video theme object from a target database; wherein the video theme object is an object described by a video draft;
[0010] generating the video draft based on the media material input by the user and the object detail information of the video theme object;
[0011] wherein the video draft includes a plurality of video segments, the plurality of video segments and a plurality of text segments have a corresponding relationship, the plurality of text segments are generated based on at least one of the object detail information of the video theme object and information extracted from the media material, and the plurality of video segments are determined for the plurality of text segments from the media material input by the user.
[0012] In an optional implementation, before the generating the video draft based on the media material input by the user and the object detail information of the video theme object, the method further comprises:
[0013] obtaining media description content input for the video theme object;
[0014] Correspondingly, the generating the video draft based on the media material input by the user and the object detail information of the video theme object comprises:
[0015] The generating the video draft based on the media material input by the user, the object detail information of the video theme object and the media description content, wherein the text segment is generated based on at least one of the object detail information of the video theme object, information extracted from the media material and the media description content.
[0016] In an optional implementation, the receiving the video theme object input by the user comprises:
[0017] In response to input of a target search word, displaying candidate objects corresponding to the target search word;
[0018] In response to a selection operation on a target candidate object, determining the target candidate object as the video theme object.
[0019] In an optional implementation, the determining the target candidate object as the video theme object in response to the selection operation on the target candidate object comprises:
[0020] In response to the selection operation on the target candidate object, displaying a resource object list corresponding to the target candidate object, wherein the resource object list comprises resource objects bound to the target candidate object;
[0021] In response to a selection operation on a first target resource object in the resource object list, determining the first target resource object as the video theme object.
[0022] In an optional implementation, after the determining the first target resource object as the video theme object in response to the selection operation on the first target resource object in the resource object list, the method further comprises:
[0023] In response to an editing trigger operation on the video theme object, displaying the resource object list corresponding to the target candidate object;
[0024] In response to a selection operation on a second target resource object in the resource object list, determining the second target resource object as the video theme object.
[0025] In an optional implementation, the method further comprises:
[0026] In response to failure to search for candidate objects corresponding to the target search word, displaying a content adding control.
[0027] In response to a triggering operation on the content adding control, the target search word is determined as object description information of a video subject object; wherein the object description information of the video subject object is used to generate the text segment.
[0028] In an optional implementation, the obtaining of the media description content input for the video subject object comprises:
[0029] displaying candidate description content determined based on the video subject object;
[0030] In response to a selection operation on a target candidate description content, the target candidate description content is determined as the media description content input for the video subject object.
[0031] In an optional implementation, before the obtaining of the media material input by the user, the method further comprises:
[0032] receiving a selection operation on a target script template;
[0033] Correspondingly, the generating of the video draft based on the media material input by the user and the object detail information of the video subject object comprises:
[0034] generating the video draft based on the media material input by the user, the target script template and the object detail information of the video subject object; wherein the text segment is generated based on at least one of the object detail information of the video subject object, the structural feature of the target script template and information extracted from the media material.
[0035] In an optional implementation, the generating of the video draft based on the media material input by the user and the object detail information of the video subject object comprises:
[0036] generating a video script based on at least one of the object detail information of the video subject object and information extracted from the media material;
[0037] performing text segmentation on the video script to obtain a plurality of text segments; wherein the plurality of text segments have a preset order relationship.
[0038] cutting out corresponding video segments for the plurality of text segments from the media material input by the user;
[0039] based on the preset order relationship, performing splicing processing on the video segments corresponding to the plurality of text segments respectively to obtain the video draft.
[0040] In an alternative implementation, the cutting out of the corresponding video clip for each of the plurality of text clips from the user-input media material comprises:
[0041] performing highlight identification on the user-input media material to determine a highlight clip in the media material;
[0042] cutting out the highlight clip from the media material and determining a text clip corresponding to the highlight clip from the plurality of text clips.
[0043] In an alternative implementation, after the cutting out of the corresponding video clip for each of the plurality of text clips from the user-input media material, the method further comprises:
[0044] generating a voice clip for the corresponding video clip based on the text clip; wherein the voice clip and the video clip have the same play time information;
[0045] Correspondingly, the splicing of the video clips corresponding to the plurality of text clips based on the preset sequence relationship to obtain the video draft comprises:
[0046] splicing the video clip having the voice clip based on the preset sequence relationship of the plurality of text clips to obtain the video draft.
[0047] In an alternative implementation, the target database is configured to store the object detail information of the video theme object in a data structure.
[0048] In an alternative implementation, the object detail information of the video theme object comprises information to be displayed on an object detail page of the video theme object.
[0049] In a second aspect, the present disclosure provides a video generation apparatus, the apparatus comprising:
[0050] a first obtaining module configured to obtain user-input media material;
[0051] a first receiving module configured to receive a user-input video theme object and obtain object detail information of the video theme object from a target database; wherein the video theme object is an object described by a video draft;
[0052] a first generating module configured to generate the video draft based on the user-input media material and the object detail information of the video theme object;
[0053] The video draft includes a plurality of video clips, the plurality of video clips have a corresponding relationship with a plurality of text clips, the plurality of text clips are generated based on at least one of object detail information of the video theme object and information extracted from the media material, and the plurality of video clips are determined for the plurality of text clips from the media material input by the user.
[0054] In a third aspect, the present disclosure provides a computer readable storage medium, the computer readable storage medium stores instructions, when the instructions are executed on a terminal device, the terminal device implements the method described above.
[0055] In a fourth aspect, the present disclosure provides a video generation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, when the processor executes the computer program, the method described above is implemented.
[0056] In a fifth aspect, the present disclosure provides a computer program product, the computer program product comprises computer programs / instructions, when the computer programs / instructions are executed by a processor, the method described above is implemented. BRIEF DESCRIPTION OF DRAWINGS
[0057] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, together with the description.
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed in the embodiment description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0059] FIG. 1 is a flowchart of a video generation method according to an embodiment of the present disclosure;
[0060] FIG. 2 is a schematic diagram of a video generation page according to an embodiment of the present disclosure;
[0061] FIG. 3 is a schematic diagram of a search page according to an embodiment of the present disclosure;
[0062] FIG. 4 is a schematic diagram of another video generation page according to an embodiment of the present disclosure;
[0063] FIG. 5 is a schematic diagram of another video generation page according to an embodiment of the present disclosure;
[0064] FIG. 6 is a schematic diagram of another video generation page according to an embodiment of the present disclosure;
[0065] FIG. 7 is a schematic diagram of another video generation page according to an embodiment of the present disclosure;
[0066] FIG. 8 is a structural schematic diagram of a video generation apparatus according to an embodiment of the present disclosure; and
[0067] FIG. 9 is a structural schematic diagram of a video generation device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0068] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0069] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments described in the specification are only some embodiments of the present disclosure, not all embodiments.
[0070] With the continuous development of video generation technology, video generation related functions are also more diversified. For example, generating a video by using a video template.
[0071] In order to meet people's diversified video generation function needs, how to further enrich the video generation method is a technical problem to be solved at present.
[0072] Therefore, an embodiment of the present disclosure provides a video generation method, media materials input by a user are acquired, a video theme object input by the user is received, and object detail information of the video theme object is acquired from a target database, wherein the video theme object is an object described by a video draft, a video draft is generated based on the media materials input by the user and the object detail information of the video theme object, wherein the video draft includes a plurality of video segments, the plurality of video segments and a plurality of text segments have a corresponding relationship, the plurality of text segments are generated based on at least one of the object detail information of the video theme object and information extracted from the media materials, and the plurality of video segments are determined for the plurality of text segments from the media materials input by the user.
[0073] After receiving the video theme object input by the user, an embodiment of the present disclosure supports generating the text segment based on the object detail information of the video theme object acquired from the target data and the information extracted from the media materials, and generating the video draft based on the text segment and the media materials input by the user. It can be seen that the present disclosure enriches the video generation method, thereby meeting the diversified video generation function needs of the user.
[0074] Based on this, the embodiment of the disclosure provides a video generation method, referring to FIG. 1, a flowchart of a video generation method provided by the embodiment of the disclosure, the method specifically includes:
[0075] S101: Obtain media material input by a user.
[0076] The video generation method provided by the embodiment of the disclosure can be applied to a client, for example, the client can include a client deployed on a smart phone, a client deployed on a tablet computer, etc.
[0077] In the embodiment of the disclosure, the media material is one or more media materials uploaded by a user, which can include media material imported from a user album, specifically, the media material can include audio media material, picture media material, etc.
[0078] For the way of obtaining the media material input by the user, the media material input by the user can be obtained in a way that the user triggers a material uploading control.
[0079] S102: Receive a video theme object input by a user, and obtain object detail information of the video theme object from a target database.
[0080] The video theme object is an object described in a video draft.
[0081] That is, the video theme object is the main description object in the video draft. Specifically, the video theme object can be POI information, specifically, the video theme object can be a store object, and can also be a commodity object under the store object, for example, the video theme object can be a commodity object of the store.
[0082] In an optional implementation, the video theme object can be determined through a target search term. Specifically, when receiving an input for a target search term, display candidate objects corresponding to the target search term, and when receiving a selection operation for a target candidate object, determine the target candidate object as the video theme object. The target search term can include a keyword related to the video theme object, specifically, the target search term can be store object information. The candidate objects corresponding to the target search term can include search results based on the target search term.
[0083] The input operation for the target search term can include a trigger operation for a search box, display a search page, and the search page displays a search box for inputting the target search term.
[0084] As shown in FIG. 2, a schematic diagram of a video generation page is provided according to an embodiment of the present disclosure. The video generation page displays a search box 201. When the search box 201 is triggered, a search page is displayed. When receiving an input of a target search term 202 acting on the search page, a candidate object corresponding to the target search term is displayed, such as a candidate object 203.
[0085] The target candidate object is one of the candidate objects corresponding to the target search term. Specifically, the target candidate object is a store object.
[0086] As shown in the video generation page of FIG. 2, when the candidate object 203 is selected, the candidate object is taken as the target candidate object. The target candidate object is determined as a video theme object, as shown by 204 in the video generation page.
[0087] In actual application, on the basis of the above selection of the target candidate object, if the target candidate object is bound with a resource object, the resource object bound with the target candidate object can be determined as the video theme object. The resource object can be a commodity object under the store object.
[0088] In an optional implementation, when receiving a selection operation on the target candidate object, if it is detected that the target candidate object is bound with a resource object, a resource object list corresponding to the target candidate object is displayed. When receiving a selection operation on a first target resource object in the resource object list, the first target resource object is determined as the video theme object.
[0089] The resource object list is used to display the resource object bound with the target candidate object. The first target resource object is one of the resource objects.
[0090] As shown in FIG. 3, a schematic diagram of a search page is provided according to an embodiment of the present disclosure. The search page displays a candidate object corresponding to a target search term. When receiving a selection operation on the target candidate object 301, a resource object list corresponding to the target candidate object is displayed. The resource object list displays a resource object bound with the target candidate object. When receiving a selection operation on a first target resource object 302, the first target resource object is determined as the video theme object.
[0091] The operation of determining the first target resource object as the video theme object can include setting a determination control. Specifically, when receiving a trigger operation on the determination control, the first target resource object is determined as the video theme object.
[0092] After the first target resource object is determined as the video theme object, the user experience is further improved, and the embodiment of the present disclosure also supports an edit trigger operation for the video theme object. The edit trigger operation is used to change the first target resource object under the target candidate object, that is, to replace the video theme object.
[0093] In an optional implementation, when the edit trigger operation for the video theme object is received, a resource object list corresponding to the target candidate object is displayed, and when a selection operation for a second target resource object in the resource object list is received, the second target resource object is determined as the video theme object. The second target resource object is a resource object in the resource object bound to the target candidate object except the first target resource object.
[0094] The edit trigger operation for the video theme object can include setting a drop-down control. Specifically, when a click operation for the drop-down control of the video theme object is received, the resource object list corresponding to the target candidate object is displayed.
[0095] As shown in FIG. 4, it is a schematic diagram of a video generation page provided by the embodiment of the present disclosure. The video generation page currently displays a video theme object 401 and a drop-down control 402. When the drop-down control 402 is triggered, a resource object list corresponding to the target candidate object is displayed. When a second target resource object 403 in the resource object list is selected and a determination control is triggered, the second target resource object is determined as the video theme object, as shown in 404.
[0096] In actual application, the embodiment of the present disclosure also supports changing the target candidate object. In one application scenario, if the target candidate object is a video theme object, when an edit trigger operation for the video theme object (target candidate object) is received, a search page of the target candidate object is displayed. When a modification operation for the target search term is received, the corresponding candidate object is displayed based on the modified target search term. When an operation of selecting one of the candidate objects from the above candidate objects is received, the candidate object is used as the video theme object.
[0097] In another application scenario, after determining the target candidate object and binding the first target resource object in the resource object of the target candidate object as the video theme object, i.e., the first target resource object under the target candidate object is the video theme object, when receiving an editing trigger operation on the video theme object (the first target resource object), displaying a resource object list corresponding to the target candidate object, when triggering a sliding operation on the resource object list, slidingly displaying the resource objects in the resource object list, and displaying a search input box of the target search word in the sliding display process, when receiving a modification input operation on the target search word, searching based on the modified target search word, and displaying the candidate objects corresponding to the modified target search word for determining the target candidate object, and further determining the resource objects from the resource object list corresponding to the target candidate object for being used as the video theme object.
[0098] In the embodiments of the present disclosure, based on the received video theme object, object detail information of the video theme object is obtained from a target database, wherein the target database is used for data structured storage of the object detail information of the video theme object, i.e., the object detail information of the video theme object is classified, structured stored and managed by using the target database.
[0099] The object detail information is used to generate a video draft, and the object detail information can include information displayed on an object detail page of the video theme object. Specifically, the object detail information can include name information, address information, etc. on the object detail page.
[0100] S103: generating the video draft based on the media material input by the user and the object detail information of the video theme object;
[0101] The video draft includes a plurality of video segments, and the plurality of video segments have a corresponding relationship with a plurality of text segments. The plurality of text segments are generated based on at least one of the information extracted from the media material and the object detail information of the video theme object. The plurality of video segments are determined from the media material input by the user for the plurality of text segments, respectively.
[0102] In the embodiments of the present disclosure, the video draft can be an unedited video draft. To generate a video draft meeting the user's demand, a plurality of video drafts can be generated based on the media material input by the user and the object detail information of the video theme object, so as to be selected by the user. For example, five video drafts are generated.
[0103] Since the video draft is spliced by multiple video clips, in the generated multiple video drafts, the video clips in each video draft can be the same or different, the number of video clips contained in each video draft can be the same or different, and the video duration of each video draft can be different.
[0104] The triggering operation for generating the video draft can include a triggering operation for the video generation control. Specifically, the video generation control can be displayed.
[0105] As shown in FIG. 5, a schematic diagram of a video generation page provided by an embodiment of the present disclosure is shown. The video generation page displays the media material uploaded by the user, the video theme object 502, and the video generation control 503, as shown in 501. When a triggering operation for the video generation control 503 is received, a video draft is generated, as shown in 504.
[0106] In addition, the video draft in the embodiment of the present disclosure is composed of multiple video clips, and the display order of the multiple video clips in a video draft is determined based on the preset order relationship between the corresponding multiple text clips. The multiple video clips are obtained by cropping from the media material input by the user.
[0107] In an optional implementation, the text clip can be generated based on the object detail information of the video theme object. The object detail information of the video theme object can include the information displayed on the object detail page of the video theme object. The object detail information can be used to represent the relevant features of the video theme object. The video theme object can specifically include the name information of the video theme object, the attribute resource value information, the evaluation information, and the like.
[0108] In an application scenario, if the video theme object is a store object, the object detail information of the video theme object can include the store name information, the store address information, the store environment information, the store evaluation information, and the like.
[0109] In another application scenario, if the video theme object is a commodity object under the store object, the object detail information of the video theme object can include the detail information of the commodity object, specifically the commodity name information, the commodity evaluation information, the commodity price information, the commodity description information, the commodity raw material information, and the like.
[0110] In another optional implementation, the text clip can also be generated based on the media material uploaded by the user. Specifically, the picture content in the media material can be analyzed to generate the text clip.
[0111] Optionally, the text segment can be generated based on the object detail information of the video theme object, or can be generated based on the user-uploaded media material.
[0112] In an optional implementation, the video script is generated based on at least one of the information extracted from the media material based on the object detail information of the video theme object, and then the generated video script is subjected to text segmentation to obtain a plurality of text segments, and then corresponding video segments are respectively cropped for the plurality of text segments from the user-input media material, the video segments corresponding to the plurality of text segments are subjected to splicing processing based on a preset order relationship between the plurality of text segments, and a video draft is obtained.
[0113] In the text segmentation of the video script, the video script can be subjected to text segmentation based on a preset order relationship to obtain a plurality of text segments, and the preset order relationship can include a paragraph relationship and a context relationship.
[0114] The corresponding video segments are respectively cropped for the plurality of text segments from the user-input media material, and specifically, the highlight recognition can be performed on the user-input media material to determine a highlight segment in the media material, and then the highlight segment is cropped from the media material, and a text segment corresponding to the highlight segment is determined from the plurality of text segments.
[0115] In actual application, since a plurality of highlight segments can be matched from the media material based on the text segments, the highlight segments with higher similarity that are matched can be preferentially cropped from the media material for splicing to generate a video draft.
[0116] In actual application, after the video draft is generated, the embodiments of the present disclosure also support previewing the video draft to meet the user demand and further improve the user experience.
[0117] As shown in the above FIG. 5, the video generation page is displayed. At least one video draft is displayed on the video generation page, including a video draft 504, and a video segment corresponding to the video draft 504 is displayed in a video preview area 505 of the video generation page.
[0118] The video generation method provided in the embodiments of the present disclosure includes: obtaining media material input by a user; receiving a video theme object input by the user; and obtaining object detail information of the video theme object from a target database, wherein the video theme object is an object described by a video draft, and the video draft is generated based on the media material input by the user and the object detail information of the video theme object, wherein the video draft includes a plurality of video segments, the plurality of video segments have a corresponding relationship with a plurality of text segments, the plurality of text segments are generated based on at least one of the object detail information of the video theme object and information extracted from the media material, and the plurality of video segments are determined for the plurality of text segments from the media material input by the user.
[0119] After receiving the video theme object input by the user, the embodiments of the present disclosure support generating the text segments based on the object detail information of the video theme object obtained from the target database and the information extracted from the media material, and generating the video draft based on the text segments and the media material input by the user. It can be seen that the present disclosure enriches the video generation method, thereby meeting the diversified video generation function needs of the user.
[0120] In actual application, in order to enrich the video generation method and further improve the user experience, the embodiments of the present disclosure also support obtaining media description content input for the video theme object. The media description content is used to determine the text segments.
[0121] The media description content can include recommended object information of the video theme object. Specifically, if the video theme object is a store object, the media description information can include recommended object information of the store object, such as a store selling point and the like. If the video theme object is a product object under the store, the media description content can include recommended object information of the product object under the store, such as a product selling point and the like.
[0122] Specifically, the media description content input for the video theme object is obtained, and the video draft is generated based on the media material input by the user, the object detail information of the video theme object and the media description content of the video theme object. The text segments contained in the video draft can be generated based on at least one of the object detail information of the video theme object, the information extracted from the media material and the media description content.
[0123] The media description content input for the video theme object can include object information generated intelligently based on an object detail page of the video theme object, and can also include media description content input manually by the user.
[0124] In an alternative implementation, after determining the video theme object, candidate description content of the video theme object is intelligently analyzed based on object information on an object detail page of the video theme object, and the candidate description content is displayed. When a selection operation for target candidate description content is received, the target candidate content is determined as the media description content of the video theme object.
[0125] Before displaying the candidate description content of the video theme object, the embodiment of the present disclosure can also display recommendation prompt information, which is used to prompt that the candidate description content is intelligently generated.
[0126] As shown in FIG. 6, it is a schematic diagram of a video generation page provided by the embodiment of the present disclosure. When the video theme object 601 is determined, the recommendation prompt information 602 is displayed on the video generation page. When the candidate description information of the video theme object is analyzed, the recommendation prompt information is hidden, and the candidate description content is displayed, such as “fresh and delicious” 603, “high-quality food materials” 604, etc. When a selection operation for the target candidate description content “high-quality food materials” is received, the target candidate content is determined as the media description content input for the video theme object.
[0127] In actual application, when the video theme object is determined based on the target search term, since there is a case that the target search term does not have a corresponding candidate object, the embodiment of the present disclosure also supports taking the target search term as the object description information of the video theme object.
[0128] In an alternative implementation, when an input operation for a target search term is received, if it is detected that no corresponding candidate object is searched for the target search term, a content addition control is displayed. When a trigger operation for the content addition control is received, the target search term is determined as the object description information of the video theme object, and then a video draft is generated based on the object description information of the video theme object and the media material input by the user.
[0129] As shown in FIG. 7, it is a schematic diagram of a video generation page provided by the embodiment of the present disclosure. The search box 701 is displayed on the video generation page. When the search box 701 is triggered, a search page is displayed. When a target search term is input on the search page, if it is detected that no corresponding candidate object is searched for the target search term, a content addition control 702 is displayed. When a trigger operation for the content addition control is received, the target search term is determined as the object description information of the video theme object, as shown in 703.
[0130] By taking the target search term as the object description information of the video theme object, the video generation mode can be enriched, and the user experience can be improved.
[0131] In actual application, to further improve user experience, the embodiment of the present disclosure also supports generating a video draft based on a script template corresponding to a video of interest of a user.
[0132] In an optional implementation, a selected operation for a target script template is received, and a video draft is generated based on the media material input by the user, the target script template, and the object detail information of the video theme object. The video draft includes video segments and text segments, and the video segments and the text segments have a corresponding relationship. The text segments are generated based on at least one of the object detail information of the video theme object, the structural features of the target script template, and information extracted from the media material.
[0133] The target script template is a script template carrying a video theme object. The target script template carrying the video theme object can refer to the target script template being a script template for promoting the video theme object. The structural features of the target script template include structural layout modes of the target script template.
[0134] In actual application, to further improve user experience, a human voice reading can also be added to the video draft.
[0135] In an optional implementation, a plurality of text segments are used to generate a plurality of voice segments corresponding to a plurality of video segments. The voice segments and the video segments have the same playback time information. Specifically, the plurality of text segments are used to generate the plurality of voice segments corresponding to the plurality of video segments by human voice reading. The voice segments and the video segments have the same playback time information, that is, the voice segments and the video segments are kept synchronous.
[0136] For the voice segments and the video segments having the same playback time information, the playback time of the voice segments generated by human voice reading can be adjusted. Specifically, the speed of the human voice reading is slowed down or accelerated according to the playback time of the video segments.
[0137] Then, the video segments with the voice segments are spliced based on a preset order relationship of the plurality of text segments, to obtain a video draft.
[0138] In actual application, after the video draft is generated, the embodiment of the present disclosure also supports generating more video drafts for the video theme object and the media material input by the user, to meet user demand and further improve user experience.
[0139] Specifically, at least one video draft is displayed on a video generation page. When a video generation operation acting on the video generation page is received, at least one candidate video draft is generated based on the object detail information of the video theme object and the media material input by the user, and the candidate video draft is displayed on the video preview page.
[0140] In the video generation page, the video drafts displayed can be ranked according to picture relevance and other ranking strategies. For the video generation operation on the video generation page, a generation control can be set on the video generation page, and when the generation control is triggered, a candidate video draft is generated and displayed on the video generation page.
[0141] For displaying the candidate video draft on the video generation page, the candidate video draft can be displayed through a preset sliding operation on the video generation page.
[0142] In actual application, in the process of displaying the video draft on the video preview page, editing operation and export operation on the video draft can also be supported to facilitate user editing and publishing. Specifically, when receiving an editing trigger operation on at least one video draft, the video draft is displayed on a video editing page, video editing information on the video draft is received, and when receiving an export operation on the video draft, a result video corresponding to the video draft is generated based on the video editing information.
[0143] Specifically, the editing trigger operation on the video draft can include displaying an editing control on the video preview page. Specifically, first, a selection operation on at least one video draft is received, and when the user triggers the editing control, a video editing page is displayed, each video segment in the video draft is displayed on the video editing page, the video editing operation and the multiple video segments in the selected video draft are displayed on the video editing page, and when receiving the video editing information on the video draft, the video editing operation is performed on the video draft. Specifically, the editing operation can be performed on each video segment in the video draft.
[0144] In an optional implementation, when the user selects the video draft and triggers the editing control on the video preview page, the video editing page is displayed, the multiple video segments of the video draft are displayed on the video editing page, the user can select one of the video segments as a video segment to be edited, and when receiving the video editing information on the video segment to be edited, the video editing operation is performed on the video segment to be edited. The video editing operation can include text editing operation, script editing operation, and music editing operation, and the like. The video editing page can be, for example, a lightweight editing page.
[0145] In another optional implementation, when the user triggers the editing control on the video preview page, the video editing page is displayed, and the video track, audio track, and the like are displayed on the video editing page. The video editing page can be, for example, a multi-track video editing page. Through the multi-track video editing page, more editing operations can be performed on the video draft to meet user needs.
[0146] In actual application, when the user selects a video draft to trigger the editing control on the video preview page, an editing page with light editing operation is first displayed to facilitate the user to perform preliminary editing operation on the video draft, and then the user can trigger the editing control on the video editing page to display the video editing page with multiple tracks to further edit the video draft, thereby meeting the user demand.
[0147] For the export operation on the video draft, an export control can be displayed. Specifically, the export control can be displayed on the video editing page. In an optional implementation, when a trigger operation on the export control is received, a result video corresponding to the video draft is generated based on the video editing information.
[0148] In actual application, the export control can also be displayed on the video preview page. In an optional implementation, during display of at least one video draft on the video preview page, a target video draft is selected, and when the export control is triggered for the target video, a result video corresponding to the target video draft is generated.
[0149] Based on the above method embodiments, the present disclosure further provides a video generation apparatus. Referring to FIG. 8, a structural schematic diagram of a video generation apparatus according to an embodiment of the present disclosure is shown. The apparatus includes:
[0150] The first obtaining module 801 is configured to obtain media material input by a user.
[0151] The first receiving module 802 is configured to receive a video subject object input by a user, and obtain object detail information of the video subject object from a target database. The video subject object is an object described by a video draft.
[0152] The first generating module 803 is configured to generate the video draft based on the media material input by the user and the object detail information of the video subject object.
[0153] The video draft includes a plurality of video segments, and the plurality of video segments have a corresponding relationship with a plurality of text segments. The plurality of text segments are generated based on at least one of the object detail information of the video subject object and information extracted from the media material. The plurality of video segments are determined from the media material input by the user for the plurality of text segments, respectively.
[0154] In an optional implementation, the apparatus further includes:
[0155] The second obtaining module is configured to obtain media description content input for the video subject object.
[0156] Correspondingly, the first generating module is specifically configured to:
[0157] generate the video draft based on the media material, the object detail information of the video theme object and the media description content; wherein the text segment is generated based on at least one of the object detail information of the video theme object, the information extracted from the media material and the media description content.
[0158] In an optional implementation, the first receiving module comprises:
[0159] The first display sub-module is configured to display candidate objects corresponding to a target search word in response to an input of the target search word.
[0160] The first determining sub-module is configured to determine a target candidate object as a video theme object in response to a selection operation on the target candidate object.
[0161] In an optional implementation, the first determining sub-module comprises:
[0162] The second display sub-module is configured to display a list of resource objects corresponding to a target candidate object in response to a selection operation on the target candidate object; wherein the list of resource objects comprises resource objects bound to the target candidate object.
[0163] The second determining sub-module is configured to determine a first target resource object in the list of resource objects as a video theme object in response to a selection operation on the first target resource object.
[0164] In an optional implementation, the apparatus further comprises:
[0165] The first display module is configured to display a list of resource objects corresponding to a target candidate object in response to an editing trigger operation on the video theme object.
[0166] The first determining module is configured to determine a second target resource object in the list of resource objects as a video theme object in response to a selection operation on the second target resource object.
[0167] In an optional implementation, the apparatus further comprises:
[0168] The second display module is configured to display a content adding control in response to a failure to search for candidate objects corresponding to a target search word.
[0169] The second determining module is configured to determine the target search word as object description information of a video theme object in response to a trigger operation on the content adding control; wherein the object description information of the video theme object is used to generate the text segment.
[0170] In an optional implementation, the second obtaining module comprises:
[0171] displaying the candidate description content determined based on the video theme object;
[0172] in response to a selection operation on a target candidate description content, determining the target candidate description content as the media description content input for the video theme object.
[0173] In an optional implementation, the apparatus further comprises:
[0174] a second receiving module configured to receive a selection operation on a target script template;
[0175] correspondingly, the first generating module is specifically configured to:
[0176] generate the video draft based on the media material input by the user, the target script template, and the object detail information of the video theme object; wherein the text segment is generated based on at least one of the object detail information of the video theme object, the structural feature of the target script template, and the information extracted from the media material.
[0177] In an optional implementation, the first generating module comprises:
[0178] a first generating submodule configured to generate a video script based on at least one of the object detail information of the video theme object and the information extracted from the media material;
[0179] a segmentation module configured to perform text segmentation on the video script to obtain a plurality of text segments; wherein the plurality of text segments have a preset order relationship;
[0180] a clipping module configured to clip a corresponding video segment for each of the plurality of text segments from the media material input by the user;
[0181] a splicing module configured to perform splicing processing on the video segments corresponding to the plurality of text segments based on the preset order relationship, to obtain the video draft.
[0182] In an optional implementation, the clipping module comprises:
[0183] a highlight identification module configured to perform highlight identification on the media material input by the user to determine a highlight segment in the media material;
[0184] a third determining module configured to clip the highlight segment from the media material, and determine a text segment corresponding to the highlight segment from the plurality of text segments.
[0185] In an optional implementation, the apparatus further includes:
[0186] a second generation module, configured to generate a voice segment for a corresponding video segment based on the text segment; wherein the voice segment has the same playing time information as the video segment;
[0187] Correspondingly, the splicing module is specifically configured to:
[0188] a third determination module, configured to splice the video segments with the voice segments based on the preset order relationship of the text segments, to obtain the video draft.
[0189] In an optional implementation, the target database is configured to store the object detail information of the video theme object in a data structure.
[0190] In an optional implementation, the object detail information of the video theme object includes information to be displayed on an object detail page of the video theme object.
[0191] In the video generation apparatus provided by the embodiments of the present disclosure, a video theme object input by a user is received, and object detail information of the video theme object is obtained from a target database, wherein the video theme object is an object described by a video draft, a video draft is generated based on media material input by the user and the object detail information of the video theme object, wherein the video draft includes a plurality of video segments, the plurality of video segments have a corresponding relationship with a plurality of text segments, the plurality of text segments are generated based on at least one of the object detail information of the video theme object and information extracted from the media material, and the plurality of video segments are determined from the media material input by the user for the plurality of text segments respectively.
[0192] After receiving the video theme object input by the user, the embodiments of the present disclosure support generating the text segments based on the information extracted from the media material and the object detail information of the video theme object obtained from the target database, and generating the video draft based on the text segments and the media material input by the user. It can be seen that the present disclosure enriches the video generation method, thereby meeting the diversified video generation function needs of the user.
[0193] In addition to the method and apparatus described above, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores instructions, and when the instructions run on a terminal device, the terminal device implements the video generation method provided by the embodiments of the present disclosure.
[0194] The embodiments of the present disclosure further provide a computer program product, which comprises computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the video generation method provided by the embodiments of the present disclosure.
[0195] In addition, the embodiments of the present disclosure further provide a video generation device, as shown in FIG. 9, which can include:
[0196] The video generation device can include a processor 901, a memory 902, an input device 903, and an output device 904. The number of processors 901 in the video generation device can be one or more, and one processor is taken as an example in FIG. 9. In some embodiments of the present disclosure, the processor 901, the memory 902, the input device 903, and the output device 904 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 9.
[0197] The memory 902 can be used to store software programs and modules, and the processor 901 executes the software programs and modules stored in the memory 902 to perform various functional applications and data processing of the video generation device. The memory 902 can mainly include a program storage area and a data storage area, and the program storage area can store an operating system, at least one application program required by a function, and the like. In addition, the memory 902 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. The input device 903 can be used to receive input digital or character information, and generate signal input related to the user settings and function control of the video generation device.
[0198] Specifically, in the present embodiment, the processor 901 loads the executable file corresponding to the process of one or more application programs into the memory 902 according to the following instructions, and executes the application program stored in the memory 902 by the processor 901, thereby realizing the various functions of the video generation device described above.
[0199] It should be noted that, in this paper, the relationship terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a…" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0200] The foregoing detailed description has set forth various embodiments of the present disclosure via the use of specific terminology. However, embodiments thereof can be practiced without the specific details (e.g., quantities, values, etc.) set forth in the specification and drawings. The present disclosure can be practiced with or without the specific details set forth.
Claims
A video generation method comprises: obtaining media material input by a user; receiving a video subject object input by the user, and obtaining object detail information of the video subject object from a target database; wherein the video subject object is an object described by a video draft; generating the video draft based on the media material input by the user and the object detail information of the video subject object; wherein the video draft comprises a plurality of video segments, the plurality of video segments have a corresponding relationship with a plurality of text segments, the plurality of text segments are generated based on at least one of the object detail information of the video subject object and information extracted from the media material, and the plurality of video segments are determined from the media material input by the user for the plurality of text segments respectively. The method of claim 1, wherein, Before the generating the video draft based on the media material input by the user and the object detail information of the video subject object, the method further comprises: obtaining media description content input for the video subject object; correspondingly, the generating the video draft based on the media material input by the user and the object detail information of the video subject object comprises: generating the video draft based on the media material input by the user, the object detail information of the video subject object and the media description content; wherein the text segments are generated based on at least one of the object detail information of the video subject object, information extracted from the media material and the media description content. The method according to claim 1 or 2, wherein The receiving the video subject object input by the user comprises: in response to input of a target search term, displaying candidate objects corresponding to the target search term; in response to a selection operation for a target candidate object, determining the target candidate object as a video subject object. The method of claim 3, wherein, The determining the target candidate object as a video subject object in response to the selection operation for the target candidate object comprises: in response to the selection operation for the target candidate object, displaying a resource object list corresponding to the target candidate object; wherein the resource object list comprises resource objects bound to the target candidate object; in response to a selection operation for a first target resource object in the resource object list, determining the first target resource object as a video subject object. The method of claim 4, wherein, After the determining the first target resource object as a video subject object in response to the selection operation for the first target resource object in the resource object list, the method further comprises: in response to an edit trigger operation for the video subject object, displaying the resource object list corresponding to the target candidate object; in response to a selection operation for a second target resource object in the resource object list, determining the second target resource object as a video subject object. The method according to any one of claims 3-5, further comprising: in response to no candidate object corresponding to the target search term being searched, displaying a content adding control; In response to a triggering operation on the content adding control, the target search word is determined as object description information of the video subject object; wherein the object description information of the video subject object is used to generate the text segment. The method of claim 2, wherein, The media description content input for the video subject object is obtained, including: Displaying the candidate description content determined based on the video subject object; In response to a selection operation on the target candidate description content, the target candidate description content is determined as the media description content input for the video subject object. The method according to any one of claims 1 to 7, wherein Before the media material input by the user is obtained, the method further includes: Receiving a selection operation on the target script template; Correspondingly, the video draft is generated based on the media material input by the user and the object detail information of the video subject object, including: The video draft is generated based on the media material input by the user, the target script template and the object detail information of the video subject object; wherein the text segment is generated based on at least one of the object detail information of the video subject object, the structural features of the target script template and the information extracted from the media material. The method according to any one of claims 1 to 8, wherein The video draft is generated based on the media material input by the user and the object detail information of the video subject object, including: Video scripts are generated based on at least one of the object detail information of the video subject object and the information extracted from the media material; Text segmentation is performed on the video scripts to obtain a plurality of text segments; wherein the plurality of text segments have a preset order relationship; Corresponding video segments are respectively cropped from the media material input by the user for the plurality of text segments; Based on the preset order relationship, the video segments corresponding to the plurality of text segments are spliced to obtain the video draft. The method of claim 9, wherein, The corresponding video segments are respectively cropped from the media material input by the user for the plurality of text segments, including: High light recognition is performed on the media material input by the user to determine a high light segment in the media material; The high light segment is cropped from the media material, and a text segment corresponding to the high light segment is determined from the plurality of text segments. The method according to claim 9 or 10, wherein After the corresponding video segments are respectively cropped from the media material input by the user for the plurality of text segments, the method further includes: Voice segments are generated for the corresponding video segments based on the text segments; wherein the voice segments and the video segments have the same playback time information; Correspondingly, the video draft is generated based on the preset order relationship, including: Based on the preset order relationship of the plurality of text segments, the video segments with the voice segments are spliced to obtain the video draft. The method according to any one of claims 1-11, wherein, The target database is used to store the object detail information of the video subject object in a data structure. The method of claim 12, wherein, The object detail information of the video subject object includes information displayed on the object detail page of the video subject object. A video generation apparatus comprises: a first obtaining module configured to obtain media material input by a user; a first receiving module configured to receive a video theme object input by the user and obtain object detail information of the video theme object from a target database, wherein the video theme object is an object described by a video draft; a first generating module configured to generate the video draft based on the media material input by the user and the object detail information of the video theme object; wherein the video draft comprises a plurality of video segments, the plurality of video segments and a plurality of text segments have a corresponding relationship, the plurality of text segments are generated based on at least one of the object detail information of the video theme object and information extracted from the media material, and the plurality of video segments are determined from the media material input by the user for the plurality of text segments respectively. A computer-readable storage medium storing instructions, wherein, When the instructions run on a terminal device, the terminal device implements the video generation method according to any one of claims 1-13. A video generation device comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the video generation method according to any one of claims 1-13 is implemented. A computer program product comprising computer programs / instructions, wherein, The computer program / instructions implement the video generation method according to any one of claims 1-13 when executed by a processor.
Citation Information
Patent Citations
Video synthesis method and device, electronic equipment and computer readable storage medium
CN111083396A
Video generation method and device, equipment and storage medium
CN112866798A
Video generation method and device, equipment and storage medium
CN117979088A