Video generation method, apparatus, device and storage medium
By generating videos based on text information and multimedia materials using video editing templates and operations, the method addresses diverse user requirements, enhancing user experience.
Patent Information
- Application Number
- JP2024520785
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-04-23
- Filing Date
- 2023-12-06
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2043-12-06
AI Technical Summary
Existing video generation methods fail to meet diverse user requirements, leading to suboptimal user experiences.
A method that involves obtaining text information describing video effect requirements and multimedia materials to generate a target video, ensuring the video effect complies with the specified requirements by using video editing templates and editing operations.
Enriches video generation methods by producing videos that meet user-defined effects, thereby improving user experience.
Smart Images

Figure 0007803636000001 
Figure 0007803636000002 
Figure 0007803636000003
Abstract
Description
[Technical Field]
[0001] This application claims priority to a Chinese invention patent application filed on April 23, 2023, entitled "Video generating method, apparatus, device and storage medium," application number 202310446304.X, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates to the field of data processing, and in particular to video generation methods, apparatus, devices and storage media. [Background technology]
[0003] With the continuous development of video processing technology, users' requirements for video generation methods are becoming more and more diverse. Therefore, how to enrich video generation methods to meet users' diversified requirements for video generation methods and improve user experience is an urgent technical problem to be solved. Summary of the Invention
[0004] To solve the above technical problems, the present disclosure provides a video generation method, apparatus, device and storage medium, which can enrich the video generation method and improve the user experience.
[0005] According to a first aspect, the present disclosure provides a video generation method, the method comprising: obtaining first text information, wherein the first text information is used to describe video effect requirements; Obtaining at least one multimedia material; generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, a video effect of the target video complies with the video effect requirements described in the first text information, and a combination of at least one video clip is displayed in the target video, each of the at least one video clips being formed based on a respective video material in the at least one multimedia material, and each of the video materials including video material and / or image material.
[0006] In an alternative embodiment, generating a target video based on the first text information and the at least one multimedia material includes: generating a video editing draft based on the first text information and the at least one multimedia material, wherein the video editing draft includes the at least one multimedia material and editing information, the editing information is used to instruct editing operations on the at least one multimedia material, the editing operations are used to edit at least each video material in the at least one multimedia material into the at least one video clip, respectively, and the video editing effects corresponding to the editing operations and / or the at least one multimedia material comply with the video effect requirements described in the first text information; generating a target video based on the video editing draft.
[0007] In an alternative embodiment, generating a video edit draft based on the first text information and the at least one multimedia material includes: determining at least one video editing template based on the first text information and the at least one multimedia material, where an editing effect of the at least one video editing template complies with a video effect requirement described in the first text information; applying editing operations indicated by a target video editing template in the at least one video editing template to the at least one multimedia material to generate a video edit draft.
[0008] In an alternative embodiment, determining at least one video editing template based on the first text information and the at least one multimedia material comprises: extracting feature tags of the first text information and the at least one multimedia material, respectively; and obtaining at least one video editing template by matching available video editing templates based on the first text information and feature tags of the at least one multimedia material, wherein the at least one video editing template includes a first video editing template that matches the feature tags of the first text information and a second video editing template that matches the feature tags of the at least one multimedia material.
[0009] In an alternative embodiment, obtaining the at least one multimedia material comprises: Matching a first multimedia material among at least one multimedia material from a user material set based on the analysis result of the first text information; and / or generating a second multimedia material among the at least one multimedia material based on the analysis result of the first text information, wherein the at least one multimedia material complies with the video effect requirements described in the first text information.
[0010] In an optional embodiment, before obtaining the first text information: further comprising displaying a text entry box in response to an introduction operation of the at least one multimedia material; Correspondingly, acquiring the first text information includes: Receiving first text information based on the text input box.
[0011] In an optional embodiment, before receiving first text information based on the text entry box: further comprising displaying at least one video tag, wherein the video tag is used to characterize a video effect; Correspondingly, receiving first text information based on the text input box includes: Obtaining first text information based on an operation of adding a target video tag from the at least one video tag to the text input box.
[0012] In an optional embodiment, after determining at least one video editing template based on the first text information and the at least one multimedia material, selecting a third video editing template from the at least one video editing template and displaying it on a video editing effect preview page, which is used to preview a video effect obtained by introducing the at least one multimedia material into the third video editing template on the preview page, and setting an update recommendation control on the preview page; and further including, in response to a trigger operation of the update recommendation control, selecting a fourth video editing template from the at least one video editing template, replacing the third video editing template displayed on the preview page with the fourth video editing template, and using the preview page to preview a video effect obtained by introducing the at least one multimedia material into the fourth video editing template.
[0013] In an optional embodiment, after determining at least one video editing template based on the first text information and the at least one multimedia material, displaying a fifth video editing template of the at least one video editing template on a preview page; obtaining adjusted text information in response to a text adjustment operation on the first text information on the preview page; determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; The method further includes replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second set of video editing templates.
[0014] In an optional embodiment, prior to determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material, receiving a material adjustment operation for the at least one multimedia material to obtain an adjusted multimedia material; Correspondingly, determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material includes: determining a second set of video editing templates based on the adjusted text information and the adjusted multimedia material;
[0015] According to a second aspect, the present disclosure provides a video generation device, the device comprising: a first obtaining module for obtaining first text information, where the first text information is used to describe video effect requirements; a second acquisition module for acquiring at least one multimedia material; and a generation module for generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, a video effect of the target video complies with the video effect requirements described in the first text information, and a combination of at least one video clip is displayed in the target video, each of the at least one video clips being formed based on a respective video material in the at least one multimedia material, and each of the video materials including video material and / or image material.
[0016] According to a third aspect, the present disclosure provides a computer-readable storage medium having instructions stored thereon that, when executed on a terminal device, cause the terminal device to implement the above method.
[0017] According to a fourth aspect, the present disclosure provides a video generation device, comprising: a memory; a processor; and a computer program stored in the memory and adapted to be executed on the processor, the computer program, when executed by the processor, realizing the method described above.
[0018] According to a fifth aspect, the present disclosure provides a computer program product, said computer program product comprising computer programs / instructions which, when executed by a processor, implement the above method.
[0019] Compared with the prior art, the technical solutions provided by the embodiments of the present disclosure have at least the following advantages:
[0020]
[0013] An embodiment of the present disclosure provides a video generation method, specifically, obtaining first text information describing video effect requirements and at least one multimedia material, and then generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, and the video effect of the target video complies with the video effect requirements described in the first text information, and the target video displays a combination of at least one video clip, each of which is formed based on a respective video component in the at least one multimedia material, and each video component includes video components and / or image components.
[0014] As can be seen from this, the embodiment of the present disclosure can generate a target video that complies with the video effect requirements described in the first text information based on the obtained first text information and multimedia material, thereby enriching the video generation method and improving user experience. [Brief explanation of the drawings]
[0021] The accompanying drawings herein, which are incorporated herein as part of this specification, illustrate suitable embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0022] In order to more clearly describe the technical solutions in the embodiments of the present disclosure or the prior art, the drawings that need to be used in the description of the embodiments or the prior art will be briefly described below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative work.
[0023] [Figure 1] 1 is a flowchart of a video generation method provided by an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram of a material selection page provided by an embodiment of the present disclosure. [Figure 3] FIG. 10 is a schematic diagram of another material selection page provided by an embodiment of the present disclosure. [Figure 4]1 is a flowchart of another video generation method provided by an embodiment of the present disclosure. [Figure 5] 1 is a flowchart of another video generation method provided by an embodiment of the present disclosure. [Figure 6] FIG. 2 is a schematic diagram of a preview page provided by an embodiment of the present disclosure. [Figure 7] 1 is a structural schematic diagram of a video generation device provided by an embodiment of the present disclosure; [Figure 8] 1 is a structural schematic diagram of a video generation device provided by an embodiment of the present disclosure; DETAILED DESCRIPTION OF THE INVENTION
[0024] In order to more clearly describe the above objectives, features and advantages of the present disclosure, the solutions of the present disclosure are further described below. It should be noted that, unless mutually contradictory, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0025] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure; however, the present disclosure may be embodied in other forms different from the embodiments herein, and it is apparent that the examples in this specification are merely some examples of the present disclosure, but not all examples.
[0026] With the development of video processing technology, user requirements for video generation methods are becoming more and more diverse. Therefore, the embodiments of the present disclosure provide a video generation method, which can analyze text information describing video effect requirements to generate a target video that meets the video effect requirements, thereby enriching the video generation method and improving the user experience.
[0027] Specifically, the video generation method provided by the embodiments of the present disclosure includes obtaining first text information describing video effect requirements and at least one multimedia material, and then generating a target video based on the first text information and the at least one multimedia material. The at least one multimedia material is displayed in the target video, and the video effect of the target video complies with the video effect requirements described in the first text information. Furthermore, the target video includes a combination of at least one video clip, each of which is formed based on video material from the at least one acquired multimedia material, and the video material includes video material and / or image material. As can be seen, the embodiments of the present disclosure can generate a target video that complies with the video effect requirements described in the first text information based on the acquired first text information and multimedia material, thereby enriching the video generation method and improving the user experience.
[0028] On this basis, an embodiment of the present disclosure provides a video generation method, and referring to FIG. 1, there is shown a flowchart of the video generation method provided by the embodiment of the present disclosure, which includes the following steps:
[0029] S101: First text information is acquired.
[0030] Here, the first text information is used to describe video effect requirements.
[0031] In the embodiments of the present disclosure, the first text information may be text information input by a user. Specifically, the input method of the text information is not limited, and may be, for example, the first text information input by voice input, the first text information input by keyboard input, or the first text information input by text information introduction.
[0032] The first text information is text information that can describe a video effect requirement, and optionally, the video effect requirement described in the first text information may be a video style type requirement, for example, the first text information may be "cartoon style." The video effect requirement described in the first text information may be a requirement for video display content, for example, the first text information may be "warm summer day, warm afternoon." In the embodiments of the present disclosure, the video effect requirement described in the first text information is not particularly limited.
[0033] S102: At least one multimedia material is obtained.
[0034] In an embodiment of the present disclosure, before generating a target video, multimedia material needs to be acquired, where the multimedia material may include images, video, audio, etc.
[0035] In an optional embodiment, multimedia materials may be acquired through user introduction. As shown in FIG. 2, a schematic diagram of a material selection page provided by an embodiment of the present disclosure shows each multimedia material in the user material collection on the material selection page. When an operation to introduce at least one multimedia material is received, at least one introduced multimedia material is acquired. In addition, when an operation to introduce at least one multimedia material is received, the display of a text input box is triggered. As shown in FIG. 2, after introducing multimedia material 201, a text input box 202 is displayed on the material selection page. First text information is input into the text input box 202, allowing the first text information to be acquired.
[0036] Furthermore, at the same time when the text input box is displayed on the material selection page, at least one video tag is displayed, and multiple video tags are displayed below the text input box 301 shown in FIG. 3, for example, the video tag is "cartoon style". A target video tag is selected from the displayed video tags and added to the text input box 301, thereby obtaining first text information. Here, the target video tag added to the text input box 301 may include one or more video tags displayed on the material selection page.
[0037] Specifically, the first text information may include only the target video tag, may include only the text information input by the user, or may include the target video tag and the text information input by the user.
[0038] In another optional embodiment, at least one multimedia material may be obtained based on an analysis of the first text information. Optionally, at least one multimedia material is matched from the user material set based on the analysis result of the first text information. Specifically, the first text information is semantically analyzed using a natural language analysis algorithm, and at least one multimedia material is matched from the user material set based on the analysis result to generate the target video.
[0039] In addition, at least one multimedia material may be generated based on the analysis result of the first text information. Specifically, the first text information is semantically analyzed using a natural language analysis algorithm, and multimedia material such as images, video clips, and audio is generated based on the analysis result to generate the target video.
[0040] In an embodiment of the present disclosure, based on the analysis result of the first text information, the multimedia material matched from the user material set and the generated multimedia material both comply with the video effect requirements described in the first text information.
[0041] S103: Generate a target video based on the first text information and the at least one multimedia material.
[0042] Here, the at least one multimedia material is displayed in the target video, the video effect of the target video complies with the video effect requirements described in the first text information, and a combination of at least one video clip is displayed in the target video, and the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and the video material includes video material and / or image material.
[0043] In the embodiment of the present disclosure, after obtaining the first text information and at least one multimedia material, a target video is generated based on the first text information and the at least one multimedia material. A specific video generation method will be specifically described in the following embodiment, and will not be repeated here.
[0044] The embodiments of the present disclosure provide a video generation method, which obtains first text information describing video effect requirements and at least one multimedia material, and then generates a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, and the video effect of the target video complies with the video effect requirements described in the first text information. The embodiments of the present disclosure also provide a video generation method, which obtains first text information and at least one multimedia material, and generates a target video based on the first text information and at least one multimedia material, and the target video includes at least one combination of video clips, each of which is formed based on video material from the at least one multimedia material. The video material includes video material and / or image material. As can be seen, the embodiments of the present disclosure can generate a target video that complies with the video effect requirements described in the first text information based on the obtained first text information and multimedia material, thereby enriching the video generation method and improving the user experience.
[0045] Based on the above embodiments, an embodiment of the present disclosure further provides a video generation method. Referring to FIG. 4, there is shown a flowchart of another video generation method provided by an embodiment of the present disclosure, which specifically includes the following steps:
[0046] S401: Obtain first text information, where the first text information is used to describe video effect requirements.
[0047] S402: At least one multimedia material is obtained.
[0048] In the embodiments of the present disclosure, the method for obtaining the first text information and at least one multimedia material can be understood with reference to the above embodiments, and will not be repeated here.
[0049] S403: Generate a video editing draft based on the first text information and the at least one multimedia material.
[0050] Here, the video editing draft includes the at least one multimedia material and editing information, the editing information is used to instruct editing operations on the at least one multimedia material, the editing operations are used to edit at least each video material in the at least one multimedia material into the at least one video clip, and the video editing effects corresponding to the editing operations and / or the at least one multimedia material comply with the video effect requirements described in the first text information.
[0051] In an embodiment of the present disclosure, after obtaining first text information and at least one multimedia material, a video editing draft is generated based on an analysis of the first text information or a comprehensive analysis of the first text information and the at least one multimedia material.
[0052] The video editing draft includes at least one acquired multimedia material and editing information, the editing information is used to instruct editing operations for the at least one multimedia material, and the editing operations are used to edit at least each video material in the at least one multimedia material into one or more video clips, where one video clip may include one video material or a combination of multiple video materials. The video editing effects corresponding to the editing operations instructed by the editing information comply with the video effect requirements described in the first text information, and the multimedia material in the video editing draft also comply with the video effect requirements described in the first text information.
[0053] In an optional embodiment, the editing information included in the video editing draft is used to direct an editing operation determined based on an analysis of the first text information, for example, if the first text information is "warm summer day, warm afternoon," then based on the analysis of the first text information, it is determined that the editing operation directed by the editing information is to add an A filter to the single-stage or multi-stage video clip.
[0054] In another alternative embodiment, the editing information included in the video edit draft is used to direct editing operations determined based on an analysis of the multimedia material introduced by the user. For example, if the multimedia material introduced by the user includes summer vacation images, video clips, etc., based on the analysis of the multimedia material introduced by the user, the editing operation directed by the editing information may be determined to add a B filter to a single-stage or multi-stage video clip.
[0055] Combining the above two embodiments, the editing operations instructed by the editing information contained in the video editing draft may include editing operations analyzed and determined based on the first text information and the multimedia material introduced by the user, and specific methods may refer to the descriptions of the above two embodiments and will not be repeated here.
[0056] In an alternative embodiment, the editing operations indicated by the editing information included in the video editing draft include editing operations indicated by a target video editing template, where the editing operations indicated by the target video editing template are used to edit the acquired multimedia material. The target video editing template may be a video editing template selected by a user or may be a video editing template determined based on an analysis of the first text information, and content related to the target video editing template will be described in detail in the following examples.
[0057] S404: Generate a target video based on the video editing draft.
[0058] In an embodiment of the present disclosure, after generating a video editing draft based on the acquired first text information and multimedia material, further editing operations are performed on the video editing draft, such as adjusting all or part of the editing information in the video editing draft.
[0059] In an optional embodiment, a video edit draft is displayed on the preview page, and in response to a video editing operation deriving operation, a target video is generated based on the video edit draft, where the video effects of the derived target video comply with the video effect requirements described in the first text information.
[0060] In a video generation method provided by an embodiment of the present disclosure, after first text information describing video effect requirements and multimedia materials are obtained, a video editing draft is generated based on the first text information and the multimedia materials, and a target video that conforms to the video effect requirements described in the first text information is generated based on the video editing draft. As can be seen from this, the embodiment of the present disclosure can generate video editing operations based on the first text information and the multimedia materials, and can further generate a target video that conforms to the video effect requirements described in the first text information based on the video editing draft, thereby enriching the video generation method and improving the user experience.
[0061] Based on the above embodiments, an embodiment of the present disclosure further provides a video generation method. Referring to FIG. 5, there is shown a flowchart of another video generation method provided by an embodiment of the present disclosure, in which the video generation method includes the following steps:
[0062] S501: Obtaining first text information, where the first text information is used to describe video effect requirements; S502: At least one multimedia material is obtained.
[0063] In the embodiments of the present disclosure, the method for obtaining the first text information and at least one multimedia material can still be understood with reference to the above embodiments, and will not be repeated here.
[0064] S503: Determine at least one video editing template based on the first text information and the at least one multimedia material.
[0065] Wherein, the editing effect of the at least one video editing template complies with the video effect requirement described in the first text information.
[0066] In an optional embodiment, after obtaining first text information and multimedia material, feature tags of the first text information and the multimedia material are extracted respectively, and then, based on the first text information and the feature tags of the at least one multimedia material, at least one video editing template is obtained by matching with available video editing templates, wherein the at least one video editing template includes a first video editing template matching with the feature tags of the first text information and a second video editing template matching with the feature tags of the at least one multimedia material.
[0067] In an optional embodiment, based on the first text information and the feature tags corresponding to each of the multimedia material, matching with available video editing templates from a template library is performed to obtain at least one successfully matched video editing template, where the editing effect of the successfully matched video editing template complies with the video effect requirement described in the first text information.
[0068] Furthermore, after obtaining the video editing template that matches the feature tag of the first text information and the video editing template that matches the feature tag of the multimedia material, the video editing template corresponding to both is mixed and arranged, and the mixed and arranged video editing template is displayed on the preview page.
[0069] 6 is a schematic diagram of a preview page provided by an embodiment of the present disclosure, in which a preset number of video editing templates are displayed in a lower area 601 of the preview page, and the preset number of video editing templates are determined based on the acquired first text information and at least one multimedia material. A user can trigger the display of more video editing templates on the preview page by swiping within the lower area 601, and can trigger the display of more video editing templates from the right side of the preview page by swiping, for example, to the left. Here, the video editing templates displayed by the pull-out are also determined based on the acquired first text information and at least one multimedia material.
[0070] In an optional embodiment, a third video editing template is selected from the at least one video editing template obtained and displayed on a video editing effect preview page, a video effect obtained by introducing the at least one multimedia material obtained into the third video editing template on the preview page is previewed on the preview page, and an update recommendation control is set on the preview page, wherein the third video editing template is any one of the selected video editing templates displayed on the preview page.
[0071] In response to a trigger operation of the update recommendation control, a fourth video editing template is selected from the at least one video editing template, and the fourth video editing template is used to replace the third video editing template displayed on the preview page, and to preview, on the preview page, a video effect obtained by introducing the at least one multimedia material into the fourth video editing template, wherein the fourth video editing template for replacing the third video editing template belongs to the at least one video editing template determined based on the first text information and the acquired at least one multimedia material.
[0072] In an optional embodiment, if the video editing template displayed on the preview page cannot meet the current user's usage needs for the video editing template, the current user triggers an updated display of the video editing template displayed on the preview page by triggering the "global change" control 602 set on the preview page.
[0073] Specifically, a first video editing template from a first set of video editing templates is displayed on the preview page. Here, the first set of video editing templates is composed of at least one video editing template determined based on the first text information and at least one multimedia material. A predetermined number of video editing templates from the first set of video editing templates are displayed on the preview page, and the predetermined number of video editing templates are displayed in the lower area 601 of the preview page shown in FIG. 6, including a first video editing template, which may be any one of the video editing templates displayed on the preview page. In response to a trigger operation of an update recommendation control on the preview page (the "Bulk Change" control 602 shown in FIG. 6), the first video editing template displayed on the preview page is replaced with a second video editing template from the first set of video editing templates. That is, in response to a trigger operation of the update recommendation control on the preview page, the video editing templates displayed on the preview page are replaced with the predetermined number of video editing templates from the first set of video editing templates, thereby realizing the update of the video editing templates, and the user generates a target video using the updated video editing templates displayed on the preview page.
[0074] S504: Applying editing operations indicated by a target video editing template in the at least one video editing template to the at least one multimedia material to generate a video editing draft.
[0075] In an embodiment of the present disclosure, in response to a selection operation of a target video editing template from at least one video editing template displayed on a preview page, editing operations indicated by the target video editing template are applied to the acquired multimedia material to generate video editing operations.
[0076] In practice, the user can also trigger a video editing template switching operation, specifically, a preview effect of a video editing draft to which any video editing template (e.g., video editing template A) is applied is displayed in the preview window 603 on the preview page, and the user can also trigger a selection operation of another video editing template (e.g., video editing template B), and a preview effect of a video editing draft to which video editing template B is applied is displayed in the preview window 603.
[0077] In another optional embodiment, in the preview page, the user adjusts the first text information according to the preview effect of the video editing draft, to generate a video editing draft that meets the user's video effect requirements.
[0078] Specifically, the preview page displays a video editing template determined based on the initial first text information and the multimedia material. After receiving the text adjustment operation of the initial first text information, the adjusted text information is obtained. Then, based on the adjusted text information and the multimedia material, a video editing template that meets the video effect requirements described in the adjusted text information is determined again. The user generates a video editing draft that meets the video effect requirements described in the adjusted text information based on the determined video editing template.
[0079] In an optional embodiment, a fifth video editing template from the at least one video editing template is displayed on the preview page, where the fifth video editing template is any one of the video editing templates determined based on the first text information and the multimedia material. In response to a text adjustment operation of the first text information on the preview page, adjusted text information is obtained, and then a second set of video editing templates is determined again based on the adjusted text information and the multimedia material, and the fifth video editing template displayed on the preview page is replaced with a sixth video editing template from the second set of video editing templates, where the sixth video editing template is any one of the video editing templates determined again based on the adjusted text information and the multimedia material.
[0080] Based on the above content, on the preview page, the user can not only adjust the first text information but also adjust the multimedia materials according to the preview effect of the video editing draft, so as to generate a video editing draft that meets the user's video effect requirements.
[0081] In an alternative embodiment, a material adjustment operation of the initial multimedia material is received to obtain an adjusted multimedia material, where the adjusted multimedia material includes all or part of the multimedia material in the initial multimedia material, and the material adjustment operation includes operations such as adding, deleting, or replacing material with respect to the initial multimedia material. A second set of video editing templates is determined based on the adjusted text information and the adjusted multimedia material, where the video editing templates in the second set of video editing templates are again determined based on the first text information (or the adjusted text information) and the adjusted multimedia material.
[0082] S505: Generate a target video based on the video editing draft.
[0083] In the embodiments of the present disclosure, after generating a video editing draft, a target video can be generated by triggering a video editing draft deriving operation, and the generated target video can be stored locally or in a cloud, or a posting operation or the like can be triggered for the target video.
[0084] The video generation method provided by the embodiments of the present disclosure can determine a video editing template that meets the video effect requirements based on text information describing the video effect requirements and multimedia materials, and then generate a target video based on the video editing template, thereby enriching the video generation method and improving the user experience.
[0085] Based on the above method embodiments, the present disclosure further provides a video generating device. Referring to FIG. 7, a structural schematic diagram of a video generating device provided by an embodiment of the present disclosure is shown, which includes: a first obtaining module 701 for obtaining first text information, where the first text information is used to describe video effect requirements; a second acquisition module 702 for acquiring at least one multimedia material; and a generation module 703 for generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, a video effect of the target video complies with the video effect requirements described in the first text information, and a combination of at least one video clip is displayed in the target video, each of the at least one video clips being formed based on each video material in the at least one multimedia material, and each of the video materials including video material and / or image material.
[0086] In an alternative embodiment, the generation module: a first generating sub-module for generating a video editing draft based on the first text information and the at least one multimedia material, wherein the video editing draft includes the at least one multimedia material and editing information, the editing information is used to instruct editing operations on the at least one multimedia material, the editing operations are used to edit at least each video material in the at least one multimedia material into the at least one video clip, respectively, and the video editing effects corresponding to the editing operations and / or the at least one multimedia material comply with video effect requirements described in the first text information; and a second generating sub-module for generating a target video based on the video editing draft.
[0087] In an alternative embodiment, the second generation sub-module: a first determining sub-module for determining at least one video editing template based on the first text information and the at least one multimedia material, where the editing effect of the at least one video editing template complies with the video effect requirement described in the first text information; and a third generating sub-module for applying editing operations indicated by a target video editing template in the at least one video editing template to the at least one multimedia material to generate a video editing draft.
[0088] In an alternative embodiment, the first determination sub-module: an extraction sub-module for extracting feature tags of the first text information and the at least one multimedia material, respectively; and a first matching sub-module for matching with available video editing templates based on the first text information and feature tags of the at least one multimedia material to obtain at least one video editing template, wherein the at least one video editing template includes a first video editing template matching feature tags of the first text information and a second video editing template matching feature tags of the at least one multimedia material.
[0089] In an alternative embodiment, the second acquisition module: a second matching sub-module for matching a first multimedia material among at least one multimedia material from a user material set based on the analysis result of the first text information; and / or and a fourth generating sub-module for generating a second multimedia material among the at least one multimedia material based on the analysis result of the first text information, wherein the at least one multimedia material complies with the video effect requirements described in the first text information.
[0090] In an alternative embodiment, the apparatus comprises: a first display module for displaying a text entry box in response to an introduction operation of at least one multimedia material; Correspondingly, the first acquisition module specifically: The text input box is used to receive first text information based on the text input box.
[0091] In an alternative embodiment, the apparatus comprises: further comprising a second display module for displaying at least one video tag, wherein the video tag is used to characterize a video effect; Correspondingly, the first acquisition module specifically: Used to obtain first text information based on an operation of adding a target video tag among the at least one video tag to the text input box.
[0092] In an alternative embodiment, the apparatus comprises: a third display module for selecting a third video editing template from the at least one video editing template and displaying it on a video editing effect preview page, and previewing the video effect obtained by introducing the at least one multimedia material into the third video editing template on the preview page, and for setting an update recommendation control on the preview page; and a first replacement module for selecting a fourth video editing template from the at least one video editing template in response to a trigger operation of the update recommendation control, replacing the third video editing template displayed on the preview page with the fourth video editing template, and previewing a video effect obtained by introducing the at least one multimedia material into the fourth video editing template on the preview page.
[0093] In an alternative embodiment, the apparatus comprises: a fourth display module for displaying a fifth video editing template among the at least one video editing template on a preview page; a first adjustment module for obtaining adjusted text information in response to a text adjustment operation on the first text information on the preview page; a first determination module for determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; and a second replacement module for replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second set of video editing templates.
[0094] In an alternative embodiment, the apparatus comprises: a second adjustment module for receiving a material adjustment operation for the at least one multimedia material to obtain an adjusted multimedia material; Correspondingly, the first base determination module specifically includes: A second set of video editing templates is used to determine based on the adjusted text information and the adjusted multimedia material.
[0095] A video generation device according to an embodiment of the present disclosure obtains first text information describing video effect requirements and at least one multimedia material, and then generates a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, and the video effect of the target video complies with the video effect requirements described in the first text information, and the target video displays a combination of at least one video clip, each of which is formed based on a respective video material in the at least one multimedia material, and each video material includes video material and / or image material. From this, it can be seen that the embodiment of the present disclosure can generate a target video that complies with the video effect requirements described in the first text information based on the obtained first text information and multimedia material, thereby enriching the video generation method and improving user experience.
[0096] In addition to the above method and apparatus, an embodiment of the present disclosure further provides a computer-readable storage medium having instructions stored therein, the instructions, when executed by a terminal device, causing the terminal device to implement the video generation method described in the embodiment of the present disclosure.
[0097] An embodiment of the present disclosure further provides a computer program product, which includes computer programs / instructions that, when executed by a processor, implement the video generation method according to the embodiment of the present disclosure.
[0098] Moreover, an embodiment of the present disclosure further provides a video generation device, as shown in FIG. 8 : The video generation device includes a processor 801, a memory 802, an input device 803, and an output device 804. The number of processors 801 in the video generation device may be one or more, and Fig. 8 illustrates an example in which there is one processor. In some embodiments of the present disclosure, the processor 801, the memory 802, the input device 803, and the output device 804 are connected via a bus or other method, and Fig. 8 illustrates an example in which they are connected via a bus.
[0099] The memory 802 is used to store software programs and modules, and the processor 801 executes the software programs and modules stored in the memory 802 to realize various functional applications and data processing of the video generation device. The memory 802 mainly includes a storage program area and a storage data area, where the storage program area may store an operating system, an application program required for at least one function, etc. Furthermore, the memory 802 includes a high-speed random access memory and may also include non-volatile memory, such as at least one disk memory device, flash memory device, or other volatile solid-state memory device. The input device 803 is used not only to receive input numeric or character information but also to generate signal inputs related to user settings and function control of the video generation device.
[0100] Specifically, in this embodiment, the processor 801 loads executable files corresponding to the processes of one or more application programs into the memory 802 in accordance with the following instructions, and executes the application programs stored in the memory 802 by the processor 801, thereby realizing various functions of the above-mentioned video generation device.
[0101] It should be noted that, in this specification, relational terms such as "first" and "second" are used to distinguish one entity or operation from another, and do not require or imply the existence of any actual relationship or order between those entities or operations. Furthermore, "comprises," "includes," or any other variation thereof covers a non-exclusive inclusion, such that a process, method, article, or device comprising a set of elements includes, in addition to those elements, other elements not expressly listed or inherent in the process, method, article, or device. Unless more specific, an element defined by the phrase "comprising one or more of..." does not exclude the presence of other identical elements in the process, method, article, or device that comprises said element.
[0102] Specific embodiments of the present disclosure have been described above so that those skilled in the art can fully understand or realize the present disclosure. Many modifications of these examples will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other examples without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the examples described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. obtaining first text information, the first text information being used to describe video effects requirements; Obtaining at least one multimedia material; generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, a video effect of the target video conforms to the video effect requirements described in the first text information, and a combination of at least one video clip is displayed in the target video, each of the at least one video clips being formed based on a respective video clip in the at least one multimedia material, and each of the video clips includes video and / or image material.
2. generating a target video based on the first text information and the at least one multimedia material, generating a video editing draft based on the first text information and the at least one multimedia material, the video editing draft including the at least one multimedia material and editing information, the editing information being used to instruct editing operations on the at least one multimedia material, the editing operations being used to edit at least each video material in the at least one multimedia material into the at least one video clip, respectively, and video editing effects corresponding to the editing operations and / or the at least one multimedia material comply with video effect requirements described in the first text information; and generating a target video based on the video edit draft.
3. generating a video edit draft based on the first text information and the at least one multimedia material, determining at least one video editing template based on the first text information and the at least one multimedia material, wherein an editing effect of the at least one video editing template complies with a video effect requirement described in the first text information; and applying editing operations indicated by a target video editing template in the at least one video editing template to the at least one multimedia material to generate a video edit draft.
4. Determining at least one video editing template based on the first text information and the at least one multimedia material includes: extracting feature tags of the first text information and the at least one multimedia material, respectively; 4. The method of claim 3, further comprising: matching available video editing templates based on the first text information and feature tags of the at least one multimedia material to obtain at least one video editing template, wherein the at least one video editing template includes a first video editing template that matches the feature tags of the first text information and a second video editing template that matches the feature tags of the at least one multimedia material.
5. Obtaining the at least one multimedia material comprises: matching a first multimedia material among at least one multimedia material from a user material set based on the analysis result of the first text information; and / or 2. The method of claim 1, further comprising: generating a second multimedia material among the at least one multimedia material based on the analysis result of the first text information, wherein the at least one multimedia material conforms to the video effect requirements described in the first text information.
6. Before obtaining the first text information, further comprising displaying a text entry box in response to an introduction of at least one multimedia material; Correspondingly, acquiring the first text information includes: The method of claim 1 , comprising receiving first text information based on the text entry box.
7. before receiving first text information based on the text input box; further comprising displaying at least one video tag, the video tag being used to characterize a video effect; Correspondingly, receiving first text information based on the text input box includes: The method of claim 6 , further comprising obtaining first text information based on an operation of adding a target video tag in the at least one video tag to the text input box.
8. after determining at least one video editing template based on the first text information and the at least one multimedia material; selecting a third video editing template from the at least one video editing template and displaying it on a video editing effect preview page, which is used to preview a video effect obtained by introducing the at least one multimedia material into the third video editing template on the preview page, and setting an update recommendation control on the preview page; 4. The method of claim 3, further comprising: in response to a trigger operation of the update recommendation control, selecting a fourth video editing template from the at least one video editing template, replacing the third video editing template displayed on the preview page with the fourth video editing template, and using the preview page to preview a video effect obtained by introducing the at least one multimedia material into the fourth video editing template.
9. after determining at least one video editing template based on the first text information and the at least one multimedia material; displaying a fifth video editing template of the at least one video editing template on a preview page; obtaining adjusted text information in response to a text adjustment operation on the first text information on the preview page; determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; and replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second set of video editing templates.
10. before determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; receiving a material adjustment operation for the at least one multimedia material to obtain an adjusted multimedia material; Correspondingly, determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material includes: The method of claim 9 , further comprising determining a second set of video editing templates based on the adjusted text information and the adjusted multimedia material.
11. a first obtaining module for obtaining first text information, the first text information being used to describe video effect requirements; a second acquisition module for acquiring at least one multimedia material; a generation module for generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, a video effect of the target video complies with the video effect requirements described in the first text information, and a combination of at least one video clip is displayed in the target video, each of the at least one video clips being formed based on a respective video material in the at least one multimedia material, and each of the video materials including video material and / or image material.
12. A computer readable storage medium having stored thereon instructions that, when executed by a processor of a terminal device, cause the processor of the terminal device to perform the method of any one of claims 1 to 10.
13. A video generation device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, the computer program implementing the method of any one of claims 1 to 10 when executed by the processor.
Citation Information
Patent Citations
Method and device for generating edited video
CN110996017A
Video display and processing method, device and system, equipment and medium
CN112579826A
Video generation method and device, equipment and storage medium
CN113518160A
Video editing method and device, equipment and storage medium
CN115442539A
Template-based multimedia editor and editing method thereof
US20070083851A1