Video generation method, apparatus, device, and storage medium
The video generation method addresses the challenge of meeting diverse user requirements by generating target videos that align with specified video effect requirements, enhancing user experience through tailored video outputs.
Patent Information
- Application Number
- JP2024520785
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-23
- Filing Date
- 2023-12-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-12-06
AI Technical Summary
Existing video generation methods struggle to meet diverse user requirements for video effects, leading to a need for enriched video generation methods that improve user experience.
A video generation method that involves obtaining text information describing video effect requirements and multimedia materials, and then generating a target video by combining video clips formed from these materials, ensuring the video effect conforms to the specified requirements.
The method effectively generates target videos that meet specific video effect requirements, enriching the video generation process and enhancing user experience by providing tailored video outputs.
Smart Images

Figure 2025516416000001_ABST
Abstract
Description
Technical Field
[0001] This application was filed on April 23, 2023, claims the priority of a Chinese invention patent application with the invention title "Video Generation Method, Apparatus, Device and Storage Medium" and application number 202310446304.X, and all the contents of the application are incorporated herein by reference.
[0002] The present disclosure relates to the field of data processing, and particularly to a video generation method, apparatus, device and storage medium.
Background Art
[0003] With the continuous development of video processing technology, users' requirements for video generation methods are becoming increasingly diverse. Therefore, in order to meet users' diverse requirements for video generation methods and improve the user experience, how to enrich the video generation method is a technical problem that needs to be solved urgently.
Summary of the Invention
[0004] To solve the above technical problems, the present disclosure provides a video generation method, apparatus, device and storage medium, which can enrich the video generation method and improve the user experience.
[0005] According to a first aspect, the present disclosure provides a video generation method, the method comprising: obtaining first text information, where the first text information is used to describe video effect requirements; obtaining at least one multimedia material; Generating a target video based on the first text information and the at least one multimedia material, where the at least one multimedia material is displayed in the target video, the video effect of the target video conforms to the video effect requirements described in the first text information, at least one combination of video clips is displayed in the target video, the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and each video material includes video material and / or pixel material.
[0006] In an alternative embodiment, generating a target video based on the first text information and the at least one multimedia material comprises Generating a video editing draft based on the first text information and the at least one multimedia material, where the video editing draft includes the at least one multimedia material and editing information, the editing information is used to instruct an editing operation on the at least one multimedia material, the editing operation is at least used to edit each video material in the at least one multimedia material into the at least one video clip respectively, and the video editing effect corresponding to the editing operation and / or the at least one multimedia material conforms to the video effect requirements described in the first text information. Generating a target video based on the video editing draft.
[0007] In an alternative embodiment, generating a video editing draft based on the first text information and the at least one multimedia material comprises Determining at least one video editing template based on the first text information and the at least one multimedia material, where the editing effect of the at least one video editing template conforms to the video effect requirements described in the first text information. Applying the editing operations instructed by the target video editing template in the at least one video editing template to the at least one multimedia material to generate a video editing draft, is included.
[0008] In an alternative embodiment, determining at least one video editing template based on the first text information and the at least one multimedia material, respectively extracting the feature tags of the first text information and the at least one multimedia material, obtaining at least one video editing template by matching with available video editing templates based on the feature tags of the first text information and the at least one multimedia material, where the at least one video editing template includes a first video editing template that matches the feature tag of the first text information and a second video editing template that matches the feature tag of the at least one multimedia material.
[0009] In an alternative embodiment, obtaining the at least one multimedia material, matching the first multimedia material in the at least one multimedia material from the user material set based on the analysis result of the first text information, and / or generating a second multimedia material in the at least one multimedia material based on the analysis result of the first text information, where the at least one multimedia material conforms to the video effect requirements described in the first text information.
[0010] In an alternative embodiment, before obtaining the first text information, further including displaying a text input box in response to an introduction operation of at least one multimedia material, Correspondingly, obtaining the first text information, including receiving first text information based on the text input box.
[0011] In an alternative embodiment, before receiving the first text information based on the text input box, further including displaying at least one video tag, where the video tag is used to characterize a video effect, Correspondingly, receiving the first text information based on the text input box includes obtaining the first text information based on an operation of adding a target video tag in the at least one video tag to the text input box.
[0012] In an alternative embodiment, after determining at least one video editing template based on the first text information and the at least one multimedia material, selecting a third video editing template from the at least one video editing template and displaying it on a preview page of the video editing effect, and using the video effect obtained by introducing the at least one multimedia material into the third video editing template on the preview page for preview, and an update recommendation control is set on the preview page, further including, in response to a trigger operation of the update recommendation control, selecting a fourth video editing template from the at least one video editing template, replacing the third video editing template displayed on the preview page with the fourth video editing template, and using the video effect obtained by introducing the at least one multimedia material into the fourth video editing template on the preview page for preview.
[0013] In an alternative embodiment, after determining at least one video editing template based on the first text information and the at least one multimedia material, Displaying the fifth video editing template among the at least one video editing template on the preview page; Obtaining adjusted text information in response to a text adjustment operation on the first text information on the preview page; Determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; Further comprising replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second set of video editing templates.
[0014] In an alternative embodiment, before determining the second set of video editing templates based on the adjusted text information and the at least one multimedia material, Further comprising receiving a material adjustment operation on the at least one multimedia material and obtaining adjusted multimedia material; Correspondingly, determining the second set of video editing templates based on the adjusted text information and the at least one multimedia material comprises: Determining a second set of video editing templates based on the adjusted text information and the adjusted multimedia material.
[0015] According to a second aspect, the present disclosure provides a video generation device, the device comprising: A first acquisition module for acquiring first text information, where the first text information is used to describe video effect requirements; A second acquisition module for acquiring at least one multimedia material; A generation module for generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, the video effect of the target video conforms to the video effect requirements described in the first text information, at least one combination of video clips is displayed in the target video, the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and each video material includes a video material and / or a pixel material.
[0016] According to a third aspect, the present disclosure provides a computer-readable storage medium, in which instructions are stored, and when the instructions are executed on a terminal device, the above method is implemented on the terminal device.
[0017] According to a fourth aspect, the present disclosure provides a video generation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the above method is implemented.
[0018] According to a fifth aspect, the present disclosure provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the above method is implemented.
[0019] The technical solutions provided by the embodiments of the present disclosure have at least the following advantages compared with the prior art.
[0020] Embodiments of the present disclosure provide a video generation method. Specifically, after obtaining first text information for describing video effect requirements and at least one multimedia material, a target video is generated based on the first text information and the at least one multimedia material. Here, the at least one multimedia material is displayed in the target video, the video effect of the target video conforms to the video effect requirements described in the first text information, at least one combination of video clips is displayed in the target video, and the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and each video material includes a video material and / or a pixel material. As can be seen from this, embodiments of the present disclosure can generate a target video that conforms to the video effect requirements described in the first text information based on the obtained first text information and multimedia material, enrich the video generation method, and improve the user experience.
Brief Description of the Drawings
[0021] The accompanying drawings here are incorporated herein as part of this specification, illustrate embodiments suitable for the present disclosure, and are used to interpret the principles of the present disclosure in conjunction with the specification.
[0022] To more clearly explain the technical solutions in the embodiments of the present disclosure or the prior art, the drawings that need to be used in the following description of the embodiments or the prior art are briefly described. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative labor.
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0024] To more clearly explain the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. Note that, unless they are mutually contradictory, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0025] In order to fully understand the present disclosure, many specific details will be described in the following description. However, the present disclosure may be implemented in other forms different from the embodiments herein. Obviously, the embodiments herein are only some embodiments of the present disclosure, not all embodiments.
[0026] With the development of video processing technology, users' requirements for video generation methods are becoming increasingly diverse. Therefore, the embodiments of the present disclosure provide a video generation method, which can generate a target video that meets the video effect requirements by analyzing the text information describing the video effect requirements, enrich the video generation method, and improve the user experience.
[0027] Specifically, in the video generation method provided by the embodiments of the present disclosure, after obtaining the first text information for describing the video effect requirements and at least one multimedia material, a target video is generated based on the first text information and the at least one multimedia material. Here, the at least one multimedia material is displayed in the target video, and the video effect of the target video conforms to the video effect requirements described in the first text information. Further, a combination of at least one video clip is displayed in the target video, and the at least one video clip is respectively formed based on the video materials in the at least one obtained multimedia material, and the video materials include video materials and / or pixel materials. As can be seen from this, the embodiments of the present disclosure can generate a target video that conforms to the video effect requirements described in the first text information based on the obtained first text information and multimedia materials, enrich the video generation method, and improve the user experience.
[0028] Based on this, the embodiments of the present disclosure provide a video generation method. Referring to FIG. 1, it is a flowchart of the video generation method provided by the embodiments of the present disclosure, and the method includes the following steps.
[0029] S101: Obtain the first text information.
[0030] Here, the first text information is used to describe the video effect requirements.
[0031] In the embodiments of the present disclosure, the first text information may be text information input by the user. Specifically, the input method of the text information is not limited. For example, the first text information input by voice input may be used, or the first text information input by keyboard input may be used, or the first text information input by text information introduction may be used.
[0032] The first text information is text information capable of describing video effect requirements. Optionally, the video effect requirements described by the first text information may be requirements of a video style type. For example, the first text information may be "comic style". The video effect requirements described by the first text information may also be requirements for video display content. For example, the first text information may be "a warm summer day, a warm afternoon". In the embodiments of the present disclosure, the video effect requirements described by the first text information are not particularly limited.
[0033] S102: Obtain at least one multimedia material.
[0034] In the embodiments of the present disclosure, it is necessary to obtain multimedia materials before generating the target video. Here, the multimedia materials may include images, videos, audios, etc.
[0035] In an alternative embodiment, the multimedia materials may be obtained by user introduction. As shown in FIG. 2, which is a schematic diagram of a material selection page provided by the embodiments of the present disclosure, each multimedia material in the user material set is displayed on the material selection page. When an introduction operation of at least one multimedia material is received, the introduced at least one multimedia material is obtained. Also, when an introduction operation of at least one multimedia material is received, the display of the text input box is triggered. As shown in FIG. 2, after introducing the multimedia material 201, the text input box 202 is displayed on the material selection page. The first text information is input into the text input box 202, and the first text information can be obtained.
[0036] Furthermore, while a text input box is displayed on the material selection page, at least one video tag is displayed, and a plurality of video tags are displayed below the text input box 301 shown in FIG. 3. For example, the video tag is "comic style". Selecting a target video tag from the displayed video tags and adding the target video tag into the text input box 301 is triggered to obtain first text information. Here, the target video tag added into the text input box 301 may include one or more video tags displayed on the material selection page.
[0037] Specifically, the first text information may include only the target video tag, may include only the text information input by the user, or may include the target video tag and the text information input by the user.
[0038] In another selectable embodiment, at least one multimedia material may be obtained based on the analysis of the first text information. Optionally, at least one multimedia material is matched from the user material set based on the analysis result of the first text information. Specifically, the first text information is semantically analyzed by a natural language analysis algorithm, and at least one multimedia material is matched from the user material set based on the analysis result and used to generate a target video.
[0039] Also, at least one multimedia material may be generated based on the analysis result of the first text information. Specifically, the first text information is semantically analyzed by a natural language analysis algorithm, and multimedia materials such as images, video clips, and audio are generated based on the analysis result and used to generate a target video.
[0040] In the embodiments of the present disclosure, the multimedia materials matched from the user material set and the generated multimedia materials based on the analysis result of the first text information both conform to the video effect requirements described in the first text information.
[0041] S103: Generate a target video based on the first text information and the at least one multimedia material.
[0042] Here, the at least one multimedia material is displayed in the target video, the video effect of the target video conforms to the video effect requirements described in the first text information, at least one combination of video clips is displayed in the target video, and each of the at least one video clip is formed based on each video material in the at least one multimedia material, and the video material includes video material and / or pixel material.
[0043] In an embodiment of the present disclosure, after obtaining the first text information and the at least one multimedia material, a target video is generated based on the first text information and the at least one multimedia material. Since the specific video generation method will be specifically described in later embodiments, it will not be repeated here.
[0044] In the video generation method provided by the embodiment of the present disclosure, after obtaining the first text information for describing the video effect requirements and the at least one multimedia material, a target video is generated based on the first text information and the at least one multimedia material. Here, the at least one multimedia material is displayed in the target video, and the video effect of the target video conforms to the video effect requirements described in the first text information. In addition, at least one combination of video clips is displayed in the target video, and each of the at least one video clip is formed based on the video material in the at least one obtained multimedia material, and the video material includes video material and / or pixel material. As can be seen from this, the embodiment of the present disclosure can generate a target video that conforms to the video effect requirements described in the first text information based on the obtained first text information and multimedia material, enrich the video generation method, and improve the user experience.
[0045] Based on the above embodiments, embodiments of the present disclosure further provide a video generation method. Referring to FIG. 4, it is a flowchart of another video generation method provided by embodiments of the present disclosure, and the video generation method specifically includes the following steps.
[0046] S401: Obtain first text information, where the first text information is used to describe video effect requirements.
[0047] S402: Obtain at least one multimedia material.
[0048] In the embodiments of the present disclosure, the methods for obtaining the first text information and at least one multimedia material may be understood with reference to the above embodiments and will not be repeated here.
[0049] S403: Generate a video editing draft based on the first text information and the at least one multimedia material.
[0050] Here, the video editing draft includes the at least one multimedia material and editing information. The editing information is used to instruct an editing operation on the at least one multimedia material. The editing operation is at least used to edit each video material in the at least one multimedia material into the at least one video clip respectively. The video editing effect corresponding to the editing operation and / or the at least one multimedia material conforms to the video effect requirements described in the first text information.
[0051] In the embodiments of the present disclosure, after obtaining the first text information and at least one multimedia material, a video editing draft is generated based on the analysis of the first text information or the comprehensive analysis of the first text information and the at least one multimedia material.
[0052] Here, the video editing draft includes at least one acquired multimedia material and editing information. The editing information is used to instruct the editing operation of the at least one multimedia material. The editing operation is used at least to edit each video material in the at least one multimedia material into one or more video clips. One video clip may include one video material or a combination of multiple video materials. The video editing effect corresponding to the editing operation instructed by the editing information conforms to the video effect requirements described in the first text information, and the multimedia materials in the video editing draft also conform to the video effect requirements described in the first text information.
[0053] In an alternative embodiment, the editing information included in the video editing draft is used to instruct the editing operation determined based on the analysis of the first text information. For example, when the first text information is "a warm summer day, a warm afternoon", based on the analysis of the first text information, it is determined that the editing operation instructed by the editing information adds an A filter to one or multiple video clips.
[0054] In another alternative embodiment, the editing information included in the video editing draft is used to instruct the editing operation determined based on the analysis of the multimedia materials introduced by the user. For example, when the multimedia materials introduced by the user include summer vacation images, video clips, etc., based on the analysis of the multimedia materials introduced by the user, it is determined that the editing operation instructed by the editing information adds a B filter to one or multiple video clips.
[0055] Combining the above two embodiments, the editing operation instructed by the editing information included in the video editing draft may include the editing operation analyzed and determined based on the first text information and the multimedia materials introduced by the user. For the specific method, refer to the descriptions of the above two embodiments, and it will not be repeated here.
[0056] Also, in an alternative embodiment, the editing operations indicated by the editing information included in the video editing draft include the editing operations indicated by the target video editing template, where the editing operations indicated by the target video editing template are used to edit the acquired multimedia material. The target video editing template may be a video editing template selected by the user or a video editing template determined based on the analysis of the first text information. The content related to the target video editing template will be described in detail in later embodiments.
[0057] S404: Generate a target video based on the video editing draft.
[0058] In the embodiments of the present disclosure, after generating a video editing draft based on the acquired first text information and multimedia material, further editing operations are performed on the video editing draft, for example, adjusting all or part of the editing information in the video editing draft.
[0059] In an alternative embodiment, the video editing draft is displayed on the preview page, and in response to the derivation operation of the video editing operation, a target video is generated based on the video editing draft. Here, the video effect of the derived target video conforms to the video effect requirements described in the first text information.
[0060] In the video generation method provided by the embodiments of the present disclosure, after acquiring the first text information describing the video effect requirements and the multimedia material, a video editing draft is generated based on the first text information and the multimedia material. Further, a target video that conforms to the video effect requirements described in the first text information is generated based on the video editing draft. As can be seen from this, in the embodiments of the present disclosure, video editing operations can be generated based on the first text information and the multimedia material, and further a target video that conforms to the video effect requirements described in the first text information can be generated based on the video editing draft, enriching the video generation method and improving the user experience.
[0061] Based on the above embodiments, embodiments of the present disclosure further provide a video generation method. Referring to FIG. 5, it is a flowchart of another video generation method provided by embodiments of the present disclosure. Here, the video generation method includes the following steps.
[0062] S501: Obtain first text information, where the first text information is used to describe video effect requirements. S502: Obtain at least one multimedia material.
[0063] In the embodiments of the present disclosure, the methods for obtaining the first text information and at least one multimedia material may still be understood with reference to the above embodiments and will not be repeated here.
[0064] S503: Determine at least one video editing template based on the first text information and the at least one multimedia material.
[0065] Here, the editing effect of the at least one video editing template conforms to the video effect requirements described in the first text information.
[0066] In an alternative embodiment, after obtaining the first text information and the multimedia material, extract the feature tags of the first text information and the multimedia material respectively. Then, based on the feature tags of the first text information and the at least one multimedia material, match with available video editing templates to obtain at least one video editing template. The at least one video editing template includes a first video editing template that matches the feature tag of the first text information and a second video editing template that matches the feature tag of the at least one multimedia material.
[0067] In an alternative embodiment, based on the feature tags corresponding to the first text information and each of the multimedia materials, match with the available video editing templates from the template library, and obtain at least one video editing template that is successfully matched. Here, the editing effect of the successfully matched video editing template conforms to the video effect requirements described in the first text information.
[0068] Furthermore, after obtaining the video editing template that matches the feature tag of the first text information and the video editing template that matches the feature tag of the multimedia material, mix and arrange the corresponding video editing templates of both, and the video editing template after mixing and arranging is displayed on the preview page.
[0069] As shown in FIG. 6, it is a schematic diagram of the preview page provided by the embodiment of the present disclosure. Here, a preset number of video editing templates are displayed in the lower area 601 of the preview page, and the preset number of video editing templates is determined based on the obtained first text information and at least one multimedia material. The user can trigger the display of more video editing templates on the preview page by performing a swipe operation acting within the lower area 601. For example, by performing a leftward swipe operation, trigger the pull-out display of more video editing templates from the right side of the preview page. Here, the pulled-out and displayed video editing templates are also determined based on the obtained first text information and at least one multimedia material.
[0070] In an alternative embodiment, a third video editing template is selected from at least one obtained video editing template and displayed on a preview page of video editing effects. A video effect obtained by introducing at least one multimedia material obtained on the preview page into the third video editing template is previewed. An update recommendation control is set on the preview page, where the third video editing template is any one selected video editing template displayed on the preview page.
[0071] In response to a trigger operation of the update recommendation control, a fourth video editing template is selected from the at least one video editing template, the third video editing template displayed on the preview page is replaced with the fourth video editing template, and the fourth video editing template is used to preview a video effect obtained by introducing the at least one multimedia material on the preview page into the fourth video editing template. Here, the fourth video editing template for replacing the third video editing template belongs to at least one video editing template determined based on the first text information and the at least one obtained multimedia material.
[0072] In an alternative embodiment, if the video editing template displayed on the preview page cannot meet the current user's usage needs for the video editing template, the current user triggers an updated display of the video editing template displayed on the preview page by triggering an operation of the "batch change" control 602 set on the preview page.
[0073] Specifically, the first video editing template in the first set of video editing templates is displayed on the preview page. Here, the first set of video editing templates is composed of at least one video editing template determined based on the first text information and at least one multimedia material. A preset number of video editing templates in the first set of video editing templates are displayed on the preview page, and a preset number of video editing templates are displayed in the lower area 601 of the preview page shown in FIG. 6, including the first video editing template, and the first video editing template may be any one of the video editing templates displayed on the preview page. In response to a trigger operation of an update recommendation control (the "batch change" control 602 shown in FIG. 6) on the preview page, the first video editing template displayed on the preview page is replaced with the second video editing template in the first set of video editing templates. That is, in response to a trigger operation of an update recommendation control on the preview page, the video editing template displayed on the preview page is replaced with a preset number of video editing templates in the first set of video editing templates to realize the update of the video editing template, and the user generates a target video by using the updated and displayed video editing template on the preview page.
[0074] S504: Apply the editing operation indicated by the target video editing template in the at least one video editing template to the at least one multimedia material to generate a video editing draft.
[0075] In an embodiment of the present disclosure, in response to a selection operation of a target video editing template in at least one video editing template displayed on the preview page, the editing operation indicated by the target video editing template is applied to the obtained multimedia material to generate a video editing operation.
[0076] In practice, the user can also trigger the operation of switching the video editing template. Specifically, in the preview window 603 on the preview page, the preview effect of a video editing draft to which an arbitrary video editing template (for example, video editing template A) is applied is displayed. The user can also trigger the selection operation of another video editing template (for example, video editing template B), and the preview effect of the video editing draft to which video editing template B is applied is displayed in the preview window 603.
[0077] In another selectable embodiment, on the preview page, the user adjusts the first text information according to the preview effect of the video editing draft, and generates a video editing draft that meets the user's video effect requirements.
[0078] Specifically, on the preview page, a video editing template determined based on the initial first text information and the multimedia material is displayed. After receiving the text adjustment operation of the initial first text information, the adjusted text information is obtained. Then, based on the adjusted text information and the multimedia material, a video editing template that meets the video effect requirements described in the adjusted text information is determined again. The user generates a video editing draft that meets the video effect requirements described in the adjusted text information based on the video editing template determined again.
[0079] In an alternative embodiment, the fifth video editing template among the at least one video editing template is displayed on the preview page, where the fifth video editing template is any one of the video editing templates determined based on the first text information and the multimedia material. In response to a text adjustment operation on the first text information on the preview page, adjusted text information is obtained, and then, based on the adjusted text information and the multimedia material, a second set of video editing templates is determined again, and the fifth video editing template displayed on the preview page is replaced with the sixth video editing template in the second set of video editing templates. Here, the sixth video editing template is any one of the video editing templates determined again based on the adjusted text information and the multimedia material.
[0080] Based on the above content, on the preview page, the user can not only adjust the first text information but also adjust the multimedia material according to the preview effect of the video editing draft, and generate a video editing draft that meets the user's video effect requirements.
[0081] In an alternative embodiment, a material adjustment operation of the initial multimedia material is received to obtain adjusted multimedia material, where the adjusted multimedia material includes all or part of the multimedia material in the initial multimedia material, and the material adjustment operation includes operations such as material addition, deletion, and replacement for the initial multimedia material. Based on the adjusted text information and the adjusted multimedia material, a second set of video editing templates is determined. Here, the video editing templates in the second set of video editing templates are determined again based on the first text information (or adjusted text information) and the adjusted multimedia material.
[0082] S505: Generate a target video based on the video editing draft.
[0083] In an embodiment of the present disclosure, after generating a video editing draft, a target video can be generated by triggering an export operation of the video editing draft. Further, the generated target video can be saved locally or in the cloud, or operations such as posting operations can be triggered for the target video.
[0084] In the video generation method provided by the embodiment of the present disclosure, based on the text information describing the video effect requirements and the multimedia materials, a video editing template that meets the video effect requirements is determined, and further, a target video is generated based on the video editing template, which can enrich the video generation method and improve the user experience.
[0085] Based on the above method embodiments, the present disclosure further provides a video generation device. Referring to FIG. 7, it is a schematic structural diagram of the video generation device provided by the embodiment of the present disclosure. The device includes: A first acquisition module 701 for acquiring first text information, where the first text information is used to describe video effect requirements. A second acquisition module 702 for acquiring at least one multimedia material. A generation module 703 for generating a target video based on the first text information and the at least one multimedia material, where the at least one multimedia material is displayed in the target video, the video effect of the target video meets the video effect requirements described in the first text information, at least one combination of video clips is displayed in the target video, and the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and each video material includes video material and / or pixel material.
[0086] In an alternative embodiment, the generation module A first generation sub-module for generating a video editing draft based on the first text information and the at least one multimedia material, where the video editing draft includes the at least one multimedia material and editing information, and the editing information is used to instruct an editing operation on the at least one multimedia material. The editing operation is at least used to edit each video material in the at least one multimedia material into the at least one video clip respectively. The video editing effect corresponding to the editing operation and / or the at least one multimedia material conforms to the video effect requirements described in the first text information. A second generation sub-module for generating a target video based on the video editing draft.
[0087] In an alternative embodiment, the second generation sub-module A first determination sub-module for determining at least one video editing template based on the first text information and the at least one multimedia material, where the editing effect of the at least one video editing template conforms to the video effect requirements described in the first text information. A third generation sub-module for applying the editing operation instructed by the target video editing template in the at least one video editing template to the at least one multimedia material to generate a video editing draft.
[0088] In an alternative embodiment, the first determination sub-module An extraction sub-module for extracting the first text information and the feature tags of the at least one multimedia material respectively. A first matching sub-module for obtaining at least one video editing template by matching with an available video editing template based on the feature tags of the first text information and the at least one multimedia material, where the at least one video editing template includes a first video editing template that matches the feature tags of the first text information and a second video editing template that matches the feature tags of the at least one multimedia material.
[0089] In an alternative embodiment, the second acquisition module includes a second matching sub-module for matching a first multimedia material in at least one multimedia material from a user material set based on the analysis result of the first text information, and / or and / or a fourth generation sub-module for generating a second multimedia material in at least one multimedia material based on the analysis result of the first text information, where the at least one multimedia material meets the video effect requirements described in the first text information.
[0090] In an alternative embodiment, the device further includes a first display module for displaying a text input box in response to an introduction operation of at least one multimedia material, and correspondingly, the first acquisition module is specifically used for receiving first text information based on the text input box. Correspondingly, the first acquisition module is specifically used to receive the first text information based on the text input box.
[0091] In an alternative embodiment, the device further includes a second display module for displaying at least one video tag, where the video tag is used to characterize a video effect, and correspondingly, the first acquisition module is specifically Correspondingly, the first acquisition module is specifically It is used to obtain first text information based on an operation of adding a target video tag in the at least one video tag to the text input box.
[0092] In an alternative embodiment, the apparatus selects a third video editing template from the at least one video editing template and displays it on a preview page of a video editing effect, and is used to preview a video effect obtained by introducing the at least one multimedia material into the third video editing template on the preview page, and a third display module for setting an update recommendation control on the preview page; In response to a trigger operation of the update recommendation control, a fourth video editing template is selected from the at least one video editing template, the third video editing template displayed on the preview page is replaced with the fourth video editing template, and the at least one multimedia material is introduced into the fourth video editing template on the preview page. And a first replacement module used to preview the obtained video effect.
[0093] In an alternative embodiment, the apparatus a fourth display module for displaying a fifth video editing template in the at least one video editing template on a preview page; a first adjustment module for obtaining adjusted text information in response to a text adjustment operation on the first text information on the preview page; a first determination module for determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; and a second replacement module for replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second set of video editing templates.
[0094] In an alternative embodiment, the apparatus further comprises a second adjustment module for receiving a material adjustment operation for the at least one multimedia material and obtaining an adjusted multimedia material. Correspondingly, specifically, the first determination module is used to determine a second set of video editing templates based on the adjusted text information and the adjusted multimedia material.
[0095] The video generation apparatus provided by the embodiments of the present disclosure obtains first text information for describing video effect requirements and at least one multimedia material, and then generates a target video based on the first text information and the at least one multimedia material. Here, the at least one multimedia material is displayed in the target video, and the video effect of the target video conforms to the video effect requirements described in the first text information. At least one combination of video clips is displayed in the target video, and each of the at least one video clip is formed based on each video material in the at least one multimedia material, and each video material includes a video material and / or a pixel material. As can be seen from this, the embodiments of the present disclosure can generate a target video that conforms to the video effect requirements described in the first text information based on the obtained first text information and multimedia material, enrich the video generation method, and improve the user experience.
[0096] In addition to the above method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the video generation method described in the embodiments of the present disclosure is implemented on the terminal device.
[0097] Embodiments of the present disclosure further provide a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the video generation method described in the embodiments of the present disclosure is realized.
[0098] In addition, embodiments of the present disclosure further provide a video generation device. As shown in FIG. 8, it includes a processor 801, a memory 802, an input device 803, and an output device 804. The number of processors 801 in the video generation device may be one or more. In FIG. 8, the case where there is one processor is taken as an example. In some embodiments of the present disclosure, the processor 801, the memory 802, the input device 803, and the output device 804 are connected via a bus or other means. Here, in FIG. 8, the case of being connected via a bus is taken as an example.
[0099] The memory 802 is used to store software programs and modules. The processor 801 realizes various functional applications and data processing of the video generation device by executing the software programs and modules stored in the memory 802. The memory 802 mainly includes a storage program area and a storage data area. Here, the storage program area may store an operating system, at least one application program required for a function, and the like. Further, the memory 802 includes a high-speed random access memory and may also include a non-volatile memory, such as at least one disk memory device, a flash memory device, or other volatile solid-state memory devices. The input device 803 is used not only to receive input numerical or character information, but also to generate signal inputs related to user settings and function controls of the video generation device.
[0100] Specifically, in this embodiment, the processor 801 loads an executable file corresponding to the process of one or more application programs into the memory 802 according to the following instructions, and realizes various functions of the video generation device by executing the application programs stored in the memory 802 by the processor 801.
[0101] It should be noted that in this specification, relative terms such as "first" and "second" are used to distinguish one entity or operation from another entity or operation, and it is required or implied that there is no arbitrary actual relationship or order between these entities or operations. Further, "comprising", "including" or any other variation covers non-exclusive inclusion, and a process, method, article or device comprising a series of elements includes, in addition to those elements, other elements not expressly listed, or elements specific to this process, method, article or device. Unless more limited, an element defined by the expression "comprising one..." does not exclude the presence of another identical element in the process, method, article or device comprising said element.
[0102] The specific embodiments of the present disclosure have been described above for those skilled in the art to fully understand or implement the present disclosure. Multiple modifications of these examples will be obvious to those skilled in the art, and the general principles defined in this specification can be implemented in other examples without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to these examples described in this specification, but should follow the broadest scope consistent with the principles and novel features disclosed in this specification.
Claims
1. Obtaining first text information, where the first text information is used to describe video effect requirements, and obtaining at least one multimedia material, and generating a target video based on the first text information and the at least one multimedia material, where the at least one multimedia material is displayed in the target video, the video effect of the target video conforms to the video effect requirements described in the first text information, at least one combination of video clips is displayed in the target video, the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and each video material includes a video material and / or a pixel material. A video generation method.
2. Generating a target video based on the first text information and the at least one multimedia material includes generating a video editing draft based on the first text information and the at least one multimedia material, where the video editing draft includes the at least one multimedia material and editing information, the editing information is used to instruct an editing operation on the at least one multimedia material, the editing operation is at least used to edit each video material in the at least one multimedia material into the at least one video clip respectively, and the video editing effect corresponding to the editing operation and / or the at least one multimedia material conforms to the video effect requirements described in the first text information, and generating a target video based on the video editing draft. The method according to claim 1.
3. Generating a video editing draft based on the first text information and the at least one multimedia material includes determining at least one video editing template based on the first text information and the at least one multimedia material, where the editing effect of the at least one video editing template conforms to the video effect requirements described in the first text information, and Applying the editing operations instructed by the target video editing template among the at least one video editing template to the at least one multimedia material to generate a video editing draft, the method according to claim 1.
4. Determining at least one video editing template based on the first text information and the at least one multimedia material includes: Extracting the feature tags of the first text information and the at least one multimedia material respectively; Obtaining at least one video editing template by matching with available video editing templates based on the feature tags of the first text information and the at least one multimedia material, the at least one video editing template including a first video editing template that matches the feature tags of the first text information and a second video editing template that matches the feature tags of the at least one multimedia material, the method according to claim 3.
5. Obtaining the at least one multimedia material includes: Matching the first multimedia material among the at least one multimedia material from the user material set based on the analysis result of the first text information; And / or Generating the second multimedia material among the at least one multimedia material based on the analysis result of the first text information, the at least one multimedia material conforming to the video effect requirements described in the first text information, the method according to claim 1.
6. Before obtaining the first text information, Further including displaying a text input box in response to an introduction operation of at least one multimedia material; Correspondingly, obtaining the first text information includes: Receiving the first text information based on the text input box, the method according to claim 1.
7. Before receiving the first text information based on the text input box, Further including displaying at least one video tag, the video tag being used to characterize video effects; Correspondingly, receiving the first text information based on the text input box includes: The method according to claim 6, comprising obtaining first text information based on an operation of adding a target video tag among the at least one video tag to the text input box.
8. After determining at least one video editing template based on the first text information and the at least one multimedia material, selecting a third video editing template from the at least one video editing template and displaying it on a preview page of a video editing effect, and using it to preview a video effect obtained by introducing the at least one multimedia material into the third video editing template on the preview page, and an update recommendation control is set on the preview page; responding to a trigger operation of the update recommendation control, selecting a fourth video editing template from the at least one video editing template, replacing the third video editing template displayed on the preview page with the fourth video editing template, and using it to preview a video effect obtained by introducing the at least one multimedia material into the fourth video editing template on the preview page. The method according to claim 3, further comprising the above.
9. After determining at least one video editing template based on the first text information and the at least one multimedia material, displaying a fifth video editing template among the at least one video editing template on a preview page; obtaining adjusted text information in response to a text adjustment operation on the first text information on the preview page; determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material; replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second set of video editing templates. The method according to claim 3, comprising the above.
10. Before determining the second set of video editing templates based on the adjusted text information and the at least one multimedia material, Further comprising receiving a material adjustment operation for the at least one multimedia material and obtaining an adjusted multimedia material, Correspondingly, determining a second set of video editing templates based on the adjusted text information and the at least one multimedia material, The method according to claim 9, comprising determining a second set of video editing templates based on the adjusted text information and the adjusted multimedia material.
11. A first acquisition module for acquiring first text information, wherein the first text information is used to describe video effect requirements, the first acquisition module, A second acquisition module for acquiring at least one multimedia material, A generation module for generating a target video based on the first text information and the at least one multimedia material, wherein the at least one multimedia material is displayed in the target video, and the video effect of the target video conforms to the video effect requirements described in the first text information, and at least one combination of video clips is displayed in the target video, and the at least one video clip is respectively formed based on each video material in the at least one multimedia material, and each video material includes a video material and / or a pixel material, a video generation device.
12. A computer-readable storage medium having instructions stored thereon, which, when executed on a terminal device, cause the terminal device to execute the method according to any one of claims 1 to 10, a computer-readable storage medium.
13. A video processing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 10 is realized.
Citation Information
Patent Citations
Method and device for generating edited video
CN110996017A
Video display and processing method, device and system, equipment and medium
CN112579826A
Video generation method and device, equipment and storage medium
CN113518160A
Video editing method and device, equipment and storage medium
CN115442539A
Template-based multimedia editor and editing method thereof
US20070083851A1
Cited By
Information processing system, information processing method, and program
JP7886506B1