Video generation method and apparatus, device, and storage medium
By receiving video generation instructions and configuration information, and using machine learning models to select highlight content to generate video clips and material clips, the problem of high barriers to material integration in video production is solved, thereby enhancing the richness and interest of the videos and improving generation efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-04-02
AI Technical Summary
Current technologies have a high barrier to entry in integrating media materials during video production, resulting in monotonous and uninteresting video content that fails to meet users' demands for richness and entertainment, thus impacting user experience.
By receiving video generation instructions and obtaining configuration information, video clips and material clips, including video scripts and audio clips, are generated based on media content and configuration information. A machine learning model is used to select highlight content and generate the target video.
It enhances the richness and entertainment value of videos, improves video generation efficiency, and enables users to create high-quality videos more quickly.
Smart Images

Figure CN2025107978_02042026_PF_FP_ABST
Abstract
Description
Video generation method, apparatus, device, and storage medium
[0001] This application claims priority to the Chinese patent application No. 202411350663.6, filed on September 25, 2024, entitled “Video generation method, apparatus, device, and storage medium”, the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a video generation method, apparatus, device, and computer-readable storage medium. BACKGROUND
[0003] With the rapid development of computer technology, various short videos and long videos have become important media for people to share content and express opinions. On this basis, people have increasingly high requirements for the richness and interestingness of video content. Therefore, people expect to be able to produce higher quality videos, and at the same time, expect to complete video creation more efficiently. SUMMARY
[0004] In a first aspect of the present disclosure, a video generation method is provided. The method comprises: receiving a video generation indication for one or more media contents; obtaining configuration information for describing video generation requirements; in response to the video generation indication, generating at least one video segment and at least one material segment corresponding to the at least one video segment respectively based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment contained in a target media content in the one or more media contents, and a material segment in the at least one material segment includes a text segment and / or an audio segment of a video script generated for the corresponding video segment; and generating a target video based on the at least one video segment and the at least one material segment.
[0005] In a second aspect of the present disclosure, a video generation apparatus is provided. The apparatus comprises: an indication receiving module configured to receive a video generation indication for one or more media contents; a configuration information obtaining module configured to obtain configuration information for describing video generation requirements; a segment generating module configured to, in response to the video generation indication, generate at least one video segment and at least one material segment corresponding to the at least one video segment respectively based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment contained in a target media content in the one or more media contents, and a material segment in the at least one material segment includes a text segment and / or an audio segment of a video script generated for the corresponding video segment; and a target video generating module configured to generate a target video based on the at least one video segment and the at least one material segment.
[0006] In a third aspect of the disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer- executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0009] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the embodiments of the present disclosure only and are not intended to limit the scope of the present disclosure to the particular embodiment illustrated. Other BRIEF DESCRIPTION OF DRAWINGS
[0010] The above-mentioned and other features and advantages of various embodiments of the present disclosure will become more apparent by reference to the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like reference numerals designate like elements, wherein:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIGS. 2A-2O show schematic diagrams of example interfaces of video generation according to some embodiments of the present disclosure;
[0013] FIG. 3 shows a schematic diagram of an example timing of utilizing a machine learning model to assist in selecting highlight content according to some embodiments of the present disclosure;
[0014] FIGS. 4A-4C show schematic diagrams of examples of a scene and a corresponding subset of explanatory information according to some embodiments of the present disclosure;
[0015] FIG. 5 shows a flowchart of a video generation process according to some embodiments of the present disclosure;
[0016] FIG. 6 shows a schematic structural block diagram of a video generation apparatus according to certain embodiments of the present disclosure; and
[0017] FIG. 7 shows a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION
[0018] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0019] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, the application program, the server or the storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0020] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of a pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, also carry a selection control for the user to select “agree” or “disagree” to provide the personal information to the electronic device.
[0021] It can be understood that the above notification and obtaining of the authorization of the user are only illustrative, and do not limit the implementation manners of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0022] It can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solutions should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, rather, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0024] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or in a different section / subsection in any manner.
[0025] In this document, unless explicitly stated otherwise, performing a step "in response to" an A does not mean that the step is performed immediately in response to A, but can include one or more intervening steps that are performed in response to A.
[0026] In the description of embodiments of the disclosure, the term "includes" and its derivatives, should be understood to be open terms, i.e., "including, but not limited to." The term "based on" should be understood as "based, at least in part, on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." The following can also include other explicit and implicit definitions. The terms "first", "second", etc. can refer to different or the same objects. The following can also include other explicit and implicit definitions.
[0027] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. In this document, "model" can also be referred to as "machine learning model", "machine learning network" or "network", which are used interchangeably in this document. A model can also include different types of processing units or networks.
[0028] As used herein, the term "component" can refer to any suitable model, module, unit, etc. used to achieve a specific effect. Such a component can provide a corresponding output based on the provided input, and can include any suitable operation, operation, etc. One example of a component is an algorithm. In the following, some embodiments of the disclosure will be mainly described with reference to the algorithm, but it should be understood that such embodiments are also applicable to other types of components.
[0029] As briefly mentioned above, people have higher and higher requirements for the richness, interest, etc. of video content, and expect video production to be more efficient. However, at present, in the process of producing videos, the threshold for integrating media materials is high, and the finished content is monotonous and boring, which makes it difficult to fully meet the needs of users, affecting the user experience.
[0030] In view of this, embodiments of the present disclosure propose an improved solution for video generation. According to various embodiments of the present disclosure, a video generation indication is received for one or more media contents. Configuration information is obtained for describing video generation requirements. In response to the video generation indication, at least one video segment and at least one material segment respectively corresponding to the at least one video segment are generated based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment contained in a target media content in the one or more media contents, and a material segment in the at least one material segment includes a text segment and / or an audio segment of a video script generated for the corresponding video segment. A target video is generated based on the at least one video segment and the at least one material segment. In this way, the one or more media contents are clipped into the target video according to the configuration information. In this way, the user can create a rich and interesting video. This can improve the richness and interest of the generated video, effectively improve the quality of the generated video, and enable the user to conveniently and quickly generate the video, thereby improving the video generation efficiency.
[0031] Example Environment
[0032] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, an electronic device 110 has an application 120 installed therein. A user 140 can interact with the application 120 via the electronic device 110 and / or an attached device of the electronic device 110.
[0033] In some embodiments, the application 120 can be a content sharing application, a content editing application, a content creating application, etc. The application 120 can provide various services related to media contents (also referred to as media content items, content items, media items, etc.) to the user 140, including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), publishing, etc. of the media contents.
[0034] In the environment 100 of FIG. 1, the electronic device 110 can present an interface 150 of the application 120 if the application 120 is in an active state. The interface 150 can include various interfaces that the application 120 can provide, such as a media content presentation interface, a media content creation interface, a media content publishing interface, etc. The application 120 can provide a media content editing function (e.g., the application 120 can be a clipping application) to support editing (e.g., clipping) of media contents in the application 120.
[0035] In some embodiments, the electronic device 110 communicates with the server 130 to enable provisioning of services of the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combinations of the aforementioned which include accessories and peripherals or any combinations thereof. In some embodiments, the electronic device 110 can also be capable of supporting any type of interface for user interaction (such as "wearable" circuitry, etc.). The server 130 can be various types of computing systems / servers capable of providing computing power including, but not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, etc.
[0036] It should be appreciated that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and are not intended to imply any limitations on the scope of the present disclosure.
[0037] Some example embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings. It should be appreciated that the pages shown in the accompanying drawings are merely examples and various page designs can actually exist. The various graphical elements in the pages can have different arrangements and different visual representations, one or more elements among them can be omitted or replaced, and one or more other elements can also exist. The embodiments of the present disclosure are not limited in this respect. In addition, in the following, the example embodiments will be mainly described with respect to the electronic device 110. It should be appreciated that the actions described with respect to the electronic device 110 can be performed by the application 120 on the electronic device 110, or can be performed by the application 120 in coordination with its server (e.g., the server 130).
[0038] Example Interactions
[0039] Example interaction processes according to embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings.
[0040] FIGS. 2A-2O show example interfaces 200A-200O of video generation according to some embodiments of the present disclosure. In embodiments of the present disclosure, the interfaces 200A-200O can be provided by the electronic device 110 shown in FIG. 1.
[0041] As shown in FIGS. 2A-2C, the electronic device 110 receives a video generation indication for one or more media content. The media content discussed herein includes, but is not limited to, a video, an image, a set of images, etc. In some examples, the one or more media content corresponding to the video generation indication can include one or more of a video, an image, a set of images, etc. The specific ways in which the electronic device 110 receives the video generation indication will be discussed in detail below.
[0042] Referring to FIG. 2A, in some embodiments, the electronic device 110 can present an entry 210 for video generation, which can be associated with the video generation indication. In the interface 200A, the entry 210 can be arranged together with entries for other services provided by the electronic device 110 for media content, such as picture editing, video effects, etc. It should be understood that the text identification (e.g., the text “Video with Voiceover”), style, and location, etc. of the entry 210 presented on the electronic device 110 can not be limited to that presented in the interface 200A.
[0043] Referring to FIG. 2B, in some embodiments, the entry 210 can be configured to present a media content selection interface 200B in response to a preset operation by the user 140. The preset operation discussed herein can include, but is not limited to, a click operation, a long press operation, a swipe operation, etc., without limitation. In some embodiments, the media content selection interface 200B can present at least one media content for the user 140 to select one or more media content therefrom.
[0044] In some embodiments, the media content selection interface 200B can include a plurality of media item labels corresponding to different media items. The plurality of media item labels can include, but are not limited to, a video label 220 corresponding to a video media item, a photo label corresponding to a photo media item, and a live label corresponding to a live photo (i.e., a dynamic photo) media item, etc. By distinguishing different media items, the user can be more convenient and efficient in selecting media content. In some embodiments, the media content selection interface 200B can present media content of one media item (e.g., the video media item shown in FIG. 2B) by default, and can present corresponding media content in response to a selection of other media item labels.
[0045] In some embodiments, the media content selection interface 200B can include a region 222 configured to place thumbnails of the selected media content. Illustratively, the electronic device 110 can sequentially mark numbers on the selected media content in the order of selection. It should be understood that the selected media content on the media content selection interface 200B can be deselected. For example, the user 140 can deselect the selected media content by clicking on the selected media content again, or can deselect the media content by the "-" mark corresponding to each media content thumbnail in the region 222. In some embodiments, the plurality of media content thumbnails in the region 222 can receive a drag operation of the user 140 to change the order.
[0046] In some embodiments, at least one of the following requirements can be met for the one or more media content: the total duration of the one or more media content can not exceed a preset total duration (e.g., 10 minutes or other preset total duration), the duration of each media content in the one or more media content can be within a preset duration range (e.g., a duration range of 1 second to 10 minutes, or other preset duration range), and the number of the one or more media content can not exceed a preset number (e.g., 50 or other number). It should be understood that other requirements to be met for the one or more media content can also be set according to actual conditions, which are not limited herein. In addition, when the aforementioned requirements are not met, corresponding prompt content can be presented on the media content selection interface 200B to remind the user 140.
[0047] Referring to FIG. 2C, in some embodiments, the electronic device 110 can present a guide page 230 for a user (for ease of discussion, can be referred to as a new user) who clicks on the portal 210 for the first time. The guide page 230 in the interface 200C can be an independent interface, a window, a panel, or a specific display region in the interface, etc. Illustratively, the guide page 230 can be used to display, for example, a guide example of the specific operation of generating a video on the electronic device 110 according to the embodiments of the present disclosure. The guide page 230 can also be used to display, for example, a web page address of the guide example, so as to jump to another page for presenting the guide example based on a click operation of the new user on the web page address.
[0048] In some embodiments, the interface 200C can include a portal 235 configured to present the media content selection interface 200B of FIG. 2B based on a received preset operation (including but not limited to a click operation, etc.) of the new user. Thus, this is conducive to the new user to understand the interactive operation for video generation more quickly, and can improve the interactive efficiency of the new user and enhance the user experience. Further, the electronic device 110 obtains configuration information for describing the video generation requirement. The configuration information can include configuration information corresponding to the target video to be generated.
[0049] In some embodiments, the configuration information can indicate at least one of a theme of the target video, or a style of the target video. In some other embodiments, the configuration information can further indicate other information such as a visual effect of the target video, a special effect of the picture, etc.
[0050] In some embodiments, to obtain the theme of the target video, the electronic device 110 can present prompt information about the theme of the video, and determine the theme of the target video based on the received user input, as at least part of the configuration information.
[0051] An example is described in combination with FIGS. 2B and 2D. After receiving a preset operation (including but not limited to a click operation, etc.) of the user 140 on the control 224 in the interface 200B, the electronic device 110 can present the configuration panel 240 in the interface 200D. The configuration panel 240 can include prompt information 242 about the theme of the video and an input control 244. Based on the prompt information 242 and the input control 244, the user 140 can input the theme of the target video to be generated to the input control 244. In some embodiments, the input control 244 can support text input and voice input, as well as any other appropriate type of input. In some embodiments, the theme of the video needs to meet preset requirements, including but not limited to a word limit, compliance requirements of the content of the theme, etc.
[0052] In some embodiments, to obtain the style of the target video, the electronic device 110 can present style information about at least one video style. The electronic device 110 can receive a selection of a video style from the at least one video style, and determine the selected video style as the style of the target video, as at least part of the configuration information.
[0053] Continuing to refer to FIG. 2D, the configuration panel 240 can further include style information. The style information can include information corresponding to at least one video style respectively, such as a style name 246, a style image 248, etc. The at least one video style can include a default video style and a plurality of preset video styles. The plurality of preset video styles can include any appropriate style such as a “Everything Can Rap” style, a “Daily Life with Children” style, etc. without limitation. It should be understood that if all the style images 248 are in a disabled state, or the electronic device 110 does not configure any style image 248, the style information can not be displayed on the configuration panel 240.
[0054] In some embodiments, the style image can be a dynamic image, which can be used to play a dynamic picture upon receiving a preset operation of the user 140. In this way, the user can view the display effect of the style in a timely manner, and the interactive experience can be improved.
[0055] Further, if the video generation instruction is received, the electronic device 110 generates at least one video clip based on the one or more media contents and the configuration information. A target video clip in the at least one video clip is a video clip contained in a target media content in the one or more media contents.
[0056] Here, one or more video clips can be selected for each media content in the one or more media contents. Specifically, for a target media content (or any media content) in the one or more media contents, a corresponding target video clip can include one or more video clips. Such a target video clip may, for example, include important clips, pictures, etc. in the target media content, and can also be referred to simply as a highlight clip. In other embodiments, the target video clip of the target media content can also include content having other meanings, which is not limited here.
[0057] In some embodiments, the at least one video clip can be generated by the electronic device 110 receiving a selection of one or more first contents included in the one or more media contents, and comparing a first total duration of the one or more first contents with a preset target duration. If the first total duration does not exceed the preset target duration, the electronic device 110 can select one or more second contents from the one or more media contents using a machine learning model, the one or more second contents being different from the one or more first contents. Then, the electronic device 110 can generate the at least one video clip based on the one or more first contents and the one or more second contents. Depending on the type of media content, the first and second contents described here can have different types. For example, the first and / or second contents can be clips of videos, one or more images, etc.
[0058] In such embodiments, the one or more first contents can be highlight contents selected by the user 140 from the one or more media contents. In some embodiments, the preset target duration can be determined based on a total duration of the one or more media contents multiplied by a preset ratio. It should be understood that the preset target duration can also be determined according to any appropriate manner, which is not limited here. In addition, the type of machine learning model is also not limited in the embodiments of the present disclosure. For example, the machine learning model herein can include one or more of a deep learning model, a supervised learning model, or an unsupervised learning model, etc.
[0059] In some embodiments, if the first total duration of the one or more first contents selected by the user 140 exceeds the preset target duration, a corresponding prompt information can be issued to remind the user 140. In some embodiments, if the duration of the one or more media contents is less than a preset lower limit duration, the user 140 can be refused to select highlight contents from the one or more media contents.
[0060] In some embodiments, to select the one or more second contents, the electronic device 110 can determine a second total duration based on a difference between the preset target duration and the first total duration, and select one or more second contents from the one or more media contents with a duration not exceeding the second total duration by using the machine learning model. The specific manner of selecting the one or more second contents by using the machine learning model will be discussed in detail below in connection with FIG. 3.
[0061] FIG. 3 shows a schematic diagram of an example timing 300 of assisting in selecting highlight contents by using a machine learning model according to some embodiments of the present disclosure. In the timing 300, it is assumed that the total duration of the one or more media contents 310 is 1 minute, and the highlight contents selected by the user 140 include the content 322, the content 324, and the content 326 (which can be collectively referred to as the highlight contents 320), then when the machine learning model performs semantic segmentation on the one or more media contents 310, the highlight contents 320 can be identified and retained. Then, the machine learning model can select other highlight contents from the one or more media contents 310 with a duration not exceeding a second total duration (i.e., the preset target duration minus the first total duration) as the one or more second contents. In this way, the electronic device 110 can combine the one or more first contents and the one or more second contents to generate at least one video clip.
[0062] It should be understood that when the indication of not manually selecting the highlight contents is received, the highlight contents can be directly selected by using the machine learning model based on the authorization of the user 140. As an example, the first contents can be a clip of a video, or an image (e.g., including a picture or a photo, etc.). It should be noted that if the user 140 selects a separate picture or photo as part of the highlight contents 320, the electronic device 110 can convert such contents into a video clip of 3 seconds or other shorter duration by default. In this way, the present embodiment can effectively improve the interaction efficiency by assisting in selecting the highlight contents by using the machine learning model.
[0063] In some embodiments, the target video clip can be determined in the following manner: if the electronic device 110 receives a content selection indication for a target media content, the electronic device 110 can present at least a part of the target media content and a marker of a selection range, the selection range including a video clip of the target media content that is selected by default. Based on a received adjustment indication of the selection range, the electronic device 110 can determine an adjusted selection range, and determine a video clip of the target media content falling within the adjusted selection range as the target video clip.
[0064] With reference to FIGS. 2B and 2E, at the media content selection interface 200B of FIG. 2B, based on the media content selected by the user 140, the electronic device 110 can receive a preset operation of the user 140 on the control 226 (e.g., the control with the label "mark highlight") in response to the content selection indication for the target media content, and in turn can present the mark highlight interface 200E of FIG. 2E. In the mark highlight interface 200E, the selection box 250 can be used to determine the selection range of the image. The left and right edges of the selection box 250 can be used as the markers of the selection range.
[0065] In some embodiments, the selection box 250 can appear by default from the beginning of the video, and the selection box 250 can have a default width (i.e., a default selection range) corresponding to the video segment that is selected by default. Illustratively, the default width can be determined according to the length of the target media content. For example, if the length of the target media content exceeds 2 seconds or other length, the default width of the selection box 250 can be half of the time axis. If the length of the target media content is less than 2 seconds, the selection box 250 can have a default width of 1 second. In some embodiments, the selection box 250 can receive a drag operation of the user 140 to adjust the width of the selection box 250, i.e., to adjust the selection range.
[0066] The adjustment of the selection range based on the selection box 250 is discussed in detail with reference to the interfaces 200F to 200I. At the interface 200F, the user 140 can click the "mark highlight" button 255 to mark the currently selected segment 262, which can be overlaid with a color different from the color of the selection box 250. Subsequently at the interface 200G, the selection box 250 appears on the right side with a width of 1 / 2 or other proportion of the remaining segment on the right side. If the length of the remaining segment on the right side is less than 2 seconds or other length, the selection box 250 appears on the left side with a width of 1 / 2 or other proportion of the remaining segment on the left side. It should be noted that the same time period cannot be added with a highlight mark twice.
[0067] At the interface 200H, the selection box 250 can be invoked again by clicking the previously marked highlight segment 262. At this time, the "cancel highlight" button 258 can be clicked to cancel the highlight mark of the current segment 262. The presentation color of the segment 262 after the highlight mark is canceled can be the color of the selection box 250, as presented at the interface 200I. In such embodiments, the edge of the selection box 250 can be dragged to adjust the length, and the drag operation is to lengthen the segment with the highlight mark. In addition, the video can be paused when the selection box 250 is dragged on the time axis. The video can be played from the position of the pointer when the play button is clicked. After the highlight mark is selected and then unselected, the highlight segment data can be discarded and does not need to be remembered.
[0068] Based on the above embodiments, by selecting highlight content for the media content, the richness and interest of the generated target video can be improved, effectively improving the quality of the generated video. Through the above specific selection method, the efficiency of selecting highlight content can be improved, thereby improving the efficiency of video generation while ensuring video quality.
[0069] In addition, if the video generation instruction is received, based on the one or more media contents and the configuration information, the electronic device 110 further generates at least one material segment corresponding to the at least one video segment respectively. The material segment in the at least one material segment includes a text segment and / or an audio segment of the video script generated for the corresponding video segment. Further, based on the at least one video segment and the at least one material segment, the electronic device 110 generates a target video. Illustratively, the material segment can include commentary information or voiceover information of the corresponding video segment. For example, the text segment can be the text of the voiceover, and the audio segment can be the sound of the voiceover.
[0070] In some embodiments, the audio segment can include audio of reading the text segment in a target voice. In some embodiments, the text segment can include text that corresponds to be presented on the target video with the scene picture of the target video, such text can be referred to as subtitles. The at least one material segment can also be referred to as voiceover content. The target voice can include, for example, a male voice, a female voice, a child voice, a deep voice, an electronic voice, etc., which is not limited herein. In some embodiments, the electronic device 110 can set an appropriate speech rate for the audio segment. In addition, a speech rate upper limit value and a speech rate lower limit value can also be set for the speech rate for selection by the user 140.
[0071] In some embodiments, the target voice can be determined in the following manner: if the configuration information indicates that the style of the target video is a default video style, based on the theme of the target video indicated by the configuration information, the electronic device 110 can determine a voice that matches the theme of the target video as the target voice. In this way, the target voice can still be adapted to the target video in the case of the default video style.
[0072] In some embodiments, if the configuration information indicates that the style of the target video is a preset video style, the electronic device 110 can determine a voice that matches the preset video style as the target voice. For example, if the preset video style of the target video is a "with children daily life" style, the matching voice can be set to a female voice.
[0073] In some embodiments, the target video can include multiple scenes, and the multiple scenes can correspond to multiple sub-material segments included in the at least one material segment respectively. In other words, each scene needs to be aligned with the respective text sub-segment and audio sub-segment. The alignment manner of the scene and the corresponding sub-material segment is discussed in detail below with reference to FIG. 4.
[0074] FIG. 4A illustrates a diagram of an example 400A of a scene and corresponding sub- material segments, according to some embodiments of the present disclosure. The example 400A includes an example 410 and an example 420, the example 410 can be an example where a scene and corresponding sub-material segments are not completely aligned, and the example 420 can be an example of a solution to the example 410.
[0075] In the example 410, the scene 412 and the corresponding sub-material segments 414 (including a text sub-segment 414-1 and an audio sub-segment 414-2) are not completely aligned in time. As can be seen, the sub-material segments 414 span across two scenes. At this time, the text sub-segment 414-1 and the audio sub-segment 414-2 need to be cut according to the timestamps respectively to ensure that they are aligned with the scene 412, according to the solution of the example 420. It should be noted that since the commentary audio can be generated based on the commentary text, the audio sub-segment and the corresponding text sub-segment have a correlation relationship and can be aligned in the timestamps.
[0076] In addition, it should be noted that the video segments can maintain the original segments uploaded by the user 140. For example, the first segment of the original media content uploaded by the user 140 has a length of 10 seconds, and after splitting and extracting the highlight content, it becomes two segments of 5 seconds each. For the original media content that can still be seen during post-editing (such as replacement), the replacement can only frame 5 seconds.
[0077] In some embodiments, the length of each scene in the plurality of scenes can be greater than or equal to the length of the corresponding sub-material segment. In some embodiments, if the length of a scene is less than the length of the corresponding sub-material segment, the length of a number of picture frames of the scene can be increased. In other embodiments, the corresponding sub-material segment of the scene can be regenerated to match the scene. The two cases will be discussed in detail below in connection with FIGS. 4B and 4C.
[0078] FIG. 4B illustrates a diagram of an example 400B of a scene and corresponding sub- material segments, according to some embodiments of the present disclosure. The example 400B includes an example 430, an example 440, and an example 450. The example 430 can be an example where a scene and corresponding sub-material segments are completely aligned. The example 440 can be an example where the length of a scene is less than the length of the corresponding sub-material segment. The example 450 can be an example where the length of a scene is greater than the length of the corresponding sub-material segment.
[0079] In the example 440, if the length of the sub-material segment 444 exceeds the length of the corresponding scene 442, the video media content can be automatically lengthened to adapt to the length first. If the lengthened media content is still not enough to fill, the last frame (a freeze frame) can be lengthened to fill the length.
[0080] In example 450, if the duration of sub-clip 446 is less than the duration of corresponding scene 442, and if there is no freeze frame, the duration of scene 442 can not be adjusted. If there is a freeze frame, the duration is adapted to the duration of sub-clip 446.
[0081] In some other embodiments, if the duration of the target video is changed, the duration of the text sub-clip is adjusted to adapt to the duration of the corresponding audio sub-clip. In this case, if the duration of the audio sub-clip exceeds the duration of the scene, the scene picture needs to be readjusted in the manner discussed above.
[0082] FIG. 4C shows a schematic diagram of an example 400C of scenes and corresponding sub-clip, according to some embodiments of the present disclosure. Example 400C includes example 430 in example 400B and an example 460 for representing a regenerated sub-clip corresponding to a scene. In such an example, if the target video regenerates at least one clip, the duration of the sub-clip and the corresponding scene can be re-adapted using a machine learning model to realign.
[0083] In some embodiments, electronic device 110 can receive an edit to the generated target video and the at least one clip. The edit may, for example, include but is not limited to replacement of media content, modification of the at least one clip, etc. Specific embodiments regarding the edit will be discussed in detail below.
[0084] In some embodiments, electronic device 110 can present respective summary information of a plurality of clips in the target video, each clip of the plurality of clips including a video clip of the at least one video clip and a clip corresponding to the video clip. Electronic device 110 can receive a replacement indication for a target clip of the plurality of clips and a selection of second media content, and replace the target clip with the second media content. The duration of the second media content can be greater than or equal to the duration of the target clip. In some embodiments, the summary information of a clip can include a picture frame of the clip and a duration of the clip.
[0085] Referring to FIGS. 2J-2K, at interface 200J, multiple segments of the target video can be displayed. For the located segment, a play indicator can be displayed to represent the play status. For the selected segment, the play can be stopped, and the display of the target video can be positioned to the selected segment. In the paused state, the electronic device 110 can receive a preset operation of the user 140 on the replace button 272 to present interface 200K. At interface 200K, the user 140 can select media content to replace. It should be noted that only media content longer than the time length of the segment to be replaced can be selected. If the media content is shorter than the segment to be replaced, the media content is presented in a non-selectable state. For the qualified media content, the user 140 can directly replace the original segment after clicking one of the media content.
[0086] In addition, interface 200J can also include a crop button 274. For the selected segment, the electronic device 110 can receive a preset operation of the user on the crop button 274 to present interface 200L. The user can adjust the picture ratio at interface 200L. Interface 200L includes a selection box 276, through which multiple picture frames corresponding to a fixed time length can be framed to enable the user 140 to crop the framed picture frames.
[0087] In some embodiments, the electronic device 110 can present a text segment included in at least one material segment. Based on a selection of a text unit in the presented text segment, the electronic device 110 can receive an edit of the selected text unit to update the selected text unit.
[0088] Referring to FIG. 2M, interface 200M can include a subtitle panel 280. The subtitles (i.e., commentary text) on the subtitle panel 280 can be displayed in time-stamped segments (or can be displayed in other forms, which are not limited herein) to form a text segment. In the case of segmented display, each text unit can include one or more subtitles. The electronic device 110 can highlight the currently played subtitle and support scrolling up and down the subtitles, and the highlight area is automatically positioned to the first subtitle segment during scrolling. The electronic device 110 can receive a click operation of the user 140 on the subtitles to trigger a subtitle editing state. The subtitle panel 280 can include an edit button 282, which can also trigger the subtitle editing state by receiving a preset operation of the user 140. In the subtitle editing state, the electronic device 110 can receive an edit of the subtitles to update the subtitles. It should be noted that a single subtitle on the subtitle panel does not support direct deletion.
[0089] In some embodiments, after the subtitle is updated, the electronic device 110 can synchronize and update the audio according to the new subtitle. This is conducive to improving the quality of the generated target video. Based on the above embodiments, the electronic device 110 can improve the interactivity by editing and modifying the media content and at least one material segment of the target video, so that the user can more conveniently and quickly generate a target video that meets the needs.
[0090] In some embodiments, the target video can include target music. The target music may, for example, include but is not limited to background music of the target video. In some embodiments, the electronic device 110 can determine the target music based on the style of the target video indicated by the configuration information.
[0091] In some embodiments, if the configuration information indicates that the style of the target video is a preset video style, the electronic device 110 can determine music matching the preset video style as the target music. In some embodiments, if the configuration information indicates that the style of the target video is a default video style, the electronic device 110 can determine recommended music as the target music. As an example, the electronic device 110 can use a machine learning model to recommend music. Illustratively, the target music can be selected from a preset music library. It should be understood that the music in the music library is music with copyright.
[0092] In some embodiments, the target music can be automatically adapted to the length of the target video. If the length of the target music is greater than the length of the target video, the end of the target music can be truncated. If the length of the target music is less than the length of the target video, the target music can be looped until the length of the target video is reached. It should be noted that if a machine learning model is used to match the target music for the target video, the length of the music is ensured to be greater than or equal to the length of the target video.
[0093] Referring back to FIG. 2D, the interface 200D further includes a “generate video” button 245. After the electronic device 110 receives a preset operation of the button 245 by the user 140, a video loading page is entered. The loading page prompt text (the loading progress is estimated by the server 130 to the electronic device 110 for a preset time) can display “Parsing media content... x%”, “Generating text... x%”, “Intelligent dubbing... x%”, “Intelligent music matching... x%”, “Almost done... x%”, and the like. The video loading page can include a “check later” button, so that the user 140 clicks the “check later” button to load the video generation task offline and exit the loading page. The video loading page can also include a “cancel” button, so that the user 140 clicks the “cancel” button to terminate the task and return to the interface 200D. In addition, if the target video generation fails (the reasons may, for example, include video theme non-compliance, network anomalies, etc.), the corresponding prompt information can be displayed on the loading page.
[0094] In some embodiments, for the generated video, the electronic device 110 can present the corresponding draft content. The draft content may, for example, include a cover image of the corresponding video, a video theme, style information, and the like. In addition, the draft content can be deleted.
[0095] Referring back to FIG. 2M, for the generated target video, the electronic device 110 can receive a preset operation of the user 140 on the “export” button 284 to present an export panel. On the export panel, the target video can be directly exported, exported and shared to other applications or shared to other users, and the like. In addition, the electronic device 110 can receive a quality setting of the target video by the user 140, such as a resolution setting, and the like.
[0096] In some embodiments, the electronic device 110 can receive a regeneration indication for the target video. If the electronic device 110 receives the regeneration indication, the electronic device 110 can generate another target video based on the modification related to the target video. Specific embodiments in which the electronic device 110 generates another target video will be discussed in detail below.
[0097] In some embodiments, to generate another target video, the electronic device 110 can receive at least one of a replacement indication, a deletion indication, or an addition indication for one or more media contents to determine another one or more media contents. Based on the other one or more media contents and the configuration information, the electronic device 110 can generate another target video.
[0098] Referring to FIG. 2N, the video regeneration interface 200N can include controls related to modifying the target video, including but not limited to at least one of a replacement control, a deletion control, or an addition control 292. Taking the addition control 292 as an example, the electronic device 110 can receive a preset operation of the user 140 on the addition control 292 to obtain another target video identical to the target video. Based on the modification of the one or more media contents and the configuration information in the other target video, the electronic device 110 can generate another target video.
[0099] In some embodiments, to generate another target video, the electronic device 110 can also receive a modification of the configuration information of the target video, and generate another target video based on the modified configuration information and the one or more media contents.
[0100] With reference to FIGS. 2N-2O, the video regeneration interface 200N can further include a configuration information modification control 294 configured to receive a preset operation of the user 140 to present a configuration information modification interface 200O. The configuration information modification interface 200O can include relevant information of a video theme and relevant information of a video style of another target video to be generated. The user 140 can click a regeneration control 296 after re-inputting the video theme and re-selecting the video style to generate another target video.
[0101] In summary, according to the embodiments of the present disclosure, one or more media contents can be clipped into a target video according to configuration information. In this way, the user can be helped to create a rich and interesting video. This can improve the richness and interest of the generated video, effectively improve the quality of the generated video, and enable the user to conveniently and quickly generate a video, thereby improving the video generation efficiency.
[0102] Example process
[0103] FIG. 5 illustrates a flowchart of a video generation process 500 according to some embodiments of the present disclosure. The process 500 can be implemented at the electronic device 110. The process 500 is described below with reference to FIG. 1.
[0104] At block 510, the electronic device 110 receives a video generation indication for one or more media contents.
[0105] At block 520, the electronic device 110 obtains configuration information describing a video generation requirement.
[0106] At block 530, the electronic device 110 generates, in response to the video generation indication, at least one video segment and at least one material segment corresponding to the at least one video segment, based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment contained in a target media content in the one or more media contents, and a material segment in the at least one material segment includes a text segment and / or an audio segment of a video script generated for the corresponding video segment.
[0107] At block 540, the electronic device 110 generates a target video based on the at least one video segment and the at least one material segment.
[0108] In some embodiments, the at least one video clip is generated by: receiving a selection of one or more first contents included in the one or more media contents; comparing a first total time length of the one or more first contents with a preset target time length; in response to the first total time length not exceeding the preset target time length, selecting one or more second contents from the one or more media contents using a machine learning model, the one or more second contents being different from the one or more first contents; and generating the at least one video clip based on the one or more first contents and the one or more second contents.
[0109] In some embodiments, selecting the one or more second contents from the one or more media contents using the machine learning model comprises: determining a second total time length based on a difference between the preset target time length and the first total time length; and selecting the one or more second contents from the one or more media contents using the machine learning model, a time length of the one or more second contents not exceeding the second total time length.
[0110] In some embodiments, the target video clip is determined by: in response to a content selection indication for a target media content, presenting a marker of at least a portion of the target media content and a selection range, the selection range including a video clip in the target media content that is selected by default; determining an adjusted selection range based on a received adjustment indication of the selection range; and determining a video clip in the target media content that falls within the adjusted selection range as the target video clip.
[0111] In some embodiments, the configuration information indicates at least one of: a theme of the target video, or a style of the target video.
[0112] In some embodiments, obtaining the configuration information describing requirements for video generation comprises: presenting prompt information about a theme of a video; and determining, based on a received user input, the theme of the target video as at least a portion of the configuration information.
[0113] In some embodiments, obtaining the configuration information describing requirements for video generation comprises: presenting style information about at least one video style; receiving a selection of a video style from the at least one video style; and determining the selected video style as a style of the target video as at least a portion of the configuration information.
[0114] In some embodiments, the target video includes target music, and the method further comprises: determining the target music based on the style of the target video indicated by the configuration information.
[0115] In some embodiments, determining the target music comprises: in response to the configuration information indicating that the style of the target video is a preset video style, determining music matching the preset video style as the target music; or in response to the configuration information indicating that the style of the target video is a default video style, determining recommended music as the target music.
[0116] In some embodiments, the audio clip comprises audio of reading the text clip in the target voice.
[0117] In some embodiments, the process 500 further comprises determining the target voice by: in response to the configuration information indicating that the style of the target video is a preset video style, determining a voice matching the preset video style as the target voice; or in response to the configuration information indicating that the style of the target video is a default video style, determining a voice matching a theme of the target video as the target voice based on the theme of the target video indicated by the configuration information.
[0118] In some embodiments, the process 500 further comprises: presenting respective summary information of a plurality of clips in the target video, each clip in the plurality of clips comprising a video clip in the at least one video clip and a material clip corresponding to the video clip; receiving a replacement indication for a target clip in the plurality of clips and a selection of second media content, a time length of the second media content being greater than or equal to a time length of the target clip; and replacing the target clip with the second media content.
[0119] In some embodiments, the process 500 further comprises: presenting a text clip included in the at least one material clip; and receiving an edit of a selected text unit in the presented text clip based on a selection of the selected text unit to update the selected text unit.
[0120] In some embodiments, the target video comprises a plurality of scenes, the plurality of scenes respectively corresponding to a plurality of sub-material clips included in the at least one material clip, and a time length of each scene in the plurality of scenes being greater than or equal to a time length of a corresponding sub-material clip.
[0121] In some embodiments, the process 500 further comprises: receiving a regeneration indication for the target video; and in response to the regeneration indication, generating another target video based on a modification related to the target video.
[0122] In some embodiments, generating the another target video comprises: receiving at least one of a replacement indication, a deletion indication, or an addition indication for one or more media contents to determine another one or more media contents; and generating the another target video based on the another one or more media contents and the configuration information.
[0123] In some embodiments, generating the another target video comprises: receiving a modification of the configuration information of the target video; and generating the another target video based on the modified configuration information and the one or more media contents.
[0124] Example apparatus and device
[0125] FIG. 6 shows a schematic structural block diagram of a video generation apparatus 600 according to certain embodiments of the present disclosure. The apparatus 600 can be implemented as or included in the terminal device 110. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.
[0126] As shown, the apparatus 600 includes an indication receiving module 610 configured to receive a video generation indication for one or more media contents.
[0127] The apparatus 600 further includes a configuration information obtaining module 620 configured to obtain configuration information used to describe a video generation requirement.
[0128] The apparatus 600 further includes a segment generating module 630 configured to, in response to the video generation indication, generate at least one video segment and at least one material segment corresponding to the at least one video segment respectively based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment included in a target media content in the one or more media contents, and a material segment in the at least one material segment comprises a text segment and / or an audio segment of a video script generated for the corresponding video segment.
[0129] The apparatus 600 further includes a target video generating module 640 configured to generate the target video based on the at least one video segment and the at least one material segment.
[0130] In some embodiments, the apparatus 600 further includes a video segment generating module configured to: receive a selection of one or more first contents included in the one or more media contents; compare a first total time length of the one or more first contents with a preset target time length; in response to the first total time length not exceeding the preset target time length, select one or more second contents from the one or more media contents using a machine learning model, the one or more second contents being different from the one or more first contents; and generate the at least one video segment based on the one or more first contents and the one or more second contents.
[0131] In some embodiments, the apparatus 600 is further configured to: determine a second total time length based on a difference between the preset target time length and the first total time length; and select one or more second contents from the one or more media contents using the machine learning model, a time length of the one or more second contents not exceeding the second total time length.
[0132] In some embodiments, the apparatus 600 further includes a target video segment determination module configured to, in response to the content selection indication for the target media content, present a marker of at least a portion of the target media content and a selection range, the selection range including a video segment in the target media content that is selected by default; determine an adjusted selection range based on a received adjustment indication of the selection range; and determine a video segment in the target media content that falls within the adjusted selection range as the target video segment.
[0133] In some embodiments, the configuration information indicates at least one of: a theme of the target video, or a style of the target video.
[0134] In some embodiments, the configuration information acquisition module 620 is further configured to present prompt information about a theme of a video; and determine, based on a received user input, the theme of the target video as at least a portion of the configuration information.
[0135] In some embodiments, the configuration information acquisition module 620 is further configured to present style information about at least one video style; receive a selection of a video style from the at least one video style; and determine the selected video style as a style of the target video as at least a portion of the configuration information.
[0136] In some embodiments, the target video includes target music, and the apparatus 600 further includes a target music determination module configured to determine the target music based on the style of the target video indicated by the configuration information.
[0137] In some embodiments, the apparatus 600 is further configured to, in response to the configuration information indicating that the style of the target video is a preset video style, determine music matching the preset video style as the target music; or in response to the configuration information indicating that the style of the target video is a default video style, determine recommended music as the target music.
[0138] In some embodiments, the audio segment includes audio of reading the text segment in the target vocal tone.
[0139] In some embodiments, the apparatus 600 further includes a target vocal tone determination module configured to, in response to the configuration information indicating that the style of the target video is a preset video style, determine a vocal tone matching the preset video style as the target vocal tone; or in response to the configuration information indicating that the style of the target video is a default video style, determine a vocal tone matching the theme of the target video as the target vocal tone based on the theme of the target video indicated by the configuration information.
[0140] In some embodiments, the apparatus 600 further includes a replacement module configured to present respective summary information of a plurality of segments in the target video, each of the plurality of segments comprising a video segment in the at least one video segment and a material segment corresponding to the video segment; receive a replacement indication for a target segment in the plurality of segments and a selection of second media content, a time length of the second media content being greater than or equal to a time length of the target segment; and replace the target segment with the second media content.
[0141] In some embodiments, the apparatus 600 further includes a text update module configured to present a text segment included in the at least one material segment; and receive an edit of a selected text unit in the presented text segment based on a selection of the selected text unit to update the selected text unit.
[0142] In some embodiments, the target video comprises a plurality of scenes, the plurality of scenes respectively corresponding to a plurality of sub-material segments included in the at least one material segment, and a time length of each of the plurality of scenes being greater than or equal to a time length of a corresponding sub-material segment.
[0143] In some embodiments, the apparatus 600 further includes a regeneration module configured to receive a regeneration indication for the target video; and generate another target video based on a modification related to the target video in response to the regeneration indication.
[0144] In some embodiments, the apparatus 600 is further configured to receive at least one of a replacement indication, a deletion indication, or an addition indication for one or more media contents to determine another one or more media contents; and generate another target video based on the another one or more media contents and the configuration information.
[0145] In some embodiments, the apparatus 600 is further configured to receive a modification of the configuration information of the target video; and generate another target video based on the modified configuration information and the one or more media contents.
[0146] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting on the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can be used to implement the electronic device 110 of FIG. 1.
[0147] As shown in FIG. 7, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 can include, but are not limited to, one or more processors 710 or processing units, memory 720, storage 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processor 710 can be a real or virtual processor and capable of performing various processes in accordance with programs stored in memory 720. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 700.
[0148] Electronic device 700 typically includes a plurality of computer storage media. Such media can be volatile and / or nonvolatile storage media and removable and / or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Memory 720 can be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), EEPROM, flash memory), or some combination of the two. Storage 730 can be removable or non-removable and can include machine-readable media such as flash drives, magnetic disks, or any other medium that can be used to store information and / or data and that can be accessed by electronic device 700.
[0149] Electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile computer-readable medium. In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0150] Communication unit 740 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.
[0151] Input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 740, as desired, in order to communicate with one or more devices that enable a user to interact with electronic device 700, or to communicate with any device (e.g., a network card, a modem, etc.) that enables electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0152] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0153] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0154] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0155] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0156] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The machine-readable medium can be a single medium, or multiple media, of the same or different type. The computer program product can be one or more computer program components embodied in medium and / or transmission signals. The computer program product can have one or more computer program components embodied in medium and / or transmission signals.
[0157] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although the implementations of the disclosure have been described with regard to one or more implementations, it will be recognized that a variety of modifications and changes can be made to these implementations without departing from the broader spirit and scope of the implementations as set forth in the preceding disclosure. For example, certain aspects of the implementations can be performed using hardware, software, and / or firmware, or any combination thereof. The above-described implementations should therefore be regarded as merely illustrative, and not as narrowing the scope of the disclosure, which is defined by the appended claims and their equivalents.
Claims
1. A method for video generation, comprising: receiving a video generation indication for one or more media contents; obtaining configuration information for describing video generation requirements; generating, in response to the video generation indication, at least one video segment and at least one material segment respectively corresponding to the at least one video segment based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment included in a target media content in the one or more media contents, and a material segment in the at least one material segment comprises a text segment and / or an audio segment of a video script generated for a corresponding video segment; and generating a target video based on the at least one video segment and the at least one material segment.
2. The method of claim 1, wherein the at least one video segment is generated by: receiving a selection of one or more first contents included in the one or more media contents; comparing a first total duration of the one or more first contents with a preset target duration; in response to the first total duration not exceeding the preset target duration, selecting one or more second contents from the one or more media contents using a machine learning model, the one or more second contents being different from the one or more first contents; and generating the at least one video segment based on the one or more first contents and the one or more second contents.
3. The method of claim 2, wherein selecting one or more second contents from the one or more media contents using a machine learning model comprises: determining a second total duration based on a difference between the preset target duration and the first total duration; and selecting the one or more second contents from the one or more media contents using the machine learning model, the one or more second contents having a duration not exceeding the second total duration.
4. The method of claim 1, wherein the target video segment is determined by: in response to a content selection indication for the target media content, presenting a marker of at least a portion of the target media content and a selection range, the selection range including a video segment in the target media content that is selected by default; determining an adjusted selection range based on a received adjustment indication for the selection range; and determining a video segment in the target media content falling within the adjusted selection range as the target video segment.
5. The method of claim 1, wherein the configuration information indicates at least one of: a theme of the target video, or a style of the target video.
6. The method of claim 1, wherein obtaining configuration information for describing video generation requirements comprises: presenting prompt information about a video theme; and determining, based on a received user input, a theme of the target video as at least a portion of the configuration information.
7. The method of claim 1, wherein obtaining configuration information for describing video generation requirements comprises: presenting style information about at least one video style; and receiving a selection of a video style among the at least one video style; and determining the selected video style as a style of the target video as at least a part of the configuration information.
8. The method of claim 1, wherein the target video comprises target music, and the method further comprises: determining the target music based on the style of the target video indicated by the configuration information.
9. The method of claim 8, wherein determining the target music comprises: in response to the configuration information indicating that the style of the target video is a preset video style, determining music matching the preset video style as the target music; or in response to the configuration information indicating that the style of the target video is a default video style, determining recommended music as the target music.
10. The method of claim 1, wherein the audio clip comprises audio of reading the text clip in a target voice tone.
11. The method of claim 10, further comprising determining the target voice tone by: in response to the configuration information indicating that the style of the target video is a preset video style, determining a voice tone matching the preset video style as the target voice tone; or in response to the configuration information indicating that the style of the target video is a default video style, determining a voice tone matching a theme of the target video as the target voice tone based on the theme of the target video indicated by the configuration information.
12. The method of claim 1, further comprising: presenting respective summary information of a plurality of clips in the target video, each clip in the plurality of clips comprising a video clip among the at least one video clip and a material clip corresponding to the video clip; receiving a replacement indication for a target clip among the plurality of clips and a selection of second media content, a duration of the second media content being greater than or equal to a duration of the target clip; and replacing the target clip with the second media content.
13. The method of claim 1, further comprising: presenting a text clip comprising the at least one material clip; and receiving an edit of a selected text unit in the presented text clip based on a selection of the selected text unit to update the selected text unit.
14. The method of claim 1, wherein the target video comprises a plurality of scenes, the plurality of scenes respectively corresponding to a plurality of sub-material clips comprising the at least one material clip, and a duration of each scene in the plurality of scenes being greater than or equal to a duration of a corresponding sub-material clip.
15. The method of claim 1, further comprising: receiving a regeneration indication for the target video; and in response to the regeneration indication, generating another target video based on modifications related to the target video.
16. The method of claim 15, wherein generating another target video comprises: receiving at least one of a replacement indication, a deletion indication, or an addition indication for the one or more media content to determine another one or more media content; and generating the other target video based on the one or more media contents and the configuration information.
17. The method of claim 15, wherein generating the other target video comprises: receiving a modification of the configuration information of the target video; and generating the other target video based on the modified configuration information and the one or more media contents.
18. A video generation apparatus, comprising: an indication receiving module configured to receive a video generation indication for one or more media contents; a configuration information obtaining module configured to obtain configuration information describing video generation requirements; a segment generating module configured to, in response to the video generation indication, generate at least one video segment and at least one material segment corresponding to the at least one video segment respectively based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment included in a target media content in the one or more media contents, and a material segment in the at least one material segment comprises a text segment and / or an audio segment of a video script generated for a corresponding video segment; and a target video generating module configured to generate a target video based on the at least one video segment and the at least one material segment.
19. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-17.
20. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method according to any one of claims 1-17.
21. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1-17.
Citation Information
Patent Citations
System and method for automatically generating complete video based on fragmented video
CN110740379A
Method and device for generating mixed video, computer equipment and medium
CN117176981A
Video generation method and device, computer equipment, storage medium and program product
CN117729296A
Video generation method and device, equipment and storage medium
CN117979088A
Video generation method and device, equipment and storage medium
CN118509666A