Video generation method and device, equipment and storage medium

CN121750893APending Publication Date: 2026-03-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2026-03-27

Smart Images

  • Figure CN121750893A_ABST
    Figure CN121750893A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and device, equipment and a storage medium. The method includes receiving a video generation indication for one or more media content; acquiring configuration information used for describing video generation requirements; in response to the video generation instruction, generating at least one video clip and at least one material clip respectively corresponding to the at least one video clip based on the one or more media contents and the configuration information, a target video clip in the at least one video clip is a video clip contained in a target media content in the one or more media contents, and a material clip in the at least one material clip comprises a text clip and / or an audio clip of a video copywriting generated for the corresponding video clip; and generating a target video based on the at least one video clip and the at least one material clip. Therefore, one or more media contents are edited into the target video according to the configuration information, and a user can be helped to create a rich and interesting video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to video generation methods, apparatus, devices, and computer-readable storage media. Background Technology

[0002] With the rapid development of computer technology, various short and long videos have become important mediums for people to share content and express their opinions. Based on this, people's demands for the richness and entertainment value of video content are also increasing. Therefore, people expect to produce higher-quality videos and, at the same time, hope to complete video creation more efficiently. Summary of the Invention

[0003] In a first aspect of this disclosure, a video generation method is provided. The method includes: receiving a video generation instruction for one or more media contents; obtaining configuration information describing video generation requirements; in response to the video generation instruction, generating at least one video segment and at least one material segment corresponding to each of the at least one video segment, based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment contained within the target media content of the one or more media contents, and the material segment in the at least one material segment includes a text segment and / or an audio segment of video script generated for the corresponding video segment; and generating a target video based on the at least one video segment and the at least one material segment.

[0004] In a second aspect of this disclosure, a video generation apparatus is provided. The apparatus includes: an instruction receiving module configured to receive a video generation instruction for one or more media contents; a configuration information acquisition module configured to acquire configuration information describing video generation requirements; a segment generation module configured to, in response to the video generation instruction, generate at least one video segment and at least one material segment corresponding to each of the at least one video segment based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment contained in the target media content of the one or more media contents, and the material segment in the at least one material segment includes a text segment and / or an audio segment of video script generated for the corresponding video segment; and a target video generation module configured to generate a target video based on the at least one video segment and the at least one material segment.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0010] Figures 2A to 2O A schematic diagram of an example interface for video generation according to some embodiments of the present disclosure is shown;

[0011] Figure 3 A schematic diagram of example timings using a machine learning model to assist in the selection of highlight content according to some embodiments of the present disclosure is shown;

[0012] Figures 4A to 4C A schematic diagram of example 400A showing a scenario and a corresponding subset of explanatory information according to some embodiments of the present disclosure is shown;

[0013] Figure 5 A flowchart of a video generation process according to some embodiments of the present disclosure is shown;

[0014] Figure 6 A schematic structural block diagram of a video generation apparatus according to certain embodiments of the present disclosure is shown; and

[0015] Figure 7 A block diagram illustrating an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0022] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0023] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0024] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.

[0026] As used herein, the term "component" can refer to any suitable model, module, unit, etc., used to achieve special effects. Such a component can provide corresponding outputs based on provided inputs and can include any suitable operations, calculations, etc. An example of a component is an algorithm. While some embodiments of this disclosure will be described below primarily with reference to algorithms, it should be understood that such embodiments are also applicable to other types of components.

[0027] As briefly mentioned above, people have increasingly higher demands for the richness and entertainment value of video content, and they also expect video production to be more efficient. However, currently, the process of integrating media materials in video production is quite challenging, and the finished content is often monotonous and uninteresting, making it difficult to fully meet user needs and impacting the user experience.

[0028] In view of this, embodiments of the present disclosure propose an improved scheme for video generation. According to various embodiments of the present disclosure, a video generation instruction for one or more media content is received. Configuration information describing the video generation requirements is obtained. In response to the video generation instruction, based on one or more media content and the configuration information, at least one video segment and at least one material segment corresponding to each of the at least one video segment are generated, wherein the target video segment in the at least one video segment is a video segment contained in the target media content of one or more media content, and the material segment in the at least one material segment includes text segments and / or audio segments of video script generated for the corresponding video segment. A target video is generated based on the at least one video segment and the at least one material segment. Thus, one or more media content are edited into the target video according to the configuration information. In this way, users can create rich and interesting videos. This can improve the richness and interest of the generated videos, effectively improve the quality of the generated videos, and enable users to generate videos conveniently and quickly, thus improving video generation efficiency.

[0029] Example Environment

[0030] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed on an electronic device 110. A user 140 can interact with the application 120 via the electronic device 110 and / or an attached device of the electronic device 110.

[0031] In some embodiments, application 120 may be a content sharing application, a content editing application, a content creation application, etc. Application 120 can provide user 140 with various services related to media content (also referred to as media content items, content items, media items, etc.), including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), publishing, etc. of media content.

[0032] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can display the interface 150 of application 120. Interface 150 may include various interfaces provided by application 120, such as media content presentation interfaces, media content creation interfaces, media content publishing interfaces, and so on. Application 120 can provide media content editing functions (e.g., application 120 may be a video editing application) to support editing (e.g., clipping) media content within application 120.

[0033] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0034] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0035] The following description continues with reference to the accompanying drawings, which will continue to illustrate some exemplary embodiments of this disclosure. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The various graphic elements on the page may have different arrangements and different visual representations, one or more elements may be omitted or replaced, and one or more other elements may also be present. The embodiments of this disclosure are not limited in this respect. Furthermore, in the following description, exemplary embodiments will be primarily described with respect to electronic device 110. It should be understood that the actions described with respect to electronic device 110 may be performed by application 120 on electronic device 110, or may be performed by application 120 in conjunction with its server (e.g., server 130).

[0036] Example Interaction

[0037] The following description, with reference to the accompanying drawings, illustrates an example interaction process according to an embodiment of the present disclosure.

[0038] Figures 2A to 2O Example interfaces 200A to 200O generated from video are shown according to some embodiments of the present disclosure. In embodiments of the present disclosure, interfaces 200A to 200O can be generated by... Figure 1 The electronic device 110 shown is provided.

[0039] like Figure 2A-2CAs shown, electronic device 110 receives a video generation instruction for one or more media contents. The media contents discussed herein include, but are not limited to, videos, images, and image sets. In some examples, the one or more media contents corresponding to the video generation instruction may include one or more of the following: videos, images, and image sets. The specific manner in which electronic device 110 receives the video generation instruction will be discussed in detail below.

[0040] refer to Figure 2A In some embodiments, electronic device 110 may present an entry point 210 for video generation, which may be associated with a video generation instruction. In interface 200A, entry point 210 may be arranged alongside entry points for other media content services (such as image editing, video effects, etc.) provided by electronic device 110. It should be understood that the text label (e.g., "Video with Narration"), style, and position of entry point 210 presented on electronic device 110 may not be limited to those presented in interface 200A.

[0041] refer to Figure 2B In some embodiments, entry point 210 may be configured to receive a preset operation from user 140 and present media content selection interface 200B. Preset operations discussed herein may include, for example, click operations, long press operations, swipe operations, etc., and are not limited thereto. In some embodiments, media content selection interface 200B may present at least one media content for user 140 to select one or more media content from.

[0042] In some embodiments, the media content selection interface 200B may include multiple media item labels, each corresponding to a different media item. These multiple media item labels may include, but are not limited to, a video label 220 corresponding to a video media item, a photo label corresponding to a photo media item, and a live label corresponding to a live photo (i.e., moving photo) media item. By distinguishing different media items, users can select media content more conveniently and quickly. In some embodiments, the media content selection interface 200B may display one media item by default (e.g., ...). Figure 2B The media content of the video media item shown can be displayed, and the corresponding media content can be displayed by accepting the selection of other media item tags.

[0043] In some embodiments, the media content selection interface 200B may include a region 222 configured to display thumbnails of the selected media content. Exemplarily, the electronic device 110 may sequentially number the selected media content according to the selection order. It should be understood that media content selected on the media content selection interface 200B can be deselected. For example, user 140 can deselect the selected media content by clicking it again, or by using the "-" marker corresponding to each media content thumbnail in region 222. In some embodiments, the order of the multiple media content thumbnails in region 222 can be changed by dragging by user 140.

[0044] In some embodiments, for one or more media content, at least one of the following can be satisfied: the total duration of one or more media content can not exceed a preset total duration (e.g., 10 minutes or other preset total duration); the duration of each media content in one or more media content can be within a preset duration range (e.g., a duration range of 1 second to 10 minutes, or other preset duration range); and the number of one or more media content can not exceed a preset number (e.g., 50 or other number). It should be understood that for one or more media content, other requirements can be set according to actual circumstances, and are not limited here. Furthermore, when these requirements are not met, corresponding prompts can be displayed on the media content selection interface 200B to remind the user 140.

[0045] refer to Figure 2C In some embodiments, the electronic device 110 may present a guide page 230 to a user who clicks on the entry point 210 for the first time (referred to as a new user for ease of discussion). The guide page 230 in the interface 200C may be a separate interface, window, panel, or a specific display area within the interface. For example, the guide page 230 may be used to display a guide example of a specific operation for generating a video on the electronic device 110 according to embodiments of this disclosure. The guide page 230 may also be used to display the web address of the guide example, so as to jump to another page for presenting the guide example based on the new user's click operation on the web address.

[0046] In some embodiments, the interface 200C may include an entry point 235, configured to be presented based on a preset action (including but not limited to a click action) received from a new user. Figure 2B The media content selection interface 200B is provided. This allows new users to quickly understand the interactive operations used for video generation, improving interaction efficiency and enhancing user experience. Furthermore, the electronic device 110 acquires configuration information describing the video generation requirements. The configuration information may include configuration information corresponding to the target video to be generated.

[0047] In some embodiments, the configuration information may indicate at least one of the following: the theme of the target video, or the style of the target video. In other embodiments, the configuration information may also indicate other information such as the visual effects, screen effects, etc. of the target video.

[0048] In some embodiments, in order to obtain the theme of a target video, the electronic device 110 may present prompts about the video theme and determine the theme of the target video based on received user input as at least part of the configuration information.

[0049] Combination Figure 2B and 2D Here's an example. After receiving a preset operation (including but not limited to clicks) from user 140 on control 224 in interface 200B, electronic device 110 can display configuration panel 240 on interface 200D. Configuration panel 240 may include prompt information 242 about the video theme and input control 244. Based on prompt information 242 and input control 244, user 140 can input the theme of the target video to be generated into input control 244. In some embodiments, input control 244 may support text input and voice input, as well as any other suitable type of input method. In some embodiments, the video theme needs to meet preset requirements, including but not limited to character limits and compliance requirements for the theme content.

[0050] In some embodiments, in order to obtain the style of a target video, the electronic device 110 may present style information about at least one video style. The electronic device 110 may receive a selection of a video style from at least one video style and determine the selected video style as the style of the target video as at least part of the configuration information.

[0051] Continue to refer to Figure 2D The configuration panel 240 may also include style information. Style information may include information corresponding to at least one video style, such as style name 246, style image 248, etc. At least one video style may include a default video style and several preset video styles. Several preset video styles may include, for example, any appropriate style such as the "Everything Can Be Rap" style, the "Daily Life with a Baby" style, etc., without limitation. It should be understood that if all style images 248 are disabled, or if the electronic device 110 does not have any style images 248 configured, style information may not be displayed on the configuration panel 240.

[0052] In some embodiments, the style image can be an animated graphic, which can be used to play dynamic visuals in response to preset user actions 140. This allows users to quickly view the style's display effect, enhancing the interactive experience.

[0053] Furthermore, if a video generation instruction is received, the electronic device 110 generates at least one video segment based on one or more media contents and configuration information. The target video segment in the at least one video segment is a video segment contained in the target media content of one or more media contents.

[0054] Here, one or more video segments can be selected for each of the one or more media contents. Specifically, for a target media content (or any media content) within one or more media contents, its corresponding target video segment may include one or more video segments. Such target video segments may include, for example, important segments or scenes from the target media content, and may also be simply referred to as highlight segments. In other embodiments, the target video segment of the target media content may also include content with other meanings, which is not limited here.

[0055] In some embodiments, at least one video segment may be generated by the following method: Electronic device 110 may receive a selection of one or more first contents included in one or more media contents, and compare a first total duration of the one or more first contents with a preset target duration. If the first total duration does not exceed the preset target duration, electronic device 110 may use a machine learning model to select one or more second contents from the one or more media contents, the one or more second contents being different from the one or more first contents. Then, electronic device 110 may generate at least one video segment based on the one or more first contents and the one or more second contents. Depending on the type of media content, the first contents and second contents described herein may have different types. For example, the first contents and / or the second contents may be segments of video, one or more images, etc.

[0056] In such embodiments, one or more first contents may be highlight content selected by user 140 from one or more media contents. In some embodiments, the preset target duration may be determined based on the total duration of one or more media contents multiplied by a preset ratio. It should be understood that the preset target duration can also be determined in any suitable manner, and there is no limitation herein. In addition, the type of machine learning model is not limited in the embodiments of this disclosure. For example, the machine learning model herein may include one or more of other types of machine learning models such as deep learning models, supervised learning models, or unsupervised learning models.

[0057] In some embodiments, if the total duration of one or more first contents selected by user 140 exceeds a preset target duration, a corresponding prompt message can be issued to remind user 140. In some embodiments, if the duration of one or more media contents is less than a preset lower limit duration, user 140 can be prevented from selecting highlight content from one or more media contents.

[0058] In some embodiments, in order to select one or more second content items, the electronic device 110 may determine a second total duration based on the difference between a preset target duration and a first total duration, and utilize a machine learning model to select one or more second content items from one or more media content items whose duration does not exceed the second total duration. The following will combine... Figure 3 Let's discuss in detail the specific ways to select one or more second contents using machine learning models.

[0059] Figure 3 A schematic diagram of an example timing sequence 300 using a machine learning model to assist in the selection of highlight content according to some embodiments of the present disclosure is shown. In timing sequence 300, assuming the total duration of one or more media contents 310 is 1 minute, and the highlight content selected by user 140 includes content 322, content 324, and content 326 (which can be collectively referred to as highlight content 320), the machine learning model can identify highlight content 320 and retain it when performing semantic segmentation on one or more media contents 310. Then, the machine learning model can select other highlight content from one or more media contents 310 whose duration does not exceed a second total duration (i.e., a preset target duration minus a first total duration) as one or more second contents. Thus, the electronic device 110 can merge one or more first contents and one or more second contents to generate at least one video segment.

[0060] It should be understood that upon receiving an instruction not to manually select highlight content, the highlight content can be directly selected using a machine learning model based on the user 140's authorization. As an example, the first content could be a video clip or an image (e.g., including pictures or photos). It should be noted that if user 140 selects a separate image or photo as part of the highlight content 320, the electronic device 110 can convert such content into a 3-second or other shorter video clip by default. Therefore, this embodiment effectively improves interaction efficiency by utilizing a machine learning model to assist in selecting highlight content.

[0061] In some embodiments, the target video segment can be determined as follows: if the electronic device 110 receives a content selection instruction for the target media content, it can present at least a portion of the target media content and a selection range flag, the selection range including the video segment in the target media content that is selected by default. Based on the received adjustment instruction for the selection range, the electronic device 110 can determine the adjusted selection range and identify the video segment in the target media content that falls within the adjusted selection range as the target video segment.

[0062] Combination Figure 2B and Figure 2E ,exist Figure 2B The media content selection interface 200B, based on the media content selected by the user 140, allows the electronic device 110 to receive preset operations from the user 140 on the control 226 (e.g., a control labeled "highlight mark"), thereby responding to content selection instructions for the target media content and subsequently displaying... Figure 2E The highlight marking interface 200E. In the highlight marking interface 200E, the selection box 250 can be used to define the selection area of ​​the image. The left and right edges of the selection box 250 can serve as markers for the selection area.

[0063] In some embodiments, the selection box 250 appears by default at the beginning of the video, and the selection box 250 may have a default width (i.e., a default selection range) corresponding to the video segment selected by default. For example, the default width can be determined based on the duration of the target media content. For instance, if the duration of the target media content exceeds 2 seconds or other durations, the default width of the selection box 250 may be half the timeline. If the duration of the target media content is less than 2 seconds, the selection box 250 may default to a width of 1 second. In some embodiments, the selection box 250 may accept dragging by the user 140 to adjust its width, i.e., adjust the selection range.

[0064] Refer to interfaces 200F to 200I for a detailed discussion of adjusting the selection range based on selection box 250. In interface 200F, user 140 clicks the "Mark Highlight" button 255 to mark the currently selected segment 262. Segment 262 can be covered with a color different from selection box 250. Subsequently, in interface 200G, selection box 250 appears on the right, with a width that can be half the width of the remaining segment on the right or another proportion. If the duration of the remaining segment on the right is less than 2 seconds or other durations, it appears at half the width of the remaining segment on the left or another proportion. It should be noted that highlight marks cannot be added repeatedly within the same time period.

[0065] In interface 200H, clicking on the previously marked highlight segment 262 will bring up the selection box 250 again. Clicking the "Cancel Highlight" button 258 will then cancel the highlight marking on the current segment 262. The color of segment 262 after the highlight marking is removed can be the color of the selection box 250, such as as shown in interface 200I. In this embodiment, dragging the edge of the selection box 250 adjusts the duration; dragging lengthens the segment with the marked highlight. Additionally, dragging the selection box 250 on the timeline pauses the video. Clicking the play button resumes playback from the pointer's position. Selecting and marking a highlight, then deselecting it, discards the highlight segment data, eliminating the need for memorization.

[0066] Based on the above embodiments, by selecting highlight content for media content, the richness and interest of the generated target video can be improved, effectively enhancing the quality of the generated video. Through the specific selection methods described above, the efficiency of highlight content selection can be increased, thereby improving video generation efficiency while ensuring video quality.

[0067] Additionally, upon receiving a video generation instruction, the electronic device 110 generates at least one media segment corresponding to at least one video segment, based on one or more media contents and configuration information. The media segment in the at least one media segment includes a text segment and / or an audio segment of video script generated for the corresponding video segment. Further, based on the at least one video segment and the at least one media segment, the electronic device 110 generates a target video. For example, the media segment may include narration or voice-over information for the corresponding video segment. For instance, the text segment may be the text of the voice-over, and the audio segment may be the sound of the voice-over.

[0068] In some embodiments, the audio segment may include audio of a text segment read aloud in a target voice. In some embodiments, the text segment may include text presented on the target video in accordance with the scene of the target video; such text may be referred to as subtitles. At least one material segment may also be referred to as narration content. The target voice may include, for example, a male voice, a female voice, a child's voice, a deep voice, an electronic voice, etc., without limitation. In some embodiments, the electronic device 110 may set an appropriate speech rate for the audio segment. In addition, an upper limit and a lower limit for the speech rate may be set for the user 140 to select.

[0069] In some embodiments, the target timbre can be determined as follows: if the configuration information indicates that the style of the target video is the default video style, the electronic device 110 can determine a timbre that matches the theme of the target video as the target timbre based on the theme of the target video indicated by the configuration information. This allows the target timbre to still be adapted to the target video even with the default video style.

[0070] In some embodiments, if the configuration information indicates that the style of the target video is a preset video style, the electronic device 110 can determine a timbre that matches the preset video style as the target timbre. For example, if the preset video style of the target video is "daily life with a child", the matching timbre can be set to a female timbre.

[0071] In some embodiments, the target video may include multiple scenes, each scene corresponding to a sub-scene of at least one footage segment. In other words, each scene needs to be aligned with its respective text and audio sub-scenes. The alignment of scenes with their corresponding sub-scenes will be discussed in detail below with reference to Figure 4.

[0072] Figure 4A A schematic diagram of example 400A of a scene and corresponding sub-material fragments according to some embodiments of the present disclosure is shown. Example 400A includes examples 410 and 420, where example 410 may be an example where the scene and the corresponding sub-material fragment are not perfectly aligned, and example 420 may be an example of a solution to example 410.

[0073] In Example 410, scene 412 and its corresponding sub-segment 414 (including text sub-segment 414-1 and audio sub-segment 414-2) are not perfectly aligned in time. It can be seen that sub-segment 414 spans two scenes. In this case, following the solution in Example 420, text sub-segment 414-1 and audio sub-segment 414-2 need to be cut according to their timestamps to ensure alignment with scene 412. It should be noted that since the narration audio can be generated based on the narration text, the audio sub-segment and its corresponding text sub-segment are related and can be aligned in timestamps.

[0074] Additionally, it should be noted that video clips can retain the original clips uploaded by user 140. For example, if the first original media content uploaded by user 140 is 10 seconds long, after being split and highlight content extracted, it becomes two 5-second clips. For clips that can still be seen during later editing (such as replacement), only the 5-second clip can be selected during replacement.

[0075] In some embodiments, the duration of each scene in a plurality of scenes can be greater than or equal to the duration of its corresponding sub-scene. In some embodiments, if the duration of a scene is less than the duration of its corresponding sub-scene, the duration of several frames in that scene can be increased. In other embodiments, the sub-scene corresponding to the scene can be regenerated to match the two. The following will combine... Figure 4B and 4C Let's discuss these two situations in detail.

[0076] Figure 4B A schematic diagram of example 400B, showing a scene and its corresponding sub-video clip according to some embodiments of the present disclosure, is shown. Example 400B includes examples 430, 440, and 450. Example 430 may be an example where the scene and its corresponding sub-video clip are perfectly aligned. Example 440 may be an example where the duration of the scene is shorter than the duration of the corresponding sub-video clip. Example 450 may be an example where the duration of the scene is longer than the duration of the corresponding sub-video clip.

[0077] In Example 440, if the duration of sub-segment 444 exceeds the duration of the corresponding scene 442, the video media content can be automatically lengthened to adapt to the duration. If lengthening the media content is still not enough to fill the time, the last frame (freeze frame) can be lengthened to fill the time.

[0078] In Example 450, if the duration of sub-video clip 446 is less than the duration of the corresponding scene 442, and there are no stop frames, then the duration of scene 442 does not need to be adjusted. If there are stop frames, then it is adapted to the duration of sub-video clip 446.

[0079] In other embodiments, if the narration text of the target video remains unchanged, but the target timbre of the narration audio changes, the duration of the text sub-segments is adjusted to match the duration of the corresponding audio sub-segments. In this case, if the duration of the audio sub-segment exceeds the duration of the scene, the scene visuals need to be readjusted as discussed above.

[0080] Figure 4C A schematic diagram of example 400C, illustrating a scene and its corresponding sub-video clip according to some embodiments of the present disclosure, is shown. Example 400C includes example 430 from example 400B and example 460 for representing the regeneration of sub-video clips corresponding to the scene. In such an example, if the target video regenerates at least one video clip, a machine learning model can be used to re-adapt the duration of the sub-video clips and the corresponding scene for re-alignment.

[0081] In some embodiments, the electronic device 110 can receive editing of the media content and at least one clip of the generated target video. This editing may include, for example, replacing the media content, modifying at least one clip, etc. Specific embodiments regarding this editing will be discussed in detail below.

[0082] In some embodiments, the electronic device 110 may present corresponding summary information of multiple segments in a target video, each of the multiple segments including a video segment from at least one video segment and a corresponding material segment. The electronic device 110 may receive a replacement instruction for a target segment among the multiple segments and a selection of second media content, and replace the target segment with the second media content. The duration of the second media content may be greater than or equal to the duration of the target segment. In some embodiments, the summary information of the segment may include the number of frames in the segment and the duration of the segment.

[0083] refer to Figure 2J-Figure 2K On interface 200J, multiple segments of the target video can be displayed. For the located segment, a playback indicator can be displayed to indicate the playback status. Playback can be stopped for the selected segment, and the target video display will then be positioned on the selected segment. In the paused state, electronic device 110 can receive a preset operation from user 140 on the replace button 272, thus presenting interface 200K. On interface 200K, user 140 can select media content to replace the segment. It should be noted that only media content longer than the segment to be replaced can be selected. If the media content is shorter than the segment to be replaced, it will be unselectable. For media content that meets the requirements, user 140 can directly replace the original segment by clicking on one of the media pieces.

[0084] Additionally, interface 200J may include a crop button 274. For the selected segment, electronic device 110 can receive a user's preset operation on the crop button 274, thereby displaying interface 200L. The user can adjust the screen ratio on interface 200L. Interface 200L includes a selection box 276, which allows selection of multiple frames corresponding to a fixed duration, enabling the user 140 to crop the selected frames.

[0085] In some embodiments, the electronic device 110 may present text segments included in at least one material segment. Based on the selection of text units in the presented text segment, the electronic device 110 may receive editing of the selected text units to update the selected text units.

[0086] refer to Figure 2MThe interface 200M may include a subtitle panel 280. The subtitles (i.e., narration text) on the subtitle panel 280 can be displayed in segments by timestamp (or in other forms, which are not limited here) to form text fragments. In the segmented display case, each text unit may include one or more subtitles. The electronic device 110 can highlight the currently playing subtitle and supports scrolling the subtitles up and down, with the highlighted area automatically positioned on the first subtitle segment during scrolling. The electronic device 110 can receive user 140's click operation on the subtitles to trigger a subtitle editing state. The subtitle panel 280 may include an edit button 282, which can also trigger a subtitle editing state by receiving a preset operation from user 140. In the subtitle editing state, the electronic device 110 can receive edits to the subtitles, thereby updating the subtitles. It should be noted that single subtitles on the subtitle panel cannot be directly deleted.

[0087] In some embodiments, after the subtitles are updated, the electronic device 110 can synchronously update the audio according to the new subtitles. This helps to improve the quality of the generated target video. Based on the above embodiments, the electronic device 110 can improve interactivity by editing and modifying the media content and at least one material segment of the target video, making it easier and faster for users to generate target videos that meet their needs.

[0088] In some embodiments, the target video may include target music. Target music may include, for example, but not limited to, the background music of the target video. In some embodiments, the electronic device 110 may determine the target music based on the style of the target video indicated by configuration information.

[0089] In some embodiments, if the configuration information indicates that the target video's style is a preset video style, the electronic device 110 can determine music that matches the preset video style as the target music. In some embodiments, if the configuration information indicates that the target video's style is a default video style, the electronic device 110 can determine recommended music as the target music. As an example, the electronic device 110 can utilize a machine learning model to recommend music. Exemplarily, the target music can be selected from a preset music library. It should be understood that the music in the music library is copyrighted music.

[0090] In some embodiments, the target music can be automatically adapted to the length of the target video. If the length of the target music is longer than the length of the target video, the end of the target music can be truncated. If the length of the target music is shorter than the length of the target video, the target music can be looped until the length of the target video is reached. It should be noted that if a machine learning model is used to match the target music to the target video, the length of the music should be greater than or equal to the length of the target video.

[0091] Return to reference Figure 2DThe interface 200D also includes a "Generate Video" button 245. After receiving a preset operation from the user 140 on button 245, the electronic device 110 enters the video loading page. The loading page prompt text (the loading progress is estimated by the server 130 to the electronic device 110) can display phrases such as "Parsing media content...x%", "Generating text...x%", "Intelligent dubbing...x%", "Intelligent background music...x%", "Almost done...x%", etc. The video loading page can include a "View Later" button, allowing the user 140 to click the "View Later" button to load the video generation task offline and exit the loading page. The video loading page can also include a "Cancel" button, allowing the user 140 to click the "Cancel" button to terminate the task and return to the interface 200D. In addition, if the target video generation fails (the reasons may include, for example, non-compliant video theme, network abnormality, etc.), corresponding prompt information can be displayed on the loading page.

[0092] In some embodiments, for a generated video, the electronic device 110 may display corresponding draft content. The draft content may include, for example, the video's cover image, video theme, style information, etc. Additionally, the draft content may be deleted.

[0093] Return to reference Figure 2M For the generated target video, the electronic device 110 can receive the user 140's preset operation on the "Export" button 284 to display the export panel. On the export panel, the target video can be exported directly, or it can be exported and shared to other applications or shared with other users, etc. Additionally, the electronic device 110 can receive the user 140's settings for the target video's image quality, such as resolution settings, etc.

[0094] In some embodiments, electronic device 110 may receive a regeneration instruction for a target video. If electronic device 110 receives a regeneration instruction, it may generate another target video based on modifications associated with the target video. Specific embodiments of electronic device 110 generating another target video will be discussed in detail below.

[0095] In some embodiments, in order to generate another target video, the electronic device 110 may receive at least one of a replacement instruction, a deletion instruction, or an addition instruction for one or more media contents to determine another one or more media contents. Based on the other one or more media contents and configuration information, the electronic device 110 may generate another target video.

[0096] refer to Figure 2NThe video regeneration interface 200N may include controls for modifying the target video, including but not limited to at least one of the following: replacement control, deletion control, and addition control 292. Taking the addition control 292 as an example, the electronic device 110 can receive a preset operation from the user 140 on the addition control 292, thereby obtaining another target video identical to the target video. Based on modifications to one or more media contents and configuration information in the other target video, the electronic device 110 can generate another target video.

[0097] In some embodiments, in order to generate another target video, the electronic device 110 may also receive modifications to the configuration information of the target video, and generate another target video based on the modified configuration information and one or more media contents.

[0098] Combination Figure 2N-Figure 2O The video regeneration interface 200N may also include a configuration information modification control 294, used to receive preset operations from user 140 to present the configuration information modification interface 200O. The configuration information modification interface 200O may include information related to the video theme and video style of the other target video to be generated. After re-entering the video theme and re-selecting the video style, user 140 can click the regeneration control 296 to generate another target video.

[0099] In summary, the embodiments of this disclosure enable the editing of one or more media contents into a target video based on configuration information. This helps users create rich and engaging videos. It increases the richness and appeal of the generated videos, effectively improves the quality of the generated videos, and allows users to generate videos conveniently and quickly, thus increasing video generation efficiency.

[0100] Example process

[0101] Figure 5 A flowchart of a video generation process 500 according to some embodiments of the present disclosure is shown. Process 500 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 500.

[0102] In box 510, electronic device 110 receives a video generation instruction for one or more media contents.

[0103] In box 520, electronic device 110 obtains configuration information describing the video generation requirements.

[0104] In box 530, in response to a video generation instruction, electronic device 110 generates at least one video segment and at least one material segment corresponding to the at least one video segment, based on one or more media content and configuration information, wherein the target video segment in the at least one video segment is a video segment contained in the target media content of one or more media content, and the material segment in the at least one material segment includes a text segment and / or an audio segment of video script generated for the corresponding video segment.

[0105] In box 540, electronic device 110 generates a target video based on at least one video clip and at least one source clip.

[0106] In some embodiments, at least one video segment is generated by: receiving a selection of one or more first contents included in one or more media contents; comparing a first total duration of one or more first contents with a preset target duration; in response to the first total duration not exceeding the preset target duration, selecting one or more second contents from one or more media contents using a machine learning model, wherein the one or more second contents are different from one or more first contents; and generating at least one video segment based on one or more first contents and one or more second contents.

[0107] In some embodiments, selecting one or more second contents from one or more media contents using a machine learning model includes: determining a second total duration based on the difference between the preset target duration and the first total duration; and selecting one or more second contents from the one or more media contents with a duration not exceeding the second total duration using the machine learning model.

[0108] In some embodiments, the target video segment is determined by: in response to a content selection instruction for the target media content, presenting at least a portion of the target media content and a selection range, the selection range including video segments in the target media content that are selected by default; determining an adjusted selection range based on a received adjustment instruction for the selection range; and identifying video segments in the target media content that fall within the adjusted selection range as the target video segment.

[0109] In some embodiments, the configuration information indicates at least one of the following: the theme of the target video, or the style of the target video.

[0110] In some embodiments, obtaining configuration information describing video generation requirements includes: presenting prompts about the video topic; and determining the topic of the target video as at least part of the configuration information based on received user input.

[0111] In some embodiments, obtaining configuration information describing video generation requirements includes: presenting style information about at least one video style; receiving a selection of a video style from at least one video style; and determining the selected video style as the style of the target video as at least a part of the configuration information.

[0112] In some embodiments, the target video includes target music, and the method further includes: determining the target music based on the style of the target video indicated by configuration information.

[0113] In some embodiments, determining the target music includes: in response to configuration information indicating that the style of the target video is a preset video style, determining music that matches the preset video style as the target music; or in response to configuration information indicating that the style of the target video is a default video style, determining recommended music as the target music.

[0114] In some embodiments, the audio segment includes audio of a text segment read aloud in a targeted tone.

[0115] In some embodiments, process 500 further includes determining the target timbre by: in response to configuration information indicating that the style of the target video is a preset video style, determining a timbre that matches the preset video style as the target timbre; or in response to configuration information indicating that the style of the target video is a default video style, determining a timbre that matches the theme of the target video as the target timbre based on the theme of the target video indicated by the configuration information.

[0116] In some embodiments, process 500 further includes: presenting corresponding summary information of a plurality of segments in a target video, each of the plurality of segments including a video segment in at least one video segment and a material segment corresponding to the video segment; receiving a replacement instruction for a target segment in the plurality of segments and a selection of second media content, the duration of the second media content being greater than or equal to the duration of the target segment; and replacing the target segment with the second media content.

[0117] In some embodiments, process 500 further includes: presenting a text segment included in at least one material segment; and receiving an edit of a selected text unit based on the selection of a text unit in the presented text segment to update the selected text unit.

[0118] In some embodiments, the target video includes multiple scenes, each scene corresponding to multiple sub-segments included in at least one footage segment, and the duration of each scene is greater than or equal to the duration of the corresponding sub-segment.

[0119] In some embodiments, process 500 further includes: receiving a regeneration instruction for the target video; and in response to the regeneration instruction, generating another target video based on modifications associated with the target video.

[0120] In some embodiments, generating another target video includes: receiving at least one of a replacement instruction, a deletion instruction, or an addition instruction for one or more media contents to determine another one or more media contents; and generating another target video based on the other one or more media contents and configuration information.

[0121] In some embodiments, generating another target video includes: receiving modifications to configuration information of the target video; and generating another target video based on the modified configuration information and one or more media contents.

[0122] Example devices and equipment

[0123] Figure 6 A schematic structural block diagram of a video generation apparatus 600 according to certain embodiments of the present disclosure is shown. The apparatus 600 may be implemented as or included in the terminal device 110. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0124] As shown in the figure, the device 600 includes an instruction receiving module 610 configured to receive a video generation instruction for one or more media contents.

[0125] The device 600 also includes a configuration information acquisition module 620, configured to acquire configuration information describing the video generation requirements.

[0126] The apparatus 600 also includes a segment generation module 630 configured to generate at least one video segment and at least one material segment corresponding to the at least one video segment based on one or more media content and configuration information in response to a video generation instruction, wherein the target video segment in the at least one video segment is a video segment contained in the target media content of one or more media content, and the material segment in the at least one material segment includes a text segment and / or an audio segment of video script generated for the corresponding video segment.

[0127] The apparatus 600 also includes a target video generation module 640, configured to generate a target video based on at least one video segment and at least one source segment.

[0128] In some embodiments, the apparatus 600 further includes a video segment generation module configured to receive selection of one or more first contents included in one or more media contents; compare a first total duration of one or more first contents with a preset target duration; in response to the first total duration not exceeding the preset target duration, select one or more second contents from one or more media contents using a machine learning model, wherein the one or more second contents are different from one or more first contents; and generate at least one video segment based on one or more first contents and one or more second contents.

[0129] In some embodiments, the device 600 is further configured to determine a second total duration based on the difference between a preset target duration and a first total duration; and to select one or more second contents from one or more media contents with a duration not exceeding the second total duration using a machine learning model.

[0130] In some embodiments, the apparatus 600 further includes a target video segment determination module, configured to, in response to a content selection instruction for target media content, present at least a portion of the target media content and a selection range flag, the selection range including video segments in the target media content that are selected by default; determine an adjusted selection range based on a received adjustment instruction for the selection range; and determine video segments in the target media content that fall within the adjusted selection range as target video segments.

[0131] In some embodiments, the configuration information indicates at least one of the following: the theme of the target video, or the style of the target video.

[0132] In some embodiments, the configuration information acquisition module 620 is further configured to present prompt information about the video topic; and to determine the topic of the target video as at least part of the configuration information based on the received user input.

[0133] In some embodiments, the configuration information acquisition module 620 is further configured to present style information about at least one video style; receive a selection of a video style among at least one video style; and determine the selected video style as the style of the target video as at least part of the configuration information.

[0134] In some embodiments, the target video includes target music, and the apparatus 600 further includes a target music determination module configured to determine the target music based on the style of the target video indicated by configuration information.

[0135] In some embodiments, the device 600 is further configured to determine music that matches the preset video style as target music in response to configuration information indicating that the style of the target video is a preset video style; or to determine recommended music as target music in response to configuration information indicating that the style of the target video is a default video style.

[0136] In some embodiments, the audio segment includes audio of a text segment read aloud in a targeted tone.

[0137] In some embodiments, the apparatus 600 further includes a target timbre determination module, configured to determine a timbre that matches the preset video style as the target timbre in response to a configuration information indicating that the style of the target video is a preset video style; or to determine a timbre that matches the theme of the target video as the target timbre based on the theme of the target video indicated by the configuration information in response to a configuration information indicating that the style of the target video is a default video style.

[0138] In some embodiments, the apparatus 600 further includes a replacement module configured to present corresponding summary information of a plurality of segments in a target video, each of the plurality of segments including a video segment in at least one video segment and a material segment corresponding to the video segment; receive a replacement instruction for a target segment among the plurality of segments and a selection of second media content, the duration of the second media content being greater than or equal to the duration of the target segment; and replace the target segment with the second media content.

[0139] In some embodiments, the apparatus 600 further includes a text update module configured to present a text segment included in at least one material segment; and to receive an edit of a selected text unit based on a selection of text units in the presented text segment, so as to update the selected text unit.

[0140] In some embodiments, the target video includes multiple scenes, each scene corresponding to multiple sub-segments included in at least one footage segment, and the duration of each scene is greater than or equal to the duration of the corresponding sub-segment.

[0141] In some embodiments, the apparatus 600 further includes a regeneration module configured to receive a regeneration instruction for a target video; and in response to the regeneration instruction, generate another target video based on modifications associated with the target video.

[0142] In some embodiments, the apparatus 600 is further configured to receive at least one of a replacement instruction, a deletion instruction, or an addition instruction for one or more media contents to determine additional one or more media contents; and to generate another target video based on the additional one or more media contents and configuration information.

[0143] In some embodiments, the apparatus 600 is further configured to receive modifications to configuration information of a target video; and to generate another target video based on the modified configuration information and one or more media contents.

[0144] Figure 7 A block diagram is shown illustrating an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to achieve Figure 1 Electronic devices 110.

[0145] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.

[0146] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.

[0147] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0148] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0149] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0150] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0151] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0152] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0153] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0155] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for video generation, comprising: receiving a video generation indication for one or more media contents; obtaining configuration information for describing video generation requirements; generating, in response to the video generation indication, at least one video segment and at least one material segment respectively corresponding to the at least one video segment based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment included in a target media content in the one or more media contents, and a material segment in the at least one material segment comprises a text segment and / or an audio segment of a video script generated for a corresponding video segment; and generating a target video based on the at least one video segment and the at least one material segment.

2. The method of claim 1, wherein the at least one video segment is generated by: receiving a selection of one or more first contents included in the one or more media contents; comparing a first total duration of the one or more first contents with a preset target duration; in response to the first total duration not exceeding the preset target duration, selecting one or more second contents from the one or more media contents using a machine learning model, the one or more second contents being different from the one or more first contents; and generating the at least one video segment based on the one or more first contents and the one or more second contents.

3. The method of claim 2, wherein selecting one or more second contents from the one or more media contents using a machine learning model comprises: determining a second total duration based on a difference between the preset target duration and the first total duration; and selecting the one or more second contents from the one or more media contents using the machine learning model, the one or more second contents having a duration not exceeding the second total duration.

4. The method of claim 1, wherein the target video segment is determined by: in response to a content selection indication for the target media content, presenting a marker of at least a portion of the target media content and a selection range, the selection range including a video segment in the target media content that is selected by default; determining an adjusted selection range based on a received adjustment indication for the selection range; and determining a video segment in the target media content falling within the adjusted selection range as the target video segment.

5. The method of claim 1, wherein the configuration information indicates at least one of: a theme of the target video, or a style of the target video.

6. The method of claim 1, wherein obtaining configuration information for describing video generation requirements comprises: presenting prompt information about a video theme; and determining, based on a received user input, a theme of the target video as at least a portion of the configuration information.

7. The method of claim 1, wherein obtaining configuration information for describing video generation requirements comprises: presenting style information about at least one video style; and ​ ​ ​ ​ receiving a selection of a video style among the at least one video style; and determining the selected video style as a style of the target video as at least a part of the configuration information.

8. The method of claim 1, wherein the target video comprises target music, and the method further comprises: determining the target music based on the style of the target video indicated by the configuration information.

9. The method of claim 8, wherein determining the target music comprises: in response to the configuration information indicating that the style of the target video is a preset video style, determining music matching the preset video style as the target music; or in response to the configuration information indicating that the style of the target video is a default video style, determining recommended music as the target music.

10. The method of claim 1, wherein the audio clip comprises audio of reading the text clip in a target voice tone.

11. The method of claim 10, further comprising determining the target voice tone by: in response to the configuration information indicating that the style of the target video is a preset video style, determining a voice tone matching the preset video style as the target voice tone; or in response to the configuration information indicating that the style of the target video is a default video style, determining a voice tone matching a theme of the target video as the target voice tone based on the theme of the target video indicated by the configuration information.

12. The method of claim 1, further comprising: presenting respective summary information of a plurality of clips in the target video, each clip in the plurality of clips comprising a video clip among the at least one video clip and a material clip corresponding to the video clip; receiving a replacement indication for a target clip among the plurality of clips and a selection of second media content, a duration of the second media content being greater than or equal to a duration of the target clip; and replacing the target clip with the second media content.

13. The method of claim 1, further comprising: presenting a text clip comprising the at least one material clip; and receiving an edit of a selected text unit in the presented text clip based on a selection of the selected text unit to update the selected text unit.

14. The method of claim 1, wherein the target video comprises a plurality of scenes, the plurality of scenes respectively corresponding to a plurality of sub-material clips comprising the at least one material clip, and a duration of each scene in the plurality of scenes being greater than or equal to a duration of a corresponding sub-material clip.

15. The method of claim 1, further comprising: receiving a regeneration indication for the target video; and in response to the regeneration indication, generating another target video based on modifications related to the target video.

16. The method of claim 15, wherein generating another target video comprises: receiving at least one of a replacement indication, a deletion indication, or an addition indication for the one or more media content to determine another one or more media content; and ​ ​ ​ ​ generating the another target video based on the one or more media contents and the configuration information. 17.The method of claim 15, wherein generating the another target video comprises: receiving a modification of the configuration information of the target video; and generating the another target video based on the modified configuration information and the one or more media contents. 18.A video generation apparatus, comprising: an indication receiving module configured to receive a video generation indication for one or more media contents; a configuration information obtaining module configured to obtain configuration information describing video generation requirements; a segment generating module configured to, in response to the video generation indication, generate at least one video segment and at least one material segment corresponding to the at least one video segment respectively based on the one or more media contents and the configuration information, wherein a target video segment in the at least one video segment is a video segment included in a target media content in the one or more media contents, and a material segment in the at least one material segment comprises a text segment and / or an audio segment of a video script generated for a corresponding video segment; and a target video generating module configured to generate a target video based on the at least one video segment and the at least one material segment. 19.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 17. 20.A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 17. ​