Method and device for content generation, equipment and storage medium

By generating target text and visually linking text and video clips in the video editing interface, the problem of low video generation efficiency is solved, resulting in a more efficient video editing experience.

CN121865029APending Publication Date: 2026-04-14BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, video clips and text are not tightly bound during video generation, resulting in low generation efficiency. Users need to manually add materials, which increases development costs.

Method used

By generating target text based on feature information, a video editing interface is provided, including a material clip editing area, an editing function area, and an effect preview area. Users can visually associate and present text clips and video clips, and partially overlap them on the timeline to generate the target video.

Benefits of technology

It provides a segmented video generation mode, which improves the efficiency of content generation, allows users to adjust video segments more accurately, and simplifies the video editing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865029A_ABST
    Figure CN121865029A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content generation method and device, equipment and a storage medium. The method comprises the steps of generating a target document for a to-be-generated video based on feature information for describing video content; based on the target copywriting, a video editing interface is presented, and the video editing interface at least comprises a material fragment editing area, an editing function area and an effect preview area; in response to a video clip generation instruction triggered in the editing function area, presenting a plurality of target copywriting clips and corresponding identification images of a plurality of video clips corresponding to the plurality of target copywriting clips in a material clip editing area, wherein the target copywriting segment in the plurality of target copywriting segments and the identification image of the corresponding video segment are visually presented in an associated manner; and generating a target video at least based on the plurality of video clips, in the target video, at least partially overlapping the target copywriting clip in the plurality of target copywriting clips with the corresponding video clip on the timeline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, devices and computer-readable storage media for content generation. Background Technology

[0002] With the rapid development of computer technology, short and long video content has become increasingly abundant, and people's aesthetic standards for videos are rising. Users are also increasingly inclined to edit video content to control the overall pacing. For example, by adding materials, text, and audio, they can make videos more diverse and engaging. Therefore, there is a growing demand for more efficient ways to generate content (e.g., videos) to provide users with a superior experience. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for content generation is provided. The method includes: generating target text for a video to be generated based on feature information describing video content, the target text including multiple target text fragments, each of the multiple target text fragments including a text segment; presenting a video editing interface based on the target text, the video editing interface including at least a material fragment editing area, an editing function area, and an effect preview area, the material fragment editing area for editing text fragments and corresponding video fragments, the editing function area for triggering the generation of one or more types of material, and the effect preview area for presenting an image of the edited effect; responding to a video fragment generation instruction triggered in the editing function area, presenting multiple target text fragments and corresponding identifier images of multiple video fragments corresponding to the multiple target text fragments in the material fragment editing area, wherein the thumbnails of the target text fragments and their corresponding video fragments are presented visually correlated; and generating a target video based at least on the multiple video fragments, wherein in the target video, the target text fragments and their corresponding video fragments at least partially overlap on a timeline.

[0004] In a second aspect of this disclosure, an apparatus for content generation is provided. The apparatus includes: a text generation module configured to generate target text for a video to be generated based on feature information describing video content, the target text including multiple target text fragments, each of the multiple target text fragments including a text segment; and a video editing interface presentation module configured to present a video editing interface based on the target text, the video editing interface including at least a material fragment editing area, an editing function area, and an effect preview area, the material fragment editing area for editing text fragments and corresponding video fragments, the editing function area for triggering the generation of one or more types of material, and the effect preview area for presenting the edited effect. The image; text clip and video clip presentation module is configured to, in response to a video clip generation instruction triggered in the editing function area, present multiple target text clips and corresponding identifier images of multiple video clips corresponding to the multiple target text clips in the material clip editing area, wherein the thumbnails of the target text clips and their corresponding video clips are presented in a visually related manner; and a video generation module is configured to generate a target video based on at least multiple video clips, wherein in the target video, the target text clips and their corresponding video clips at least partially overlap on the timeline.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0010] Figures 2A to 2LA schematic diagram of an example interface for generating target text according to some embodiments of the present disclosure is shown;

[0011] Figures 3A to 3C A schematic diagram of an example interface for determining at least a portion of feature information according to some embodiments of the present disclosure is shown;

[0012] Figure 4 A flowchart illustrating an example process for generating video clips according to some embodiments of this disclosure is shown;

[0013] Figures 5A to 5M A schematic diagram of an example interface for generating video clips according to some embodiments of the present disclosure is shown;

[0014] Figures 6A to 6J A schematic diagram of an example interface for generating video clips according to other embodiments of the present disclosure is shown;

[0015] Figure 7 A schematic diagram of an example architecture for content generation according to some embodiments of the present disclosure is shown;

[0016] Figure 8 A flowchart illustrating the process of generating content according to some embodiments of this disclosure is shown;

[0017] Figure 9 A block diagram of an apparatus for content generation according to some embodiments of the present disclosure is shown; and

[0018] Figure 10 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0026] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0027] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0028] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.

[0029] As used herein, a “unit,” “operational unit,” or “subunit” can consist of any suitable machine learning model or network. As used herein, a set of elements or similar expressions can include one or more such elements. For example, “a set of convolutional units” can include one or more convolutional units.

[0030] As briefly mentioned earlier, users can enrich and diversify their videos by adding materials, text, and audio. Typically, videos can be generated based on text (e.g., a piece of text). For example, in a scenario where one piece of text corresponds to one subtitle, refreshing the text requires refreshing the audio. In a scenario where one piece of text corresponds to n subtitles, each word needs a timestamp, increasing development costs. Consequently, the visuals and text in the generated video are unlinked, and users need to manually add materials to generate the video. Therefore, the efficiency of video generation is relatively low.

[0031] Embodiments of this disclosure propose an improved scheme for content generation. According to various embodiments of this disclosure, target text for a video to be generated is generated based on feature information describing video content. The target text includes multiple target text fragments, each of which includes a piece of text. Based on the target text, a video editing interface is presented. The video editing interface includes at least a material fragment editing area, an editing function area, and an effect preview area. The material fragment editing area is used to edit the text fragments and corresponding video fragments. The editing function area is used to trigger the generation of one or more types of material. The effect preview area is used to present the edited effect image. In response to a video fragment generation instruction triggered in the editing function area, multiple target text fragments and corresponding identifier images of multiple video fragments corresponding to the multiple target text fragments are presented in the material fragment editing area, wherein the target text fragments and the identifier images of the corresponding video fragments are presented visually correlated. Finally, a target video is generated based on at least the multiple video fragments, wherein in the target video, the target text fragments and the corresponding video fragments at least partially overlap on the timeline.

[0032] This provides users with a segmented video generation mode, making it easier for them to adjust different video segments more accurately, thereby improving the efficiency of content generation.

[0033] Example Environment

[0034] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.

[0035] In some embodiments, application 120 may be a content sharing application, a content editing application, a content creation application, etc. Application 120 can provide user 140 with various services related to media content (also referred to as media content items, content items, media items, etc.), including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), publishing, etc. of media content.

[0036] exist Figure 1 In environment 100, if application 120 is active, terminal device 110 can display the interface 150 of application 120. Interface 150 may include various interfaces provided by application 120, such as media content presentation interface, media content creation interface, media content publishing interface, etc. Application 120 can provide media content editing functions (for example, application 120 may be a video editing application) to support editing (e.g., clipping) media content within application 120.

[0037] In some embodiments, terminal device 110 communicates with server 130 to provide services to application 120. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0038] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0039] The following description continues with reference to the accompanying drawings, which will now include some exemplary embodiments of this disclosure. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The various graphic elements on the page may have different arrangements and different visual representations, one or more elements may be omitted or replaced, and one or more other elements may also be present. The embodiments of this disclosure are not limited in this respect. Furthermore, in the following description, exemplary embodiments will be primarily described with respect to terminal device 110. It should be understood that the actions described with respect to terminal device 110 may be performed by application 120 on terminal device 110, or may be performed by application 120 in conjunction with its server (e.g., server 130).

[0040] In some embodiments, the terminal device 110 generates target text for the video to be generated based on feature information used to describe video content. In some embodiments, the target text includes multiple target text fragments, each of which includes a piece of text. In some examples, the terminal device 110 generates target text including multiple target text fragments based on feature information used to generate the video content (e.g., key content).

[0041] In some embodiments, the feature information includes at least one content point indicating video content. The feature information also includes at least one access link corresponding to each of the at least one content point, the access link being used to obtain detailed information about the corresponding content point. In some examples, the feature information may include content point information and links used to generate content (e.g., video).

[0042] Figures 2A to 2L Schematic diagrams of example interfaces 200A to 200L for generating target text according to some embodiments of the present disclosure are shown. Figures 2A to 2B In the example interfaces 200A and 200B shown, user 140 can click the "Start Creating" control 211 to create a target video. In some examples, if user 140 is a new user, terminal device 110 will display a tutorial video. Figure 2A and 2EIn the example interfaces 200A and 200E shown, if terminal device 110 detects that a new user 140 clicks the "Start Creating" control 211, it automatically plays a tutorial video 251, and the "Start Creating" control 211 appears below the tutorial video 250. Subsequently, if terminal device 110 detects that the user clicks the "Start Creating" control 211, it presents a video editing interface 200B for generating the target video, where user 140 can input key points and links. Then, if terminal device 110 receives a click from user 140 on the "Generate Text" control 221, it generates the target text for the video to be generated based on the key points and links.

[0043] like Figure 2A and 2E In example interfaces 200A and 200E to 200F shown on 2F, if terminal device 110 determines that user 140 has created some videos, it can display recent records 262 in a blank editing interface 260 when it detects that user 140 clicks the "Start Creating" control 211. If terminal device 110 detects that the cursor is positioned in the blank editing interface 260, it cancels the display of recent records 262. In some examples, user 140 can click control 261 to view recent records 262. In some examples, user 140 can perform further editing operations on the videos in recent records 262, such as renaming, deleting, copying, etc., of video C.

[0044] In some embodiments, if the terminal device 110 receives a search instruction for a topic related to the video to be generated, it presents search results for that topic, including one or more content summaries for that topic. If the terminal device 110 detects a selection of a target content summary from the one or more content summaries, it adds the target content summary and the access link corresponding to the target content summary as at least part of the feature information. Figures 3A to 3C Schematic diagrams of example interfaces 300A to 300C for identifying at least a portion of feature information according to some embodiments of the present disclosure are shown.

[0045] like Figures 3A to 3CIn the example interfaces 300A to 300C shown, if terminal device 110 detects that user 140 clicks the "Search" control 311, it presents search results for topic A in interface 300C. The search results for topic A include citation sources 331 (e.g., article A, article B, etc.), an overview 332, and a summary of key points 333, etc. In some examples, if terminal device 110 detects that user 140 clicks the text corresponding to a key point, it presents detailed information corresponding to that key point in the interface. In some embodiments, if terminal device 110 detects that user 140 clicks the access link 334 corresponding to the "Overview" key point, it can add the "Overview" key point and the access link 334 to part of the feature information.

[0046] In some examples, the theme for the video to be generated can be freely entered by user 140. If user 140 does not enter a theme, and terminal device 110 detects that user 140 clicks the "Search" control 311, user 140 can enter the theme for generating the target video in input box 313. Figure 2B In the example interface 200B shown, if user 140 has entered topic A, and terminal device 110 detects that user 140 clicks the "Search" control 311, it can directly display the search results 321 for topic A in interface 300B. In some examples, if user 140 has entered topic A, and terminal device 110 detects that user 140 clicks the "Search" control 311 again, it can directly display the historical search records for topic A in interface 300B. In some examples, terminal device 110 can display the search results of multiple searches in a top-down order.

[0047] like Figure 2E and Figure 2B In the example interfaces 200E and 200B shown, if terminal device 110 detects that user 140 clicks the "Assistant Write Text" control 252, it can present the editing function area 220. In some embodiments, the video editing interface may include the editing function area 220. In some embodiments, the editing function area is used to trigger the generation of one or more types of materials. In some examples, multiple types of materials include text, video clips (also known as storyboards), voice-over, subtitles, music, etc. (tab keys). Figure 2B In the example interface 200B shown, if the terminal device 110 detects that the user 140 clicks the text tab key in the editing function area 220, it presents the editing function area 220-1 for editing text. In some examples, the user 140 can input feature information in the editing function area 220-1. Accordingly, such as Figure 5AIn the example interface 500A shown, if the terminal device 110 detects that the user 140 clicks the storyboard tab in the editing function area 220, then the editing function area 220-2 for selecting video segment configurations will be displayed. Figure 6A In the example interface 600A shown, if the terminal device 110 detects that the user 140 clicks the dubbing tab key in the editing function area 220, then the editing function area 220-3 for selecting the timbre will be displayed. For ease of discussion, the editing function areas 220-1, 220-2, 220-3, ... will be collectively referred to as the editing function area 220 below.

[0048] In some embodiments, the terminal device 110 may also obtain additional information about the target video. Accordingly, the terminal device 110 further generates target text based on the additional information. In some embodiments, the additional information indicates at least one of the following: the duration of the target video, the style of the target video, and user customization requirements for the target video. Figure 2C In the example interface 200C shown, user 140 can enter additional information about the target video in the editing function area 220.

[0049] In some examples, user 140 can select the duration of the target video, such as X minutes, provided by terminal device 110 in the editing function area 220. In other examples, user 140 can also enter the style of the target video in the input box 232 for entering the style of the target video included in the editing function area 220. For example, the maximum number of characters that user 140 can enter is XXX, and user 140 can also clear the entered characters with one click. In some examples, terminal device 110 may also support user 140 in selecting to extract video text to determine the style of the target video. Figure 2D In the example interface 200D shown, user 140 can select a local file to choose a video for parsing to extract its text, thereby determining the style of the target video. If terminal device 110 detects that user 140 has selected a video, it displays a progress message 241 indicating the video parsing process. Figure 2C As shown, user 140 can also enter more requirements in input box 233 to generate the target video.

[0050] The following will continue to refer to Figures 2A to 2L The description terminal device 110 generates target text for the video to be generated.

[0051] In some embodiments, the terminal device 110 presents a video editing interface based on the target text. This video editing interface includes at least a clip editing area and an effect preview area. The clip editing area is used to edit text clips and corresponding video clips, and the effect preview area is used to display the edited effect image. Figure 2GThe example interface shown is 200G. After user 140 fills in the feature information, terminal device 110 can display a clip editing area 270. User 140 can edit text clips and corresponding video clips in the clip editing area 270. For example... Figure 5F The example interface 500F shown allows terminal device 110 to display an effect preview area 560. User 140 can browse the video clip corresponding to the text fragment in the effect preview area 560, and / or browse the target video.

[0052] In some embodiments, the terminal device 110 presents multiple initial text fragments generated based on feature information in the material fragment editing area. For example... Figures 2G to 2H In the example interfaces 200G to 200H shown, after user 140 fills in the feature information, if terminal device 110 detects that user 140 clicks the "Generate Text" control 221, then the initial text fragments 271, 272, 282, etc., are streamed in the material fragment editing area (also known as the "Story Panel") 270. In some examples, during the generation of multiple initial text fragments, user 140 can click the "Stop Generation" control 273 to terminate the output of initial text fragments.

[0053] In some embodiments, the text fragments generated by the terminal device 110 can be presented in a predetermined style, such as underlining the generated text fragments. Figure 2G and Figure 2J In the example interfaces 200G and 200J shown, if terminal device 110 detects that the user clicks the "Stop Generating" control 273, terminal device 110 will display the currently generated text fragment. Terminal device 110 can also display an "Assistant Rewrite" control 281 for user 140 to select and rewrite the text fragment.

[0054] In some examples, user 140 can adjust the order of multiple text snippets presented by terminal device 110, copy, delete text snippets, etc. For example... Figure 2H In the example interface 200H shown, if the terminal device 110 detects a user's trigger on the text fragment 271, it can present an editing panel 280 for editing the text fragment 271, allowing the user 140 to adjust the order in which the text fragments 271 are presented, as well as to copy and delete the text fragments 271. In some embodiments, if a text fragment is subjected to certain operations (e.g., copying, deleting, etc.), the digital objects, audio, visuals, and subtitles corresponding to that text fragment are updated accordingly.

[0055] Accordingly, the terminal device 110 can determine multiple target text fragments based on interactive operations received in the material fragment editing area for at least one initial text fragment among multiple initial text fragments. In some embodiments, the terminal device 110 receives modification information for a first initial text fragment among multiple initial text fragments in the material fragment editing area. Subsequently, the terminal device 110 presents an updated text fragment for the first initial text fragment, generated based on modification request information, in the material fragment editing area. The terminal device 110 determines a second target text fragment as one of the multiple target text fragments based on processing instructions for the updated text fragment.

[0056] like Figures 2H to 2K In the example interfaces 200H to 200K shown, if terminal device 110 detects that user 140 clicks the "Assistant Rewrite" control 281 corresponding to the initial text fragment 282 in the material fragment editing area 270, it can present an input box 291 in the material fragment editing area 270 for inputting rewrite prompts. Terminal device 110 can receive the rewrite request information input by user 140 based on input box 291. In some examples, if terminal device 110 detects that user 140 clicks the control corresponding to polishing / expansion / abbreviation, it polishes / expans / abbreviates the rewrite request information input by user 140 and presents it in input box 291. Based on the rewrite request information, terminal device 110 presents the updated text fragment 2112 for the initial text 282 in interface 200K.

[0057] Then, if terminal device 110 detects that user 140 clicks on the regenerate control 2116, it regenerates the second result. If terminal device 110 detects that user 140 clicks on the control 2115 indicating return to the previous section, it presents the previous result. If terminal device 110 detects that user 140 clicks on the cancel control 2117, it exits the rewriting process. Understandably, at this point, the initial text fragment is the target text fragment. If terminal device 110 detects that user 140 clicks on the insert below control 2113, it does not replace the initial text fragment, but inserts the updated text fragment 2112 below the initial text 282 with a line break. If terminal device 110 detects that user 140 clicks on the replace control 2114, it replaces the initial text 282 with the updated text fragment 2112.

[0058] like Figure 2L In the example interface 200L shown, if the terminal device 110 detects that the user clicks the "Generate Text" control 221 again when there is already generated text, it will start generating the second set of multiple target segment texts separated by identifier 2121 and line breaks.

[0059] In some embodiments, the terminal device 110 may also receive a segment processing instruction for at least one of a plurality of initial text segments in the material segment editing area. Then, the terminal device 110 determines at least a portion of the plurality of target text segments according to the segment processing instruction. Figure 7 A schematic diagram of an example architecture 700 for content generation according to some embodiments of the present disclosure is shown. Figure 7 The example architecture 700 shown can support users simultaneously modifying the voice-over / digital objects / subtitles. For example, after user 140 simultaneously modifies the voice-over / digital objects / subtitles, if terminal device 110 detects that the user clicks the refresh control, it will re-combine subtitles A1, A2, B1, C1, C2, D1, D2 into subtitles a1, a2, b1, c1, c2, d1, d2, and re-combine voice-over A, B, C, D into voice-over a, b, c, d. For segments containing digital objects, the digital objects will be refreshed. In some examples, if terminal device 110 does not detect that the user clicks the refresh control, it will still display the original voice-over, digital objects (voice-over + image), and subtitles.

[0060] For example, User 140 can add a line break in the middle of the previous text to split it into two parts, thus generating two video clips. In this case, the visuals / voice-over / digital objects / subtitles, etc., corresponding to the split text remain unchanged from the previous text. Correspondingly, the next shot's visuals / voice-over / digital objects / subtitles will be empty. User 140 can also add a line break at the end of the previous text to create a blank paragraph, thus generating two video clips. In this case, the visuals / voice-over / digital objects / subtitles of the newly added video clip will be empty. As another example, User 140 can click delete at the beginning of the next text to merge paragraphs, thus merging them into one video clip. In this case, the visuals / voice-over / digital objects / subtitles of the next text are deleted, while the visuals / voice-over / digital objects / subtitles of the previous text remain unchanged. User 140 can also delete an entire text paragraph and then click delete again to delete a video clip. In this case, the text / visuals / voice-over / digital objects / subtitles of the deleted video clip are deleted. User 140 can also edit a piece of text to add, delete, or modify content. In this case, the visuals, voice-over, digital human, and subtitles remain unchanged.

[0061] In some embodiments, if the terminal device 110 receives a triggered video clip generation instruction in the editing function area, it presents multiple target text clips and corresponding identifier images of multiple video clips in the material clip editing area. In some embodiments, each target text clip is presented in visual association with the identifier image of its corresponding video clip. In some embodiments, each target text clip is used to describe the corresponding video clip. The following primarily uses a target text clip (e.g., text clip 520) as an example, and refers to... Figure 4 and Figures 5A to 5M This describes the video segment corresponding to the first target text fragment generated by the terminal device. However, this is merely an example and is not intended to limit the scope of the disclosure. For example, user 140 can select from digital objects, matching footage, and a mixture of digital objects and footage for the entire video. Figure 4 A flowchart of an example process 400 for generating video clips according to some embodiments of the present disclosure is shown. Figures 5A to 5M Schematic diagrams of example interfaces 500A to 500M for generating video clips according to some embodiments of the present disclosure are shown.

[0062] like Figure 5B and 5F In the example interfaces 500B and 500F shown, if terminal device 110 receives a click instruction from user 140 in the editing function area 220 to generate a "Generate Storyboard" control 512 for generating video clips, then in the material clip editing area 270, it displays the identifier images (i.e., thumbnails) of text clip 522 and its corresponding video clip 565, the identifier images (i.e., thumbnails) of text clip 521 and its corresponding video clip 563, and so on, for each text clip and the corresponding video clip. In some examples, terminal device 110 displays the identifier images (i.e., thumbnails) of text clips and their corresponding video clips in association.

[0063] In some embodiments, the terminal device 110 presents multiple target text snippets and corresponding identifier images of multiple video snippets in the material snippet editing area, including: the terminal device 110 determining at least one of the following based on the video snippet configuration received in the editing function area: candidate material matching the first target text snippet in the candidate material set, or media content of the first target text snippet read aloud by a target digital object. The terminal device 110 generates a video snippet corresponding to the first target text snippet as one of the multiple video snippets based on at least one of the candidate material or media content. Then, the terminal device 110 presents the identifier images of the first target text snippet and the generated corresponding video snippet in the material snippet editing area. In some examples, the video snippet configuration received by the terminal device 110 may indicate multiple video generation modes selected by the user 140 in the editing function area.

[0064] In some embodiments, the terminal device 110 presents information about multiple video generation modes in an editing function area. These multiple video generation modes include at least: a first mode utilizing digital objects, a second mode using footage matched to text snippets, or a third mode combining digital objects and footage. The terminal device 110 can receive user selections for the multiple video generation modes in the editing function area. Then, based on the user selections, the terminal device 110 determines at least one of the candidate footage or media content.

[0065] In some embodiments, terminal device 110 may receive a user selection of a target digital object from a plurality of digital objects in an editing function area. Then, terminal device 110 generates media content based on the image of the target digital object and the corresponding voice. In some examples, when user 140 selects a digital person in editing function area 270, terminal device 110 first segments the text fragment into sentences. Based on the voice of the digital object selected by user 140, terminal device 110 generates audio of the digital object reading the text fragment. Then, based on the image and audio of the digital object, terminal device 110 generates visuals of the digital object reading the text fragment.

[0066] Understandably, terminal device 110 can determine the scene of a digital object (also known as a digital human) reading text segment 520. Subsequently, based on the scene of the digital object reading text segment 520, terminal device 110 generates a video segment corresponding to text segment 520, which can then be used as one of multiple video segments. For example... Figure 5K and 5LIn the example interfaces 500K and 500L, when user 140 selects the digital human mode, terminal device 110 can generate a video clip 5121 corresponding to the text clip 520 based on the digital human A5110 selected by user 140.

[0067] In some examples, during the process of generating a screen showing the digital object reading a passage (i.e., rendering the digital object), terminal device 110 can display a prompt message indicating that the digital object is being rendered. If rendering is successful, terminal device 110 can display a prompt message indicating successful rendering. Accordingly, user 140 can preview the screen showing the digital object reading the passage. If rendering fails, terminal device 110 can display a prompt message indicating rendering failure and a prompt message indicating a retry. Figure 5C In the example interface 500C shown, user 140 can also click "Create Digital Object" to create their own digital object, such as digital human C531.

[0068] In some embodiments, the terminal device 110 can determine candidate materials that match the text fragment 520 from a candidate material set. It is understood that the terminal device 110 can generate a video fragment corresponding to the text fragment 520 based on the candidate materials that match it, so as to serve as one of multiple video fragments. For example... Figure 5M In the example interface 500M, when user 140 selects matching material, terminal device 110 can generate a description of the corresponding video segment based on text fragment 520, and then select candidate material B matching text fragment 520 from the candidate material library to generate video segment 5131 corresponding to text fragment 520. In some examples, if there is a scene of a digital object reading text fragment 520, the candidate material B matching text fragment 520 can be determined based on the duration of the digital object reading text fragment 520. Accordingly, after terminal device 110 generates the video segment corresponding to text fragment 520, user 140 can preview the video segment corresponding to text fragment 520.

[0069] The following will continue to refer to Figures 5A to 5J and Figure 4 The terminal device 110 generates a video segment corresponding to the text segment 520 based on the image of the digital object reading the text segment 520 and candidate materials, so as to serve as a segment among multiple video segments.

[0070] like Figures 5A to 5FIn the example interfaces 500A to 500F shown, in the absence of text, if the terminal device 110 detects that the user 140 clicks the generation storyboard control 512, which is used to indicate the generation of a video clip corresponding to the text fragment, then it presents a prompt message 511 indicating that the video clip (also known as a "storyboard") corresponding to the text fragment can only be generated after inputting text.

[0071] In some embodiments, when user 140 selects the digital object and material mixing mode 521, terminal device 110 generates a video clip corresponding to the text segment 520 based on the digital object selected by user 140 (e.g., digital human A) and the material determined to match the text segment 520. Accordingly, terminal device 110 presents prompt information 541 in interface 500D to indicate that it is in the text segmentation process.

[0072] Subsequently, after segmenting the text fragment 520 into sentences, the terminal device 110 generates audio of digital human A reading the text fragment 520 aloud. Correspondingly, during the generation of the video of digital human A reading the text fragment 520, the terminal device 110 displays a prompt message 551 indicating that the video fragment corresponding to the text fragment 520 is in the process of being generated. After the video fragment corresponding to the text fragment 520 is generated, the terminal device 110 displays the video fragment 561 corresponding to the text fragment 520. In some examples, upon successful generation of the video fragment, the terminal device 110 may also display a prompt message 564 "Generation successful" in the area corresponding to the video fragment 561.

[0073] like Figure 4 In the example process 400 shown, in box 410, terminal device 110 segments each of the multiple segmented segments and determines the video generation mode. In box 411, terminal device 110 generates audio corresponding to the text segment based on the timbre of the digital object selected by user 140. Accordingly, in box 412, if terminal device 110 determines that the video generation mode requires additional material, it generates description information for the video segment. In box 413, terminal device 110 determines the material matching the text segment. In box 414, terminal device 110 can obtain the resource description matching the text segment from the server based on the resource identifier corresponding to the material matching the text segment. In box 415, terminal device 110 can download the resource description matching the text segment. In box 416, terminal device 110 presents the screen corresponding to the material matching the text segment. In some embodiments, the screen on the player is rendered in real time, and user 140 can preview it.

[0074] In box 417, after generating audio and determining the video generation mode in parallel, terminal device 110 converts the audio link corresponding to the text fragment into a file resource. Accordingly, in box 418, terminal device 110 saves the file resource to storage space (e.g., cloud storage). In box 419, the browser loads the audio corresponding to the link. In box 420, terminal device 110 plays the audio loaded by the browser. In box 421, terminal device 110 renders an image of the digital object, meaning the player can support users previewing static digital humans in the video fragment. In box 422, terminal device 110 renders the digital human, and the player can preview the dynamic digital human. In some examples, when the image corresponding to the material matching the text fragment fails to be generated successfully, the corresponding video image is displayed as a black screen.

[0075] In some embodiments, if the terminal device 110 receives a content replacement instruction or content addition instruction for a third video segment among multiple video segments in the media clip editing area, it presents at least one media source. For example... Figures 5G to 5H In the example interfaces 500G to 500H shown, if the terminal device 110 detects that the user 140 clicks on the content corresponding to the text fragment 520 (e.g., the image), it will present a material source panel 580 including a "My Materials" list and a "Material Library" list.

[0076] like Figure 5G As shown, if the content corresponding to the text fragment 520 is empty, the terminal device 110 can display a plus sign 574 in the corresponding area of ​​the text fragment 520. If the content corresponding to the text fragment 520 is not empty, the terminal device 110 can display editing controls in the corresponding area of ​​the text fragment 520, such as a cropping control 572, a replacement control 571, and a clear control 572. In some examples, if the text fragment 520 contains descriptive information, the terminal device 110 displays the descriptive information corresponding to the text fragment 520.

[0077] In some embodiments, terminal device 110 updates a third video segment based on the selection of material from at least one source. For example... Figures 5H to 5IIn the example interfaces 500H to 500I and 500J shown in Figure 5J, if terminal device 110 detects that user 140 clicks on "Material A" 581 included in the "My Materials" list, it generates video segment 591 based on "Material A" 581. In some embodiments, after selecting "Material A" 581, user 140 can also trim the duration corresponding to "Material A" 581. That is, the trimming frame duration can default to the recommended duration, left-aligned with the original material 0s; if the video material duration is less than the recommended duration, the trimming frame duration defaults to the full video length. In some examples, the user can also drag the trimming frame to change the selection position, and can also stretch the trimming frame left / right to change the duration. In some examples, the user can also drag to adjust the volume percentage.

[0078] In some embodiments, user 140 can also select materials from the "Material Library" list. That is, if terminal device 110 detects that user 140 clicks on the "Material Library" list in the material source panel 580, it can display a search box 5101. In some embodiments, user 140 can search for materials based on the search box 5101.

[0079] In some embodiments, if the terminal device 110 receives an instruction to add an object to a fourth video clip among multiple video clips in the media clip editing area, it adds a digital object to the fourth video clip to read aloud the media content of the target text clip corresponding to the fourth video clip. For example... Figure 5F As shown, when a text fragment has a corresponding digital object, the terminal device 110 can add a slot 562 corresponding to the digital object in the interface. In some examples, if the terminal device 110 receives an instruction to add an object to the video fragment 563, it adds media content of digital object reading and text fragment 521 to the video fragment 563.

[0080] If terminal device 110 receives an object removal instruction for the fifth video clip among multiple video clips in the media clip editing area, it can remove the media content of the digital object reading the target text clip corresponding to the fifth video clip from the fifth video clip. For example... Figure 5F As shown, if terminal device 110 receives a removal instruction for the object displayed in slot 562 corresponding to video segment 564, it can remove the media content of digital object reading and text segment 522 from video segment 564. Figure 5G As shown, user 140 can also adjust the position of digital objects and materials presented in the video clip. Terminal device 110 can present controls for downloading, uploading, deleting, rotating, and scaling digital objects and materials in an area 575 associated with the video clip, allowing user 140 to download, upload, delete, rotate, and scale digital objects and materials.

[0081] The following is for reference only. Figures 6A to 6J This disclosure describes the process used to generate video clips. Figures 6A to 6J Schematic diagrams of example interfaces 600A to 600J for generating video clips according to other embodiments of the present disclosure are shown.

[0082] In some embodiments, terminal device 110 generates an audio segment that reads a target text segment corresponding to the first video segment, based on audio configuration information received in the editing function area, for at least a first video segment among a plurality of video segments. Then, terminal device 110 adds the audio segment to the first video segment. The following description uses video segment 610 among a plurality of video segments as an example, but this is merely exemplary and not intended to limit the scope of the disclosure. It is understood that the audio configuration information can be global (i.e., for the entire video).

[0083] like Figure 6A In the example interface 600A shown, if the terminal device 110 detects that the user 140 clicks on the application tone control 613, it can generate an audio segment corresponding to the target text segment of the video segment 610 (i.e., dubbing the target text segment) based on the C tone 611 selected by the user 140 and the speech rate set by the user 140, etc. In some examples, if the terminal device 110 detects that the user 140 clicks on the C tone 611, it can play the C tone 611 for the user 140 to listen to. In some examples, the terminal device 110 presents a prompt message 612 indicating that the generation is in progress during the process of generating the audio segment.

[0084] like Figures 6C to 6D In the example interface 600D shown, terminal device 110 can also provide user 140 with a creation entry 631 for creating timbres. User 140 can create a timbre 641 based on creation entry 631. In some examples, if terminal device 110 detects that user 140 has modified the speech rate, it compresses the already generated audio segment. If terminal device 110 detects that user 140 has modified the text, it regenerates the audio segment based on the modified text and replaces the original audio segment. In some examples, digital objects and subtitles are also updated along with the deletion, compression, and regeneration of audio segments.

[0085] In some embodiments, the terminal device 110 generates an audio segment corresponding to the text segment based on the text segment, the user-selected timbre, and the speaking speed. The audio segment is linked in real time to the timestamp, digital object, and the image of the digital object reading the text segment.

[0086] In some embodiments, the terminal device 110 may also provide the user 140 with an entry point for selecting music, so as to add music to a target video or a video clip. For example... Figures 6G to 6H In the example interfaces 600G to 600H shown, if terminal device 110 detects that user 140 clicks the "Add" control 671, it adds "Music A" to the target video or a video segment. Conversely, if terminal device 110 detects that user 140 clicks the "Delete" control 681, it removes "Music A" from the target video or a video segment.

[0087] In some embodiments, terminal device 110 identifies one or more keywords in a target text segment corresponding to at least a second video segment among a plurality of video segments. Accordingly, terminal device 110 determines a corresponding visual effect for the one or more keywords. Then, terminal device 110 renders the one or more keywords in the subtitles of the second video segment according to the corresponding visual effect.

[0088] like Figure 6F In the example interface 600F shown, terminal device 110 identifies keywords 665 in the target text segment corresponding to video clip 610. In some examples, terminal device 110 can use a model to identify keywords 665 in the target text segment corresponding to video clip 610. Subsequently, terminal device 110 renders keywords 665 in the subtitles of video clip 610 according to style 663 selected by user 140, with the effect of style 663. In some examples, terminal device 110 can present controls for downloading, uploading, deleting, rotating, and zooming subtitles in an area 662 associated with the video clip, allowing user 140 to download, upload, delete, rotate, and zoom subtitles.

[0089] In some embodiments, if the terminal device 110 receives a subtitle enhancement instruction for a target video in the editing function area, it identifies one or more keywords. For example, the user 140 can use the subtitle enhancement instruction (e.g., the user clicks the control corresponding to "One-Click Packaging") to highlight subtitles, adjust subtitle sound effects, and modify subtitle animations, etc. Figure 6E In the example interface 600E shown, if the terminal device 110 detects that the user 140 clicks the "one-click wrap" control 651 used to indicate the subtitles of the video clip 610 to beautify them, then the terminal device 110 renders the keywords in the subtitles of the video clip 610 according to the predetermined style.

[0090] In some embodiments, if the terminal device 110 detects that the user 140 clicks the "One-Click Packaging" control 651, it can transmit the subtitles and timestamps corresponding to the text fragment to the algorithm. Further, the terminal device 110 uses a model to identify keywords, packaging points, and effects in the text fragment, and presents the subtitles, keywords, and packaging elements (e.g., sound effects, text animations, stickers, text templates, special effects, etc.) corresponding to the text fragment returned by the model.

[0091] In some examples, if no voice-over is provided for a text segment, and the terminal device 110 detects that the user 140 clicks the "One-Click Packaging" control 651, it can display a prompt message indicating that the beautification indicator is unavailable. For example, "Packaging is unavailable, please add voice-over first." In some embodiments, the terminal device 110 displays the corresponding subtitles for the text segment while generating the voice-over. That is, the terminal device 110 can perform subtitle timing based on the timestamps of the text segment and the audio segment, thereby determining the appearance time of the subtitles.

[0092] In some embodiments, the terminal device 110 displays an additional status indicator for the target video in the media clip editing area. This additional status indicator indicates the type of additional content being added to the target video. For example... Figure 6B In the example interface 600B shown, if user 140 adds a timbre, terminal device 110 can display an additional status identifier 621 in the material clip editing area 270 to indicate that the timbre has been added. Figure 6I In the example interface 600I shown, if user 140 adds timbre, visual effects corresponding to subtitles, and music, terminal device 110 can display additional status identifiers 621, 691, and 692 in the target text presentation area to indicate the added timbre, visual effects corresponding to subtitles, and music.

[0093] In some embodiments, if the terminal device 110 receives a deletion operation for the additional status identifier in the material clip editing area, it removes the additional type of content from the target video. For example... Figure 6B In the example interface 600B shown, if the terminal device 110 detects that the user 140 clicks the delete control 622, the additional status identifier 621 used to indicate that the tone has been added can be deleted.

[0094] In some embodiments, if the terminal device 110 receives a detected selection of a text segment from multiple target text segments in the material segment editing area, it can play a video segment corresponding to the selected text segment in the effect preview area. For example... Figure 5FIn the example interface 500F shown, if the terminal device 110 detects that the user 140 clicks on the video clip corresponding to the text clip 520 in the material clip editing area 270, the video corresponding to the text clip 520 can be played in the effect preview area 560.

[0095] In some embodiments, terminal device 110 generates a target video based on at least a plurality of video clips. In some embodiments, in the target video, target text clips among a plurality of target text clips at least partially overlap with their corresponding video clips on the timeline. In some examples, each text clip among a plurality of target text clips included in the target video overlaps with its corresponding video clip on the timeline. Understandably, users can edit the text clips generated by terminal device 110. Subsequently, terminal device 110 updates the video clip corresponding to the text clip based on the user's editing of the text clip, thereby updating the target video. In some examples, terminal device 110 generates the target video based on at least a plurality of video clips, and the target video can be played in an effect preview area.

[0096] like Figure 6J In the example interface 600J shown, if the terminal device 110 detects that the user 140 clicks the "Start Export" control 6101, it can export the target video based on multiple video clips. The terminal device 110 can export the target video based on the cover, name, resolution, quality, frame rate, and format corresponding to the target video selected by the user 140.

[0097] In summary, based on the feature information used to describe the key content of the video, multiple text clips can be generated in a streaming manner. Furthermore, based on these multiple text clips, multiple video clips associated with each text clip can be generated, allowing users to more accurately adjust different video clips, thereby improving the efficiency of content generation.

[0098] Example process

[0099] Figure 8 A flowchart of a content generation process 800 according to some embodiments of the present disclosure is shown. Process 800 can be implemented at terminal device 110. Reference is made below. Figure 1 Describe the process 800.

[0100] In box 810, terminal device 110 generates target text for the video to be generated based on feature information used to describe video content. The target text includes multiple target text fragments, and each of the multiple target text fragments includes a piece of text.

[0101] In frame 820, terminal device 110 presents a video editing interface based on multiple target texts. The video editing interface includes at least a material clip editing area, an editing function area, and an effect preview area. The material clip editing area is used to edit text clips and corresponding video clips. The editing function area is used to trigger the generation of one or more materials. The effect preview area is used to present the edited effect image.

[0102] In frame 830, in response to a video clip generation instruction triggered in the editing function area, the terminal device 110 presents multiple target text clips and corresponding identification images of multiple video clips in the material clip editing area, wherein the target text clips and the identification images of the corresponding video clips are presented in a visually related manner.

[0103] In frame 840, terminal device 110 generates a target video based on at least a plurality of video segments, wherein in the target video, the target text segments of the plurality of target text segments at least partially overlap with the corresponding video segments on the timeline.

[0104] In some embodiments, presenting the corresponding identifier images of the plurality of target text fragments and the plurality of video fragments corresponding to the plurality of target text fragments in the material fragment editing area includes: for a first target text fragment among the plurality of target text fragments, determining at least one of the following based on the video fragment configuration received in the editing function area: a candidate material in the candidate material set that matches the first target text fragment, or media content in which a target digital object reads the first target text fragment; generating a video fragment corresponding to the first target text fragment based on at least one of the candidate material or media content, as one of the plurality of video fragments; and presenting the identifier images of the first target text fragment and the generated corresponding video fragment in the material fragment editing area.

[0105] In some embodiments, determining media content that reads a first target text segment aloud by a target digital object includes: receiving a user selection of a target digital object among a plurality of digital objects in an editing function area; and generating media content based on the image of the target digital object and the timbre corresponding to the target digital object.

[0106] In some embodiments, determining at least one of the candidate materials or media content includes: presenting information about multiple video generation modes in an editing area, the multiple video generation modes including at least: a first mode utilizing digital objects, a second mode matching materials based on text snippets, or a third mode combining digital objects and materials; receiving user selections for the multiple video generation modes in the editing area; and determining at least one of the candidate materials or media content based on the user selections.

[0107] In some embodiments, the feature information includes: at least one content point indicating video content, and at least one access link corresponding to each of the at least one content point, wherein the access link is used to obtain detailed information of the corresponding content point.

[0108] In some embodiments, process 800 further includes: in response to receiving a retrieval instruction for a topic of a video to be generated, presenting retrieval results for the topic, the retrieval results including one or more content summaries for the topic; and in response to selecting a target content summary from the one or more content summaries, adding the target content summary and the access link corresponding to the target content summary as at least part of the feature information.

[0109] In some embodiments, process 800 further includes: obtaining additional information about the target video, wherein the generation of target copy is further based on the additional information, and the additional information indicates at least one of the following: the duration of the target video, the style of the target video, or user customization requirements for the target video.

[0110] In some embodiments, generating target text for a video to be generated includes: presenting a plurality of initial text fragments generated based on feature information in a material fragment editing area; and determining a plurality of target text fragments based on interactive operations received in the material fragment editing area for at least one initial text fragment among the plurality of initial text fragments.

[0111] In some embodiments, determining a plurality of target copy fragments includes: receiving modification request information for a first initial copy fragment among a plurality of initial copy fragments in a material fragment editing area; presenting an updated copy fragment for the first initial copy fragment generated based on the modification request information in the material fragment editing area; and determining a second target copy fragment as one of the plurality of target copy fragments based on processing instructions for the updated copy fragment.

[0112] In some embodiments, determining a plurality of target text fragments includes: receiving a fragment processing instruction for at least one text fragment among a plurality of initial text fragments in a material fragment editing area; and determining at least a portion of the plurality of target text fragments based on the fragment processing instruction.

[0113] In some embodiments, process 800 further includes: for at least a first video segment among a plurality of video segments, generating an audio segment that reads a target text segment corresponding to the first video segment based on audio configuration information received in the editing function area; and adding the audio segment to the first video segment.

[0114] In some embodiments, process 800 further includes: for at least a second video segment among a plurality of video segments, identifying one or more keywords in a target text segment corresponding to the second video segment; determining a corresponding visual effect for the one or more keywords; and rendering the one or more keywords in the subtitles of the second video segment according to the corresponding visual effect.

[0115] In some embodiments, identifying one or more keywords is in response to receiving a subtitle enhancement instruction for a target video in the editing ribbon.

[0116] In some embodiments, process 800 further includes: in response to receiving a content replacement instruction or content addition instruction for a third video segment among a plurality of video segments in the media segment editing area, presenting at least one media source; and updating the third video segment based on the selection of media in the at least one media source.

[0117] In some embodiments, process 800 further includes at least one of the following: in response to receiving an object addition instruction for a fourth video segment among a plurality of video segments in the media clip editing area, adding media content of a digital object reading of a target text segment corresponding to the fourth video segment to the fourth video segment; or in response to receiving an object removal instruction for a fifth video segment among a plurality of video segments in the media clip editing area, removing media content of a digital object reading of a target text segment corresponding to the fifth video segment from the fifth video segment.

[0118] In some embodiments, process 800 further includes: in the clip editing area, presenting an additional status indicator for the target video, the additional status indicator indicating additional types of content added to the target video.

[0119] In some embodiments, process 800 further includes: removing additional type content from the target video in response to receiving a deletion operation for an additional status identifier in the material clip editing area.

[0120] In some embodiments, process 800 further includes: receiving a selection of a text segment from a plurality of target text segments in a material segment editing area; and, in response to the received selection, playing a video segment corresponding to the selected text segment in an effect preview area.

[0121] Example devices and equipment

[0122] Figure 9 A schematic structural block diagram of an apparatus 900 for content generation according to certain embodiments of the present disclosure is shown. The apparatus 900 may be implemented as or included in a terminal device 110. Various modules / components in the apparatus 900 may be implemented by hardware, software, firmware, or any combination thereof.

[0123] As shown in the figure, the device 900 includes a text generation module 910, configured to generate target text for a video to be generated based on feature information used to describe video content. The target text includes multiple target text fragments, each of which includes a piece of text. The device 900 also includes a video editing interface presentation module 920, configured to present a video editing interface based on the multiple target texts. The video editing interface includes at least a material fragment editing area, an editing function area, and an effect preview area. The material fragment editing area is used to edit text fragments and corresponding video fragments. The editing function area is used to trigger the generation of one or more materials. The effect preview area is used to present the edited effect image. The device 900 also includes a text fragment and video fragment presentation module 930, configured to, in response to a video fragment generation instruction triggered in the editing function area, present multiple target text fragments and corresponding identifier images of multiple video fragments in the material fragment editing area. The target text fragments and their corresponding video fragment identifier images are presented visually correlated. The apparatus 900 also includes a video generation module 940 configured to generate a target video based on at least a plurality of video segments, wherein in the target video, target text segments in the plurality of target text segments at least partially overlap with the corresponding video segments on the timeline.

[0124] In some embodiments, the text snippet and video snippet presentation module 930 is further configured to, for a first target text snippet among a plurality of target text snippets, determine at least one of the following based on the video snippet configuration received in the editing function area: a candidate material in a candidate material set that matches the first target text snippet, or media content in which a target digital object reads the first target text snippet; generate a video snippet corresponding to the first target text snippet as one of a plurality of video snippets based on at least one of the candidate material or media content; and present an identifier image of the first target text snippet and the generated corresponding video snippet in the material snippet editing area.

[0125] In some embodiments, the text and video clip presentation module 930 is further configured to receive a user selection of a target digital object among a plurality of digital objects in an editing function area; and to generate media content based on the image of the target digital object and the timbre corresponding to the target digital object.

[0126] In some embodiments, the text snippet and video snippet presentation module 930 is further configured to present information about multiple video generation modes in the editing function area, the multiple video generation modes including at least: a first mode utilizing digital objects, a second mode matching material according to text snippets, or a third mode combining digital objects and material; and to receive user selections of the multiple video generation modes in the editing function area; and to determine at least one of candidate material or media content based on user selections.

[0127] In some embodiments, the feature information includes: at least one content point indicating video content, and at least one access link corresponding to each of the at least one content point, wherein the access link is used to obtain detailed information of the corresponding content point.

[0128] In some embodiments, the apparatus 900 further includes an adding module configured to, in response to receiving a retrieval instruction for a topic of a video to be generated, present retrieval results for the topic, the retrieval results including one or more content summaries for the topic; and, in response to selecting a target content summary from the one or more content summaries, add the target content summary and the access link corresponding to the target content summary as at least part of the feature information.

[0129] In some embodiments, the apparatus 900 further includes an additional information generation module configured to acquire additional information about the target video, wherein the generation of the target text is further based on the additional information, and the additional information indicates at least one of the following: the duration of the target video, the style of the target video, or user customization requirements for the target video.

[0130] In some embodiments, the copywriting generation module 910 is further configured to present a plurality of initial copywriting fragments generated based on feature information in the material fragment editing area; and to determine a plurality of target copywriting fragments based on interactive operations received in the material fragment editing area for at least one initial text fragment among the plurality of initial copywriting fragments.

[0131] In some embodiments, the copy generation module 910 is further configured to receive modification request information for a first initial copy fragment among a plurality of initial copy fragments in a material fragment editing area; present an updated copy fragment for the first initial copy fragment generated based on the modification request information in the material fragment editing area; and determine a second target copy fragment as one of a plurality of target copy fragments based on processing instructions for the updated copy fragment.

[0132] In some embodiments, the copy generation module 910 is further configured to receive a fragment processing instruction for at least one of a plurality of initial copy fragments in a material fragment editing area; and to determine at least a portion of a plurality of target copy fragments based on the fragment processing instruction.

[0133] In some embodiments, the apparatus 900 further includes a video clip adding module, configured to generate an audio clip that reads a target text clip corresponding to the first video clip, based on audio configuration information received in the editing function area, for at least a first video clip among a plurality of video clips; and to add the audio clip to the first video clip.

[0134] In some embodiments, the apparatus 900 further includes a keyword rendering module configured to, for at least a second video segment among a plurality of video segments, identify one or more keywords in a target text segment corresponding to the second video segment; determine corresponding visual effects for the one or more keywords; and render one or more keywords in the subtitles of the second video segment according to the corresponding visual effects.

[0135] In some embodiments, identifying one or more keywords is in response to receiving a subtitle enhancement instruction for a target video in the editing ribbon.

[0136] In some embodiments, the apparatus 900 further includes a video segment update module configured to, in response to receiving a content replacement instruction or content addition instruction for a third video segment among a plurality of video segments, present at least one source material; and update the third video segment based on the selection of material from the at least one source material.

[0137] In some embodiments, the apparatus 900 further includes a digital human removal module, configured to add media content of a digital object reading a target text segment corresponding to the fourth video segment in response to receiving an object addition instruction for a fourth video segment among a plurality of video segments in the media clip editing area, or to remove media content of a digital object reading a target text segment corresponding to the fifth video segment from the fifth video segment in response to receiving an object removal instruction for a fifth video segment among a plurality of video segments.

[0138] In some embodiments, the device 900 further includes a status indicator presentation module configured to present additional status indicators for the target video in the clip editing area, the additional status indicators indicating additional types of content added to the target video.

[0139] In some embodiments, the apparatus 900 further includes a content removal module configured to remove additional type content from the target video in response to receiving a deletion operation for an additional status identifier in the clip editing area.

[0140] In some embodiments, the device 900 further includes a video clip playback module configured to receive a selection of a text clip from a plurality of target text clips in a material clip editing area; and in response to the received selection, to play a video clip corresponding to the selected text clip in an effect preview area.

[0141] Figure 10 A block diagram is shown illustrating an electronic device 1000 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 10 The electronic device 1000 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 10 The electronic device 1000 shown can be used to achieve Figure 1 Terminal equipment 110.

[0142] like Figure 10 As shown, electronic device 1000 is in the form of a general-purpose electronic device. Components of electronic device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processing unit 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 1000.

[0143] Electronic device 1000 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 1000.

[0144] Electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 10 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1020 may include computer program product 1025 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0145] The communication unit 1040 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 1000 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0146] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1000, or with any device that enables electronic device 1000 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0147] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0148] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0149] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0150] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0152] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A content generation method, comprising: Based on feature information used to describe video content, target text for the video to be generated is generated. The target text includes multiple target text fragments, and each of the multiple target text fragments includes a piece of text. Based on the target text, a video editing interface is presented. The video editing interface includes at least a material clip editing area, an editing function area, and an effect preview area. The material clip editing area is used to edit text clips and corresponding video clips. The editing function area is used to trigger the generation of one or more types of materials. The effect preview area is used to present the edited effect image. In response to a video clip generation instruction triggered in the editing function area, the plurality of target text clips and corresponding identification images of the plurality of video clips corresponding to the plurality of target text clips are presented in the material clip editing area, wherein the target text clips and the identification images of the corresponding video clips are presented in a visually associated manner. as well as A target video is generated based on at least the plurality of video segments, wherein in the target video, the target text segments among the plurality of target text segments at least partially overlap with the corresponding video segments on the timeline.

2. The method according to claim 1, wherein presenting the plurality of target text fragments and corresponding identifier images of the plurality of video fragments respectively corresponding to the plurality of target text fragments in the material fragment editing area includes: For the first target text fragment among the multiple target text fragments, Based on the video clip configuration received in the editing function area, at least one of the following is determined: candidate material in the candidate material set that matches the first target text clip, or media content in which the target digital object reads the first target text clip; Based on at least one of the candidate materials or the media content, a video clip corresponding to the first target text segment is generated as one of the plurality of video clips; and The first target text fragment and the corresponding generated video fragment's identifier image are displayed in the material fragment editing area.

3. The method according to claim 2, wherein determining the media content of reading the first target text fragment aloud to the target digital object includes: The editing function area receives a user selection of the target digital object from among a plurality of digital objects; as well as The media content is generated based on the image of the target digital object and the timbre corresponding to the target digital object.

4. The method of claim 2, wherein determining at least one of the candidate materials or the media content comprises: The editing function area presents information about multiple video generation modes, which include at least: Using the first pattern of digital objects, The second mode of matching materials based on text snippets, or A third mode combining virtual objects and assets; and The editing function area receives user selections for the plurality of video generation modes; and Based on the user's selection, at least one of the candidate materials or the media content is determined.

5. The method according to claim 1, wherein the feature information includes: Indicates at least one key content point of the video content, and At least one access link corresponding to each of the at least one content point, the access link being used to obtain detailed information about the corresponding content point.

6. The method according to claim 1, further comprising: In response to receiving a search instruction for a topic of the video to be generated, search results for the topic are presented, the search results including one or more content summaries for the topic; as well as In response to the selection of a target content summary from the one or more content summaries, the target content summary and the access link corresponding to the target content summary are added as at least a part of the feature information.

7. The method according to claim 1, further comprising: Obtain additional information about the target video, wherein the generation of the target text is further based on the additional information, and the additional information indicates at least one of the following: The duration of the target video, The style of the target video, or User customization requirements for the target video.

8. The method of claim 1, wherein generating the target text for the video to be generated comprises: In the material clip editing area, multiple initial text clips generated based on the feature information are presented; as well as The plurality of target text fragments are determined based on the interactive operations received in the material fragment editing area for at least one initial text fragment among the plurality of initial text fragments.

9. The method of claim 8, wherein determining the plurality of target text fragments comprises: The editing area for the material fragments receives modification request information for the first initial text fragment among the plurality of initial text fragments; In the material clip editing area, an updated copy clip generated based on the modification requirement information for the first initial copy clip is presented; as well as Based on the processing instructions for the updated copy fragment, a second target copy fragment is determined as one of the plurality of target copy fragments.

10. The method of claim 8, wherein determining the plurality of target text fragments comprises: The material segment editing area receives a segment processing instruction for at least one of the plurality of initial text segments; as well as Based on the fragment processing instructions, at least a portion of the plurality of target text fragments is determined.

11. The method according to claim 1, further comprising: For at least the first video segment among the plurality of video segments, an audio segment is generated that reads a target text segment corresponding to the first video segment, based on the audio configuration information received in the editing function area; as well as Add the audio clip to the first video clip.

12. The method according to claim 1, further comprising: For at least the second video segment among the plurality of video segments, identify one or more keywords in the target text segment corresponding to the second video segment; Determine the appropriate visual effects for the one or more keywords; as well as The one or more keywords are rendered in the subtitles of the second video segment according to the corresponding visual effects.

13. The method of claim 12, wherein identifying the one or more keywords is in response to receiving a subtitle enhancement instruction for the target video in the editing function area.

14. The method according to claim 1, further comprising: In response to receiving a content replacement instruction or content addition instruction for a third video segment among the plurality of video segments in the media segment editing area, at least one media source is presented; as well as The third video segment is updated based on the selection of material from the at least one source material.

15. The method of claim 1, further comprising at least one of the following: In response to receiving an instruction to add an object to a fourth video clip among the plurality of video clips in the media clip editing area, a digital object is added to the fourth video clip to read aloud the media content of a target text clip corresponding to the fourth video clip, or In response to receiving an object removal instruction for a fifth video clip among the plurality of video clips in the media clip editing area, the media content of the digital object reading the target text clip corresponding to the fifth video clip is removed from the fifth video clip.

16. The method according to claim 1, further comprising: In the clip editing area, additional status indicators are displayed for the target video, indicating the additional types of content added to the target video.

17. The method of claim 16, further comprising: In response to receiving a deletion operation for the additional status identifier in the media clip editing area, the additional type of content is removed from the target video.

18. The method according to claim 1, further comprising: The selection of text segments from the plurality of target text segments is received in the material segment editing area; as well as In response to the received selection, a video clip corresponding to the selected text clip is played in the effect preview area.

19. An apparatus for content generation, comprising: The text generation module is configured to generate target text for the video to be generated based on feature information used to describe the video content. The target text includes multiple target text fragments, and each of the multiple target text fragments includes a piece of text. The video editing interface presentation module is configured to present a video editing interface based on the target text. The video editing interface includes at least a material clip editing area, an editing function area, and an effect preview area. The material clip editing area is used to edit text clips and corresponding video clips. The editing function area is used to trigger the generation of one or more types of materials. The effect preview area is used to present the edited effect image. The text clip and video clip presentation module is configured to, in response to a video clip generation instruction triggered in the editing function area, present the plurality of target text clips and corresponding identifier images of the plurality of video clips respectively corresponding to the plurality of target text clips in the material clip editing area, wherein the identifier images of the target text clips and the corresponding video clips are presented in a visually associated manner. as well as A video generation module is configured to generate a target video based on at least the plurality of video segments, wherein in the target video, the target text segments among the plurality of target text segments at least partially overlap with the corresponding video segments on the timeline.

20. An electronic device, comprising: At least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 18 when executed by the at least one processing unit.

21. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 18.