Method, apparatus, device and storage medium for generating video

CN122802745APending Publication Date: 2026-09-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344951.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0007]应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802745A_ABST
    Figure CN122802745A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, apparatus, device and storage medium for generating a video. The method comprises: in response to selection of a media template, presenting a configuration panel; determining configuration information via the configuration panel, the configuration information comprising reference material obtained via a material adding control and text content obtained via a text input control; and providing video content generated based on the media template and the configuration information, wherein the video content comprises audio content corresponding to the text content and picture content, the picture content comprising a dynamic process of a virtual object, the dynamic process corresponding to the audio content, and the video content further comprising at least one material segment corresponding to the reference material, the at least one material segment being determined based on a relevance of the reference material to the text content. In this way, embodiments of the present disclosure can quickly generate video content associated with a virtual object, improving video generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, and computer-readable storage media for generating video. Background Technology

[0002] With the development of computer technology, people can share or access media content through internet platforms. Media editing tools have also gradually become important tools in people's lives; for example, people can use media editing tools to generate or edit video content and share it through media platforms. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for generating video is provided. The method includes: presenting a configuration panel in response to the selection of a media template; determining configuration information via the configuration panel, the configuration information including reference material obtained via a material addition control and text content obtained via a text input control; and providing video content generated based on the media template and the configuration information, wherein the video content includes audio content corresponding to the text content and visual content, the visual content including a dynamic process of a virtual object corresponding to the audio content, and the video content also includes at least one material segment corresponding to the reference material, the at least one material segment being determined based on the relevance between the reference material and the text content.

[0004] In a second aspect of this disclosure, an apparatus for generating video is provided. The apparatus includes: a panel presentation module, an information determination module, and a video generation module, wherein the panel presentation module is configured to present a configuration panel in response to the selection of a media template; the information determination module is configured to determine configuration information via the configuration panel, the configuration information including reference material obtained via a material addition control and text content obtained via a text input control; the video generation module is configured to provide video content generated based on the media template and the configuration information, wherein the video content includes audio content corresponding to the text content and visual content associated with a virtual object, the visual content including a dynamic process of the virtual object corresponding to the audio content, and the video content also includes at least one material segment corresponding to the reference material, the at least one material segment being determined based on the relevance between the reference material and the text content.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0010] Figures 2A to 2F Example interfaces according to some embodiments of this disclosure are shown;

[0011] Figure 3 A flowchart illustrating an example process for generating video according to some embodiments of this disclosure is shown;

[0012] Figure 4 A schematic structural block diagram of an example apparatus for generating video according to some embodiments of the present disclosure is shown; and

[0013] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0019] As mentioned above, media editing tools have gradually become important tools in people's lives. For example, people can use media editing tools to edit video content and share it through media platforms.

[0020] In the traditional editing process or before editing, people usually need to collect the necessary materials themselves, such as choosing suitable scenes or backgrounds, selecting suitable on-camera personnel and their clothing and decorations, in order to shoot satisfactory video materials. Then, they add multiple materials to the editing track in the editing interface and adjust them one by one to obtain satisfactory video content. This greatly affects the efficiency of media editing.

[0021] Embodiments of this disclosure propose a scheme for generating video. The scheme includes: presenting a configuration panel in response to the selection of a media template; determining configuration information via the configuration panel, the configuration information including reference material obtained via a material addition control and text content obtained via a text input control; providing video content generated based on the media template and the configuration information, wherein the video content includes audio content and visual content corresponding to the text content, the visual content including a dynamic process of a virtual object corresponding to the audio content; and the video content also includes at least one material segment corresponding to the reference material, the at least one material segment being determined based on the relevance between the reference material and the text content.

[0022] In this way, the embodiments of this disclosure can quickly configure reference materials and text content based on media templates, and quickly generate video content accordingly. The generated video content includes visual and audio content associated with virtual objects, effectively ensuring the correlation between text content and materials in the video and virtual objects and visual content, ensuring video generation quality, improving video generation efficiency, and enhancing user experience.

[0023] The following section provides a detailed description of various example implementations of this scheme, in conjunction with the accompanying drawings.

[0024] Example Environment

[0025] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.

[0026] In this example environment 100, electronic device 110 may run an application 120 that supports video generation. Application 120 can be any suitable type of application for generating video, examples of which may include, but are not limited to, audio / video sharing applications, media editing or creation applications, or other suitable applications that provide video generation services. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.

[0027] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting video generation.

[0028] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0029] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for the application 120 in electronic device 110 that supports video generation.

[0030] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.

[0031] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0032] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0033] Example Interaction

[0034] Figures 2A to 2F Example interfaces 200A to 200F according to some embodiments of the present disclosure are shown. Interfaces 200A to 200F may, for example, be provided by... Figure 1 The electronic device 110 shown is provided.

[0035] refer to Figure 2A As shown, Figure 2A Example user interface 200A according to some embodiments of the present disclosure is shown. As an example, electronic device 110 may provide user interface 200A. For example, electronic device 110 may present example user interface 200A based on the response to operations of its installed application 120, or based on the response to at least one operation of application 120.

[0036] As an example, the interactive interface 200A can also serve as the entry point for creating works. Through this interface, users can create works, such as graphic works, video content, and audio works.

[0037] In some embodiments, the electronic device 110 may provide a set of creation entry points 210 in the interactive interface 200A to create works or work materials via at least one of the creation entry points. For example, creating video content or video materials.

[0038] As an example, users can create video content or other media works through the "Quick Creation" entry; they can identify subtitle text in video content through the "Subtitle Recognition" entry to obtain subtitle materials or text materials; and they can obtain audio materials or audio works through the "Text Reading" entry.

[0039] In some embodiments, the electronic device 110 creates video content associated with a virtual object in response to a triggering of a target creation entry 215 in a set of creation entries 210. As an example, the virtual object may include objects such as virtual digital humans or digital animals.

[0040] In some embodiments, the electronic device 110 can, based on the triggering of the target creation entry 215, generate a set of media templates associated with virtual objects. For example... Figure 2B The interface shown is 200B.

[0041] refer to Figure 2B As shown, the electronic device 110 presents a set of media templates 220 in the template display interface 200B.

[0042] As an example, the electronic device 110 may provide at least one template tag 220a based on template content or template type in the interface 200B, such as knowledge, festivals, or information. Based on the triggering of any template tag 220a, the electronic device 110 may display a media template corresponding to that template tag 220a. Users can quickly filter media templates by triggering any of the template tags 220a.

[0043] In some embodiments, taking the first media template 225 as an example, for a media template presented in the interface 200B, the electronic device 110 can display information such as the media cover and content name of the media template.

[0044] like Figure 2B As shown, the electronic device 110 can also display the creation control 225a at a cover-related location of the first media template 225 (e.g., at the bottom or below the media cover). As an example, the creation control 225a can serve as an application control or playback control for the first media template 225.

[0045] In some embodiments, the electronic device 110 may present a playback interface of the first media template 225 based on the triggering of the creation control 225a. The electronic device 110 may also provide creation controls for converting the first media template 225 into a work within the playback interface of the first media template 225.

[0046] In some embodiments, the electronic device 110 may also, based on the triggering of the authoring control 225a, present a media creation interface or media configuration panel based on the first media template 225. As an example, the electronic device 110 may present... Figure 2C The interface shown is 200C.

[0047] refer to Figure 2C As shown, the electronic device 110 presents a configuration panel 230 based on the first media template 225 in the interface 200C. Through the configuration panel 230, the electronic device 110 can receive at least one configuration operation from the user, thereby determining the configuration information associated with the first media template 225, so as to generate corresponding video content or reference material for video content based on the configuration information.

[0048] In some embodiments, the electronic device 110 presents at least one configuration control in the configuration panel 230.

[0049] like Figure 2CAs shown, the electronic device 110 can provide at least one input control, such as a text input control 231a and an audio input control 231b. As an example, the text input control 231a can be presented in the configuration panel 230 as "Script", "Input Text", "Input Text", etc., and the audio input control 231b can be presented as "Upload Audio", "Input Voice", etc., without limitation.

[0050] As an example, electronic device 110 can receive reference text input by a user via text input control 231a. Electronic device 110 can directly use the received reference text as text content, or it can obtain text content based on the received reference text using a pre-configured generative model. The text content is text that the user can add to the video content.

[0051] As an example, via audio input control 231b, electronic device 110 can receive reference audio uploaded by a user and determine the corresponding text content based on the reference audio. For example, electronic device 110 performs speech recognition and text generation on the reference audio to obtain the corresponding reference text. As another example, electronic device 110 can input the reference audio into a pre-configured recognition model to obtain the corresponding text content.

[0052] In some embodiments, users can upload audio files from local or cloud paths using the audio input control 231b, or record audio files using the audio input control 231b and upload them.

[0053] like Figure 2C As shown, the electronic device 110 can also display a preview box 232 corresponding to the content type option in the configuration panel 230. Through the preview box 232, the electronic device 110 can display the indicator elements corresponding to the received reference text / or reference audio, such as displaying the file identifier corresponding to the text content or audio file.

[0054] In some embodiments, the electronic device 110 may also display a subtitle generation control 233 in a configuration panel 230. Based on whether the subtitle generation control 233 is selected, the electronic device 110 can determine configuration information regarding the generated subtitle content.

[0055] As an example, in response to the selection of the subtitle generation control 233, the electronic device 110 can determine that subtitle content is included in the configuration information determined based on the first media template 225, and that the subtitle content is determined based on the text content determined in the aforementioned process. The subtitle style of the subtitle content can be the same as the subtitle style in the first media template 225.

[0056] Additionally, the electronic device 110 can also display a subtitle keyword control in the configuration panel 230. For example, "Subtitle highlighting".

[0057] In some embodiments, in response to the activation of the subtitle keyword control, the electronic device 110 determines a set of keywords from the text content and, based on the subtitle styles corresponding to the keywords in the first media template 225, determines a first subtitle style corresponding to the set of keywords. Based on this, the electronic device 110 applies the first subtitle style to the generated subtitle content corresponding to the text content.

[0058] In some embodiments, text other than keywords in the text content is presented in the subtitle content in a second subtitle style that differs from the first subtitle style. As an example, the second subtitle style may be the presentation style used for subtitles other than keywords in the first media template 225.

[0059] In some embodiments, in response to the subtitle keyword being turned off, the electronic device 110 can determine the subtitle style in the first media template 225 as the presentation style corresponding to the subtitle content generated based on the text content.

[0060] like Figure 2C As shown, the electronic device 110 can also display media configuration controls in the configuration panel 230. As an example, the media configuration controls may include a media addition control 234a.

[0061] In some embodiments, via the material addition control 234a, the electronic device 110 can receive at least one reference material and display the material identifier corresponding to the reference material. For example... Figure 2D Material identifiers 234a-1 and 234a-2 are shown.

[0062] refer to Figure 2D As shown, the electronic device 110 displays material identifiers 234a-1 and 234a-2 in the configuration panel 230, and each material identifier displays a corresponding removal control. In response to the triggering of the removal control corresponding to any material identifier, the electronic device 110 removes that material identifier and the reference material corresponding to that material identifier.

[0063] Return to reference Figure 2C As shown, the material configuration control may also include a material display control 234b. Through this material display control 234b, users can configure how the materials are displayed.

[0064] In some embodiments, the first media template 225 can be any one of a real-scene template, a green screen template, or a container template.

[0065] Taking a real-scene template as an example, the virtual objects in the template image and the template background are an integrated structure. In the video content generated based on the real-scene template, the added material content can be displayed as the foreground element of the virtual object, or it can be displayed as a template image to replace the video frame.

[0066] Taking a green screen template as an example, the virtual object and the template background are independent materials. In the video content generated based on the green screen background, the added materials can be displayed as foreground elements of the virtual object, or as background elements of the virtual object between the virtual object and the template background. They can also replace the virtual object and the template background to be displayed as a complete image of the video frame.

[0067] Taking a container template as an example, the template image provides a preset display area for displaying added materials. As an example, the preset display area can be a fixed area or a variable area in the template image, it can be the foreground or background area of ​​a virtual object, or at least part of the limb area of ​​a virtual object such as a digital human, and there are no limitations here.

[0068] Comprehensive reference Figure 2C and Figure 2D As shown, the electronic device 110 can present at least one material display option 234b-1 based on the triggering of the material display control 234b. For example, the "Digital Human Foreground" option and the "Digital Human Background" option.

[0069] As an example, in response to the selection of the "digital human foreground" option, the electronic device 110 determines the configuration information of the reference material corresponding to the material identifier 234a-1 and the material identifier 234a-2 as follows: the reference material is displayed as the foreground content of the virtual object (e.g., digital human) in the screen content, that is, the reference material can occlude at least part of the screen of the virtual object in the screen content.

[0070] In some embodiments, after selecting a media display option, the electronic device 110, in response to a trigger on the media display control 234b, can collapse the media display option again, displaying only the media display control 234b. The display effect can be seen from [reference needed]. Figure 2E As shown in the image.

[0071] Return to reference Figure 2C As shown, the electronic device 110 can also display more settings controls 235 in the configuration panel 230. As an example, based on the triggering of the more settings control 235, the electronic device 110 can expand to display at least one other configuration operation control, such as... Figure 2D and Figure 2E The "Change Timbre" and "Change Image" configuration controls are shown in the image.

[0072] refer to Figure 2D and Figure 2E As shown, based on the triggering of the more settings control 235, the electronic device 110 can present a preset set of candidate sound styles. As an example, the electronic device 110 can present a set of sound styles corresponding to at least one sound indicator element 235a.

[0073] As an example, the sound indicator element 235a corresponding to a sound style may include an image icon and a style name. The style name may include a little girl, a little boy, a gentle girl, a deep male voice, an anime character, etc. The image icon may be a virtual image corresponding to the style name.

[0074] In response to the selection of a target sound indicator element 235a-1 in a set of sound indicator elements 235a, the electronic device 110 uses the sound style indicated by the target sound indicator element 235a-1 as at least part of the sound configuration information corresponding to the audio content.

[0075] Based on the triggering of the more settings control 235 in the configuration panel 230, the electronic device 110 can also present a set of virtual images 235b as visual images of virtual objects.

[0076] In some embodiments, in response to the selection of a target virtual image 235b-1 in a set of virtual images 235b, the electronic device 110 may replace the corresponding current virtual image in the first media template 225 with the target virtual image 235b-1, and determine the second virtual object and image parameters indicated by the target virtual image, so as to use the image parameters as at least part of the configuration information of the virtual object in the screen content.

[0077] Comprehensive reference Figures 2C-2E As shown, the electronic device 110 provides a video generation control 236 in the configuration panel 230. In response to triggering the video generation control 236, the electronic device 110 can generate corresponding video content by combining the configuration information determined above and the first media template 225. For example... Figure 2F The video content 201 in the interactive interface 200F shown.

[0078] In some embodiments, a generative model deployed on electronic device 110 or server 130 may be used to generate video content based on the above configuration information.

[0079] As an example, the generated video content may include audio and video content corresponding to the text content, as well as subtitle content, etc.

[0080] In some embodiments, the screen content includes the dynamic process of a virtual object, which corresponds to the audio content. For example, the screen content may include the facial movement process of the virtual object, such as dynamic changes in lip shape, eyelids, and cheeks in response to the audio content; it may also include the limb movement process of the virtual object, such as changes in body posture, gestures, and head position. The virtual object may include a virtual digital human or animal.

[0081] In some embodiments, based on the triggering of the video generation control 236, the electronic device 110 may also present an information confirmation panel before presenting the interactive interface 20, to present a prompt message and a confirmation control. As an example, the prompt message may be a prompt message about the virtual points required to generate the video content 201. In response to the triggering of the confirmation control, the electronic device 110 regenerates the video content 201 and presents the interactive interface 200F of the video content 201.

[0082] refer to Figure 2F As shown, based on the triggering of the video generation control 236, the electronic device 110 can present an interactive interface 200F of the generated video content 201. In this interactive interface 200F, the electronic device 110 can present the audio content, video content, subtitle content, etc. of the video content 201 frame by frame.

[0083] In some embodiments, the electronic device 110 may also present a set of editing controls 240 in the interactive interface 200F for editing at least a portion of the video content 201.

[0084] As an example, the electronic device 110 can receive editing operations on at least one of the following in the interactive interface 200F: text content, audio content, screen content, subtitle content, presentation effects, etc., for the video content 201, without limitation.

[0085] Virtual object driver

[0086] In some embodiments, electronic device 110 or server 130 may determine audio control signals based on text content and may provide the generative model with images of virtual objects to be used (e.g., virtual objects corresponding to media templates or virtual objects selected by the user) and audio control signals to generate video content.

[0087] Specifically, an audio generation model can be used to generate audio content corresponding to a preset timbre based on the broadcast text. As an example, the preset timbre could be the timbre used in a media template. Alternatively, the preset timbre could be another timbre selected by the user. Furthermore, the generated audio content can provide audio control signals as a generative model to drive virtual objects.

[0088] In some embodiments, the model may include multiple attention layers. The attention layers may include multimodal attention units. Specifically, where the control signal includes audio content, audio features of the audio content can be determined. As an example, the audio features may include Mel-spectral features of the audio content. Further, an encoding unit may be used to process the audio features to obtain audio tokens. As an example, the encoding unit may include, for example, a multilayer perceptron (MLP) or other suitable model.

[0089] Furthermore, the cross-attention unit can update the model's input video features based on the cross-attention mechanism and audio tokens. The updated input video features can then be provided to the multimodal attention unit in the next attention layer. In this way, after processing through multiple attention layers, the model can output the final video features to generate the final video content.

[0090] Example process

[0091] Figure 3 A flowchart of an example process 300 for generating video according to some embodiments of the present disclosure is shown. Process 300 can be implemented at electronic device 110. Reference is made below. Figure 1 To describe process 300.

[0092] like Figure 3 As shown in box 310, electronic device 110 displays a configuration panel in response to the selection of a media template.

[0093] As an example, you can refer to the media template. Figure 2B The configuration panel of the first media template 225 shown can be referenced. Figures 2C-2E The configuration panel 230 shown.

[0094] Continue to refer to Figure 3 As shown, in box 320, electronic device 110 determines configuration information via a configuration panel. This configuration information includes reference materials obtained via a material addition control and text content obtained via a text input control.

[0095] As an example, reference materials can be found here. Figure 2DThe text content of the materials indicated by material identifiers 234a-1 and 234a-2 shown can be referenced via... Figure 2C The reference text received by the “Input Text” control 231a shown, or the reference text determined based on the reference audio uploaded via the “Upload Audio” control 231b.

[0096] Continue to refer to Figure 3 As shown in box 330, electronic device 110 provides video content generated based on media templates and configuration information. The video content includes audio content and visual content corresponding to text content. The visual content includes the dynamic process of virtual objects, which corresponds to the audio content. Furthermore, the video content also includes at least one clip corresponding to reference material, the at least one clip being determined based on the relevance between the reference material and the text content.

[0097] As an example, the video content can be referenced. Figure 2F The video content shown is 201.

[0098] In this way, the embodiments of this disclosure can quickly configure reference materials and text content based on media templates, and quickly generate video content accordingly. The generated video content includes visual and audio content associated with virtual objects, effectively ensuring the correlation between text content and materials in the video and virtual objects and visual content, ensuring video generation quality, improving video generation efficiency, and enhancing user experience.

[0099] In some embodiments of this disclosure, the virtual object includes: a first virtual object corresponding to a media template; or a second virtual object indicated by configuration information.

[0100] In this way, the embodiments of this disclosure can effectively expand the diversity of virtual objects, enrich video content, and improve user experience.

[0101] In some embodiments of this disclosure, determining configuration information via a configuration panel includes: determining a material display mode via a configuration panel, wherein the material display mode indicates that at least one material fragment is displayed in the foreground or background layer of a virtual object.

[0102] In this way, the present invention can configure the display hierarchy between material clips and virtual objects, ensuring the display effect of material clips in video content and improving video quality and user experience.

[0103] In some embodiments of this disclosure, determining configuration information via the configuration panel includes: obtaining text content via a text input control in the configuration panel; or obtaining reference audio via an audio input control in the configuration panel to determine the text content corresponding to the reference audio.

[0104] As an example, a text input control can be referenced. Figure 2C The “Input Text” control 231a shown here, and the audio input control can be referenced. Figure 2C The “Upload Audio” control 231b is shown.

[0105] In this way, the embodiments of this disclosure can determine text content based on text input controls or audio input controls, expand the file types and methods for determining text content, improve the efficiency of text content determination, and enhance user experience.

[0106] In some embodiments of this disclosure, determining configuration information via a configuration panel includes: presenting a set of sound styles in the configuration panel; and receiving a selection of a first sound style from the set of sound styles, wherein the first sound style is used to generate audio content corresponding to the text content.

[0107] As an example, you can refer to the sound style. Figure 2D and Figure 2E The set of indicator elements 235a shown corresponds to the sound style.

[0108] In this way, the embodiments of this disclosure can select and configure the sound style of audio content, effectively expand the sound style of audio content, diversify the selection of sound style of audio content, and improve user experience.

[0109] In some embodiments of this disclosure, at least one material also includes subtitle content corresponding to the text content, and the subtitle style of the subtitle content is determined based on a media template.

[0110] In this way, the embodiments of this disclosure can quickly determine the subtitle content and its style corresponding to the text content based on the media template, thereby improving the richness of the material and enhancing the user experience.

[0111] In some embodiments of this disclosure, the subtitle content is generated based on the following process: determining a set of keywords from the text content; determining a first subtitle style corresponding to the set of keywords based on a media template; and generating subtitle content corresponding to the text content, wherein the set of keywords is applied to the first subtitle style.

[0112] In this way, the embodiments of this disclosure can quickly locate keywords from text content and highlight the keywords in the style of a first subtitle, thereby improving the eye-catching effect of the keywords, helping users to quickly identify the keywords, and enhancing the user experience.

[0113] In some embodiments of this disclosure, a second caption style is applied to a portion of the text content other than a set of keywords, the second caption style being determined based on a media template.

[0114] In this way, the embodiments of this disclosure can quickly distinguish the subtitle styles of keywords and other parts, thereby improving the efficiency of subtitle content generation.

[0115] In some embodiments of this disclosure, determining configuration information via a configuration panel includes: presenting at least one virtual object in the configuration panel; and receiving a selection of a second virtual object from the at least one virtual object.

[0116] As an example, at least one virtual object can be referenced. Figure 2E The virtual objects corresponding to a set of virtual images 235b shown.

[0117] In this way, the embodiments of this disclosure can configure or replace the first virtual object corresponding to the media template, effectively expanding the application methods of the media template, diversifying the selection and configuration of virtual objects, enriching video content, and improving user experience.

[0118] In some embodiments of this disclosure, at least one virtual object includes: a preset virtual object; and / or a virtual object created based on the current user's configuration operations.

[0119] In this way, embodiments of this disclosure can select preset virtual objects or allow users to freely configure and create virtual objects, effectively expanding the diversity of virtual objects, enriching video content, and improving user experience.

[0120] In some embodiments of this disclosure, the style information of at least one material segment in the video content is determined based on a media template.

[0121] In this way, the embodiments of this disclosure can automatically and quickly determine the style information of at least one material segment in the video content based on the media template, improve the configuration efficiency of the style information corresponding to the material segment, improve the generation efficiency of the video content, and enhance the user experience.

[0122] In some embodiments of this disclosure, the style information of at least one material segment in the video content indicates at least one of the following: the presentation position of at least one material segment; the dynamic effect of at least one material segment; and the display size of at least one material segment.

[0123] In this way, the embodiments of this disclosure can quickly configure at least one piece of information such as the presentation position, dynamic effects, and display size of the material clip in the video content, thereby improving the configuration efficiency of the style information corresponding to the material clip, improving the generation efficiency of the video content, and improving the user experience.

[0124] Example devices and equipment

[0125] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4A schematic structural block diagram of an example device 400 for generating video according to certain embodiments of the present disclosure is shown. Device 400 may be implemented as or included in electronic device 110. Various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0126] like Figure 4 As shown, the device 400 includes: a panel presentation module 410, an information determination module 420, and a video generation module 430. The panel presentation module 410 is configured to present a configuration panel in response to the selection of a media template; the information determination module 420 is configured to determine configuration information via the configuration panel, the configuration information including reference material obtained via a material addition control and text content obtained via a text input control; the video generation module 430 is configured to provide video content generated based on the media template and the configuration information, wherein the video content includes audio content and visual content corresponding to the text content, the visual content including a dynamic process of a virtual object corresponding to the audio content, and the video content also including at least one material segment corresponding to the reference material, the at least one material segment being determined based on the relevance between the reference material and the text content.

[0127] In some embodiments of this disclosure, the virtual object includes: a first virtual object corresponding to a media template; or a second virtual object indicated by configuration information.

[0128] In some embodiments of this disclosure, determining configuration information via a configuration panel includes: determining a material display mode via a configuration panel, wherein the material display mode indicates that at least one material fragment is displayed in the foreground or background layer of a virtual object.

[0129] In some embodiments of this disclosure, determining configuration information via the configuration panel includes: obtaining text content via a text input control in the configuration panel; or obtaining reference audio via an audio input control in the configuration panel to determine the text content corresponding to the reference audio.

[0130] In some embodiments of this disclosure, determining configuration information via a configuration panel includes: presenting a set of sound styles in the configuration panel; and receiving a selection of a first sound style from the set of sound styles, wherein the first sound style is used to generate audio content corresponding to the text content.

[0131] In some embodiments of this disclosure, at least one material also includes subtitle content corresponding to the text content, and the subtitle style of the subtitle content is determined based on a media template.

[0132] In some embodiments of this disclosure, the subtitle content is generated based on the following process: determining a set of keywords from the text content; determining a first subtitle style corresponding to the set of keywords based on a media template; and generating subtitle content corresponding to the text content, wherein the set of keywords is applied to the first subtitle style.

[0133] In some embodiments of this disclosure, a second caption style is applied to a portion of the text content other than a set of keywords, the second caption style being determined based on a media template.

[0134] In some embodiments of this disclosure, determining configuration information via a configuration panel includes: presenting at least one virtual object in the configuration panel; and receiving a selection of a second virtual object from the at least one virtual object.

[0135] In some embodiments of this disclosure, at least one virtual object includes: a preset virtual object; and / or a virtual object created based on the current user's configuration operations.

[0136] In some embodiments of this disclosure, the style information of at least one material segment in the video content is determined based on a media template.

[0137] In some embodiments of this disclosure, the style information of at least one material segment in the video content indicates at least one of the following: the presentation position of at least one material segment; the dynamic effect of at least one material segment; and the display size of at least one material segment.

[0138] In the apparatus 400 provided in this embodiment, the specific processing procedures of each module and the technical effects they bring can be referred to the relevant description of the example process 300 above, and will not be repeated here.

[0139] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0140] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0141] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0142] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0143] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0144] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0145] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0146] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0147] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0149] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating video, comprising: The configuration panel is displayed in response to the selection of a media template; Configuration information is determined via the configuration panel, including reference materials obtained via the material addition control and text content obtained via the text input control; and Provide video content generated based on the media template and the configuration information, wherein the video content includes audio content and visual content corresponding to the text content, the visual content includes the dynamic process of a virtual object, the dynamic process corresponding to the audio content, and the video content also includes at least one material segment corresponding to the reference material, the at least one material segment being determined based on the relevance between the reference material and the text content.

2. The method according to claim 1, wherein, The virtual object includes: The first virtual object corresponding to the media template; or The second virtual object indicated by the configuration information.

3. The method according to claim 1, wherein, The configuration information determined via the configuration panel includes: The material display mode is determined via the configuration panel, and the material display mode indicates that the at least one material fragment is displayed in the foreground layer or background layer of the virtual object.

4. The method according to claim 1, wherein, The configuration information determined via the configuration panel includes: The text content is obtained via the text input control in the configuration panel; or A reference audio is obtained via the audio input control in the configuration panel to determine the text content corresponding to the reference audio.

5. The method according to claim 1, wherein, The configuration information determined via the configuration panel includes: A set of sound styles is presented in the configuration panel; and The system receives a selection of a first sound style from the set of sound styles, wherein the first sound style is used to generate the audio content corresponding to the text content.

6. The method according to claim 1, wherein, The at least one material also includes subtitle content corresponding to the text content, and the subtitle style of the subtitle content is determined based on the media template.

7. The method according to claim 6, wherein, The subtitle content was generated based on the following process: Identify a set of keywords from the text content; Based on the media template, determine the first subtitle style corresponding to the set of keywords; as well as Generate subtitle content corresponding to the text content, wherein the set of keywords is applied to the first subtitle style.

8. The method according to claim 7, wherein, The portion of the text content other than the set of keywords is subject to a second subtitle style, which is determined based on the media template.

9. The method according to claim 1, wherein, The process of determining configuration information via the configuration panel includes: At least one virtual object is presented in the configuration panel; and Receive a selection of a second virtual object from the at least one virtual object.

10. The method according to claim 9, wherein, The at least one virtual object includes: Preset virtual objects; and / or Virtual objects created based on the current user's configuration operations.

11. The method according to claim 1, wherein, The style information of at least one material segment in the video content is determined based on the media template.

12. The method according to claim 11, wherein, The style information indicates at least one of the following: The presentation position of the at least one material fragment; The dynamic effects of at least one material segment; The display size of the at least one material fragment.

13. An apparatus for generating video, comprising: The panel presentation module is configured to display a configuration panel in response to the selection of a media template; The information determination module is configured to determine configuration information via the configuration panel, the configuration information including reference materials obtained via the material addition control and text content obtained via the text input control; and The video generation module is configured to provide video content generated based on the media template and the configuration information, wherein the video content includes audio content and visual content corresponding to the text content, the visual content includes a dynamic process of a virtual object, the dynamic process corresponding to the audio content, and the video content also includes at least one material segment corresponding to the reference material, the at least one material segment being determined based on the relevance between the reference material and the text content.

14. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processing unit.

15. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 12.