Video generation method and device, equipment and storage medium

By acquiring audio-related video generation requests and template resources, dynamically switching subtitle text fragments, and generating videos that match the audio, the problem of low video generation efficiency is solved, and efficient video generation is achieved.

CN122073628APending Publication Date: 2026-05-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies have low video generation efficiency, making it difficult to efficiently generate audio-related video content.

Method used

By obtaining a video generation request associated with the target audio, subtitle text fragments are generated using resources in the first video template, and video is generated by combining audio and image materials to achieve dynamic switching effects of subtitle text.

Benefits of technology

It improves the efficiency of video generation, enabling the rapid generation of high-quality video content that matches the audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073628A_ABST
    Figure CN122073628A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a video generation method and device, equipment and a storage medium. The method provided by the invention comprises the following steps: acquiring a video generation request associated with target audio; in response to the video generation request, obtaining a first group of resources corresponding to the subtitle text in the first video template; the subtitle text is text information associated with the target audio, and the first group of resources is used for indicating a first editing effect of the subtitle text; based on the first group of resources, generating a first character material fragment corresponding to the subtitle text, the first character material fragment being used for presenting an effect of dynamically switching the subtitle text fragment according to the occurrence time of the subtitle text fragment in the target audio; and generating a first video corresponding to the target audio according to the first character material segment, the target audio and a second group of resources in the first video template, the second group of resources comprising video materials and / or picture materials, which are matched with the target audio and / or the subtitle text. According to the invention, the video generation efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, and computer-readable storage media for generating video. Background Technology

[0002] With the rapid development of smart technology, various forms of electronic devices are greatly enriching people's daily lives. For example, users can interact with various electronic devices, such as generating videos from audio. Improving the efficiency of video generation is a key concern. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for generating a video is provided. The method includes: acquiring a video generation request associated with target audio; in response to the video generation request, acquiring a first set of resources in a first video template corresponding to subtitle text; the subtitle text being text information associated with the target audio, and the first set of resources being used to indicate a first editing effect of the subtitle text; based on the first set of resources, generating a first text material fragment corresponding to the subtitle text, the first text material fragment being used to present an effect of dynamically switching subtitle text fragments according to the time of appearance of the subtitle text fragment in the target audio; and generating a first video corresponding to the target audio based on the first text material fragment, the target audio, and a second set of resources in the first video template, the second set of resources including at least one first image material, the at least one first image material including video material and / or image material, and the at least one first image material matching the target audio and / or subtitle text.

[0004] In a second aspect of this disclosure, an apparatus for generating video is provided. The apparatus includes: a first acquisition module configured to acquire a video generation request associated with target audio; a second acquisition module configured to, in response to the video generation request, acquire a first set of resources in a first video template corresponding to subtitle text; the subtitle text being text information associated with the target audio, and the first set of resources used to indicate a first editing effect of the subtitle text; a first generation module configured to, based on the first set of resources, generate a first text material fragment corresponding to the subtitle text, the first text material fragment being used to present an effect of dynamically switching subtitle text fragments according to the time of appearance of the subtitle text fragment in the target audio; and a second generation module configured to, based on the first text material fragment, the target audio, and a second set of resources in the first video template, generate a first video corresponding to the target audio, the second set of resources including at least one first image material, the at least one first image material including video material and / or image material, and the at least one first image material matching the target audio and / or subtitle text.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0010] Figures 2A to 2F A schematic diagram of an example interface according to some embodiments of the present disclosure is shown;

[0011] Figure 2G A flowchart illustrating the generation of music videos according to some embodiments of this disclosure is shown;

[0012] Figure 3 A flowchart illustrating a process for generating video according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic structural block diagram of an apparatus for generating video according to certain embodiments of the present disclosure is shown; and

[0014] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0020] Embodiments of this disclosure propose a scheme for generating video. According to various embodiments of this disclosure, a video generation request associated with a target audio can be obtained; in response to the video generation request, a first set of resources corresponding to subtitle text in a first video template can be obtained; the subtitle text is text information associated with the target audio, and the first set of resources is used to indicate a first editing effect of the subtitle text; based on the first set of resources, a first text material fragment corresponding to the subtitle text is generated, the first text material fragment being used to present the effect of dynamically switching subtitle text fragments according to the time of appearance of the subtitle text fragment in the target audio; and based on the first text material fragment, the target audio, and a second set of resources in the first video template, a first video corresponding to the target audio is generated, the second set of resources including at least one first image material, the at least one first image material including video material and / or image material, and the at least one first image material matching the target audio and / or subtitle text.

[0021] Therefore, the embodiments of this disclosure can generate a video corresponding to the target audio based on the first video template, which can effectively improve the video generation efficiency.

[0022] Example Environment

[0023] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.

[0024] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction, examples of which may include, but are not limited to, video creation applications or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.

[0025] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.

[0026] In some embodiments, electronic device 110 communicates with server 130 (target device) to provide services to application 120. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0027] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0028] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.

[0029] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0030] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0031] Example Interaction

[0032] The following description, with reference to the accompanying drawings, illustrates an example interaction process according to an embodiment of the present disclosure. Figures 2A to 2F Example interfaces 200A to 200F according to some embodiments of the present disclosure are shown. Interfaces 200A to 200F can be provided by... Figure 1 The electronic device 110 shown is provided.

[0033] In some embodiments, the electronic device 110 may provide an interface 200A, which can be an audio generation interface, allowing users to input the materials needed to generate the audio, such as the text corresponding to the audio, the descriptive information corresponding to the audio, etc. The audio can correspond to any appropriate type, such as music audio, podcast audio, recitation audio, speech audio, etc.

[0034] For ease of description, this application uses audio as an example to illustrate the music generation process, wherein the materials required to generate music may include, but are not limited to, lyrics, musical styles, etc.

[0035] In some embodiments, the electronic device 110 may present an input box 201 in the interface 200A. As an example, a user can input lyrics for a piece of music to be generated based on the input box 201, at which point the electronic device 110 can receive the lyrics input by the user. As another example, a user can input keywords based on the input box 201, allowing the electronic device 110 to generate lyrics for the music to be generated based on the keywords. Specifically, the electronic device 110 may generate lyrics associated with a keyword based on a first target model. The first target model can be any suitable machine learning model, such as a generative model, etc.

[0036] In some embodiments, the electronic device 110 may display an input box 202 corresponding to the music description in the interface 200A to support users in inputting descriptive information such as the style and duration of the music to be generated. As an example, the user can input the music style corresponding to the music to be generated based on the input box 202, and the electronic device 110 can receive the music style input by the user.

[0037] In other embodiments, the electronic device 110 may also present a set of candidate controls corresponding to the music style in the interface 200A (e.g., Figure 2A The controls 203-1, 203-2, and 203-3 shown can correspond to different music styles. The electronic device 110 can, in response to the selection of a target control within this set of candidate controls, determine the music style associated with that target control as the music style corresponding to the music to be generated.

[0038] In some embodiments, the electronic device 110 can generate music associated with the received lyrics and / or music description. For example, a user can input lyrics and a music description using input boxes 201 and 202 in interface 200A, respectively, and click the "Start Generation" control displayed in interface 200A. The electronic device 110 can then generate music associated with the input lyrics and music description. Specifically, the electronic device 110 can utilize a second target model to generate music based on the lyrics and music description. The second target model can be any suitable machine learning model, such as a generative model, etc.

[0039] In some embodiments, the music generated by the electronic device 110 and associated with lyrics and / or musical descriptions can be multiple, such as... Figure 2B As shown, the electronic device 110 can present a set of generated candidate music in the interface 200B, which may specifically include music 210-1, music 210-2, and music 210-3, etc.

[0040] Furthermore, the electronic device 110 can, in response to receiving a selection of a target music from this set of candidate music, obtain a video generation request associated with the target music. Figure 2B As an example, the user can click the "Use" control corresponding to music 210-1 and click the video generation control 203 in interface 200B. At this time, electronic device 110 can obtain the video generation request associated with music 210-1.

[0041] by Figure 2B as well as Figure 2C As an example, electronic device 110 may, in response to a request to generate video associated with music 210-1, present... Figure 2C The interface 200C is shown. Interface 200C can display progress information about the music video generation, indicating that the music video is being generated (or loading). Furthermore, the electronic device 110 can also display the specific generation progress of the music video on interface 200C, such as displaying a progress bar or the percentage of the music video already generated.

[0042] It should be noted that, in addition to receiving video generation requests associated with the target music, the electronic device 110 can also receive video requests associated with other types of audio (such as recitation audio, etc.) besides the target music. This application does not restrict the way video requests are received.

[0043] To improve video generation efficiency, the electronic device 110 can use a video template (first video template) as a reference to generate a first video draft corresponding to the target audio, and then generate the video corresponding to the target audio based on the first video draft. The video template can include the basic structure of the video draft corresponding to the video to be generated, and this video template can include predetermined replaceable information. For example, taking a music-related template as an example, the author, song title, lyrics, and other information included in this video template are temporary placeholders. Subsequently, these temporary placeholders can be replaced based on the relevant information of the target music (which is associated with the video generation request) to generate a music video corresponding to the target music.

[0044] In some embodiments, the first video template can be a preset video template.

[0045] In other embodiments, electronic device 110 may send the target audio to server 130. Server 130 may determine a first video template associated with the target audio from a set of candidate video templates based on at least one attribute of the target audio. The at least one attribute may include, but is not limited to, text associated with the target audio (such as lyrics to music, poems in a poetry recitation, etc.), the audio style corresponding to the target audio, and other attributes.

[0046] In other embodiments, the electronic device 110 may present a set of candidate video templates. Further, the electronic device 110 may determine a first video template associated with the target audio in response to a user's selection of a target template from this set of candidate video templates.

[0047] In some embodiments, this set of candidate video templates can be generated based on the following process:

[0048] Server 130 can obtain a set of video drafts (whose initial form can be video) edited by the user using a pre-defined editing tool. Server 130 can reverse engineer the protocol of the corresponding video based on this set of video drafts. This protocol can at least indicate the basic structure of the video. For example, if the video draft is related to music, the protocol can at least indicate lyrics animations, background images, stickers, etc., and can also indicate information such as how to integrate music, lyrics, visual elements, and dynamic effects. Server 130 can package and generate the final video template based on the protocol of the video corresponding to this set of video drafts and related resources. These related resources can be video, images, effects, stickers, fonts, etc., which will not be elaborated here.

[0049] In some embodiments, the electronic device 110 may, in response to the video generation request, obtain a first set of resources corresponding to the subtitle text in a first video template. The subtitle text is text information associated with the target audio. Taking music content as an example, the corresponding subtitle text can be determined based on the lyrics corresponding to this music content. This subtitle text is the text that can be subsequently displayed in the generated video. The first set of resources can be used to indicate a first editing effect for the subtitle text. Specifically, this first editing effect can include editing effects of any appropriate dimension, such as the font size, color, appearance and disappearance animation effects of the subtitle text, etc. Taking music as an example, the editing effect of the lyrics can also be called lyric animation. Lyric animation is a type of lyric special effect that can display lyric content in different cool ways. Lyric animation can be composed of at least one of the following lyric animation elements: text animation (such as fade-in / fade-out, scaling, rotation, movement, etc.), color change, font style, etc. It should be noted that the lyrics animation effect corresponding to the target music is related to the lyrics animation effect corresponding to the lyrics in the first video template. The target lyrics (the lyrics corresponding to the target music) corresponding to this lyrics animation effect have synchronized dynamic effects and visual elements with the lyrics in the first video template. For example, if the first line of lyrics in the first video template corresponds to font A and color B, then the editing effect corresponding to the first line of lyrics in the target lyrics can also be displayed in the form of font A and color B.

[0050] In some embodiments, this first editing result can be sent from the target device (server 130) to the electronic device 110. Specifically, the electronic device can obtain a processing file for the subtitle text sent by the target device, which indicates the first editing effect for the subtitle text. For example, taking music as the target audio, this processing file can instruct the first line of lyrics to be displayed in font A, the second line of lyrics to be displayed in color B, and a sticker C to be displayed after the second line of lyrics finishes playing and before the third line of lyrics begins playing, etc. This processing file can be any suitable type of file, such as a JSON file, etc.

[0051] Furthermore, the electronic device 110 can generate a first text material segment corresponding to the subtitle text based on this first set of resources. In some embodiments, the subtitle text can be divided into at least one subtitle text segment, and each subtitle text segment appears at a different time in the target audio. For example, taking the target audio as target music, the subtitle text can be text that is subsequently displayed in the generated video and is associated with the lyrics in the target music. Specifically, each line of lyrics in the target music can be divided into a lyric segment. The first text material segment and the subtitle text segment are in one-to-one correspondence. The first text material segment can be used to present the effect of dynamically switching subtitle text segments according to the appearance time of the subtitle text segments in the target audio.

[0052] Furthermore, the electronic device 110 can generate a first video corresponding to the target audio based on the first text material fragment, the target audio, and the second set of resources in the first video template. The second set of resources may include at least one first image material. The at least one first image material may include video material and / or image material. The at least one first image material may be matched with the target audio and / or subtitle text. For example, the second set of resources may include background images, stickers, etc., in the first video template.

[0053] In some embodiments, the electronic device 110 can, in response to acquiring a first set of resources, create a sub-draft file corresponding to the subtitle text based on the first set of resources and the subtitle text to generate a first text material fragment. This sub-draft file may include editing effects for the subtitle text, such as text animation, color changes, synchronization effects, etc. Further, the electronic device 110 can, in response to acquiring a second set of resources, migrate the data from the sub-draft file to a main draft file. Further, the electronic device 110 can combine the first text material fragment and the second set of resources in the main draft file to generate a first video draft. The first text material fragment is recorded in the first video draft in the form of a draft, and this first video draft may include, in addition to the first text material fragment, any other appropriate sub-drafts, such as sub-drafts associated with various resources other than the subtitle text, etc. Further, the electronic device 110 can synthesize a first video based on the first video draft and target audio.

[0054] Specifically, the electronic device 110 can convert this first video draft into a composite clip, where the composite clip appears as a video segment to the user in the multitrack editor, but its internal structure is actually a complete video draft. Furthermore, the electronic device 110 can insert target audio into this first video draft to add it to the track, thereby synthesizing the first video.

[0055] To improve video generation efficiency, the process of generating the first text material fragment based on the first set of resources and the process of acquiring the second set of resources can be carried out in parallel.

[0056] Taking the target audio as the target music as an example, the following is about... Figure 2G The process of generating the music video corresponding to the target music is explained below:

[0057] In box 251, electronic device 110 can pull a list of video templates and obtain the first video template. For example, the first video template in the list can be selected as the first video template.

[0058] In box 252, electronic device 110 can prepare a sub-draft file, i.e., generate the first text material fragment. Specifically, electronic device 110 can create a video compositing link task based on a target interface. The target interface can be any suitable interface used to obtain this processed file from server 130. Electronic device 110 can download this processed file. Specifically, this processed file can be a compressed file. Electronic device 110 can decompress this processed file to obtain a decompressed processed file of type JSON. Electronic device 110 can parse this processed file and, based on this processed file and the lyrics in the target music, obtain the first set of resources corresponding to the subtitle text. Further, electronic device 110 can, based on this first set of resources and the subtitle text, create a sub-draft file corresponding to the subtitle text to generate the first text material fragment.

[0059] In box 253, electronic device 110 can obtain the second set of resources in the first video template, such as downloading background images, stickers, etc. from the first video template.

[0060] In box 254, electronic device 110 can migrate the sub-draft file to the main draft file and combine the first text material fragment and the second set of resources in the main draft file to generate the first video draft.

[0061] In box 255, electronic device 110 can convert this first video draft into a composite clip. This composite clip appears as a video segment to the user in the multitrack editor, but its internal structure is actually a complete video draft. Furthermore, electronic device 110 can add audio files; specifically, it can add target music to this composite clip to generate a target music video. In other words, electronic device 110 can insert this target music into the first video draft to add it to the track, thereby generating the target music video.

[0062] In box 256, electronic device 110 can determine that the target music video has been successfully synthesized.

[0063] It should be noted that server 130 may generate video drafts based on the latest draft version, while electronic device 110 may have requests for lower versions. In order to avoid problems such as higher version video drafts being unable to be opened on lower version electronic device 110, server 130 may perform video draft upgrade / downgrade processing before distributing the video draft.

[0064] by Figure 2DAs an example, the electronic device 110 can display the target video 230 (a video generated based on the target audio) at a predetermined location on the interface 200D. The electronic device 110 can also display information associated with this target video in the interface 200D, such as duration information, video cover information, etc. The electronic device 110 can also display playback controls for the target video in the interface 200D, allowing the user to select these controls to play the target video 230.

[0065] In some embodiments, the electronic device 110 may also present a template changing control in the interface 200D to allow users to input template changing requests, thereby enabling modifications to the target video 230. The electronic device 110 may present a set of candidate video templates in response to receiving a template changing request. Figure 2D As an example, electronic device 110 can display video template 220-1, video template 220-2, etc. in interface 200D.

[0066] The electronic device 110 can, in response to receiving a selection instruction for a second video template from a set of candidate video templates, generate a second video corresponding to the target audio based on the second video template and the target audio.

[0067] Specifically, the electronic device 110 can, in response to receiving a selection instruction for a second video template from a set of candidate video templates, obtain a third set of resources corresponding to the subtitle text in the second video template based on the second video template. The third set of resources is used to indicate a second editing effect for the subtitle text. Further, the electronic device 110 can generate a second text material fragment corresponding to the subtitle text based on the third set of resources. Further, the electronic device 110 can generate a second video corresponding to the target audio based on the second text material fragment, the target audio, and a fourth set of resources in the second video template. The fourth set of resources includes at least one second image material, which includes video material and / or image material, and the at least one second image material matches the target audio and / or subtitle text.

[0068] Specifically, the electronic device can generate a third video draft based on the second text material fragment and the fourth set of resources in the second video template. It should be noted that the process of generating a third video draft corresponding to the target audio based on the second video template is the same as the process of generating a first video draft corresponding to the target audio based on the first video template, and will not be elaborated here. Furthermore, the electronic device 110 can synthesize a second video corresponding to the target audio based on the third video draft and the target audio.

[0069] In some embodiments, the electronic device 110 may also present a media replacement control 231 in the interface 200D to allow users to input media replacement requests, thereby enabling changes to the target video 230. In response to receiving a media replacement request, the electronic device 110 may present a set of candidate media. This set of candidate media can be any suitable type of media, such as images, videos, etc. Figure 2D as well as Figure 2E As an example, electronic device 110 may, in response to receiving a user's selection of control 231, present as follows: Figure 2E The interface 200E is shown. The electronic device 110 can present a set of videos, such as video 240-1, video 240-2, and video 240-3, on the interface 200E. Furthermore, in response to receiving a selection of at least one material from this set of candidate materials, the electronic device 110 can generate a third video corresponding to the target audio based on this at least one material and the target audio.

[0070] Specifically, the electronic device 110 can use at least a portion of the selected at least one piece of material to replace the second set of resources corresponding to the first video, thereby obtaining a fifth set of resources after replacement. The duration of the at least one piece of material may exceed the duration of the target audio, or it may be less than the duration of the target audio.

[0071] Specifically, the electronic device 110 can, in response to a first duration of at least one piece of material exceeding a second duration of the target audio, trim a first target material corresponding to the second duration from the at least one piece of material. The first target material can be automatically trimmed by the electronic device 110, or it can be trimmed by the user based on an editing interface.

[0072] As an example, the electronic device 110 can present a media editing interface. Furthermore, the user can perform a trimming operation on at least one piece of media based on this media editing interface, at which point the electronic device can obtain the trimmed first target media. For example, if the total duration of the video selected by the user is 70 seconds, and the duration of the target music is 60 seconds, then the electronic device 110 can present a video trimming interface to support the user in trimming the selected video based on this interface, resulting in a 60-second video. Furthermore, the electronic device 110 can use the first target media to replace the second set of resources in the first video draft.

[0073] The electronic device 110 can also, in response to a first duration of at least one material being less than a second duration of the target audio, generate a second target material corresponding to the second duration based on the at least one material. As an example, this second target material may include multiple sets of sub-materials, each set of sub-materials including the at least one material. For instance, if the total duration of the video selected by the user is 3 seconds, and the duration of the target music is 60 seconds, the electronic device 110 can fill in the 3-second video so that the generated target video loops through the 3-second videos sequentially until playback ends after 60 seconds. Furthermore, the electronic device 110 can use the second target material to replace a second set of resources in the first video draft.

[0074] In some embodiments, the electronic device 110 may, in response to the fact that the type corresponding to the received at least one material is at least one image, set the playback rate of the at least one image in the target video until the playback ends when the last image of the at least one image has been played.

[0075] In some embodiments, the electronic device 110 can also support other editing operations by the user on the target video 230 (first video, second video, or third video), such as splitting segments of the target video 230, adjusting volume, speed, switching to picture-in-picture, placing video on the main track, using masks, and blending modes. Figure 2F As shown, the electronic device 110 can display the main track corresponding to the target video 230 on the interface 200F. The electronic device 110 can also display editing components for the target video 230 on the interface 200F, such as a material replacement component, a template replacement component, a splitting component, a volume control component, a deletion component, and a speed adjustment component. As an example, the electronic device 110 can split segments in the target video 230 in response to receiving a user's selection of the splitting component. As another example, the electronic device 110 can adjust the volume of the target video 230 in response to receiving a user's selection of the volume control component.

[0076] Therefore, the embodiments of this disclosure can generate a video corresponding to the target audio based on the first video template, which can effectively improve the video generation efficiency.

[0077] Example process

[0078] Figure 3 A flowchart of a video generation process 300 according to some embodiments of the present disclosure is shown. Process 300 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 300.

[0079] In box 310, electronic device 110 obtains a video generation request associated with the target audio.

[0080] In frame 320, electronic device 110 responds to a video generation request by acquiring a first set of resources corresponding to the subtitle text in a first video template; the subtitle text is text information associated with the target audio, and the first set of resources is used to indicate the first editing effect of the subtitle text.

[0081] In frame 330, electronic device 110 generates a first text material fragment corresponding to the subtitle text based on the first set of resources. The first text material fragment is used to present the effect of dynamically switching the subtitle text fragment according to the time when the subtitle text fragment appears in the target audio.

[0082] In box 340, electronic device 110 generates a first video corresponding to the target audio based on a second set of resources in the first text material fragment, the target audio, and the first video template. The second set of resources includes at least one first image material, which includes video material and / or image material, and at least one first image material matches the target audio and / or subtitle text.

[0083] In some embodiments, the first video template includes: a preset video template; a video template determined based on at least one attribute of the target audio; or a video template selected by the user.

[0084] In some embodiments, process 300 further includes: in response to receiving a template replacement request, presenting a set of candidate video templates; in response to receiving a selection instruction for a second video template in the set of candidate video templates, obtaining a third set of resources corresponding to the subtitle text in the second video template based on the second video template, the third set of resources being used to indicate a second editing effect of the subtitle text; and generating a second text material fragment corresponding to the subtitle text based on the third set of resources; and generating a second video corresponding to the target audio based on the second text material fragment, the target audio, and a fourth set of resources in the second video template, the fourth set of resources including at least one second image material, the at least one second image material including video material and / or image material, and the at least one second image material matching the target audio and / or subtitle text.

[0085] In some embodiments, process 300 further includes: in response to receiving a material replacement request, presenting a set of candidate materials; in response to receiving a selection instruction for at least one material in the set of candidate materials, replacing a second set of resources corresponding to the first video based on a portion of the at least one material to obtain a fifth set of resources after replacement; and generating a third video corresponding to the target audio based on the first text material fragment, the target audio, and the fifth set of resources.

[0086] In some embodiments, process 300 further includes: replacing a second set of resources corresponding to the first video based on a portion of at least one material, including: in response to a first duration of at least one material exceeding a second duration of a target audio, cropping a first target material corresponding to the second duration from the at least one material; and using the first target material to replace the second set of resources.

[0087] In some embodiments, replacing a second set of resources corresponding to a first video based on a portion of at least one material includes: in response to a first duration of at least one material being less than a second duration of a target audio, generating a second target material corresponding to the second duration based on at least one material; and using the second target material to replace the second set of resources.

[0088] In some embodiments, generating a first video corresponding to the target audio based on a first text material fragment, a target audio, and a second set of resources in a first video template includes: in response to obtaining the first set of resources, creating a sub-draft file corresponding to the subtitle text based on the first set of resources and subtitle text to generate a first text material fragment; and in response to obtaining the second set of resources, migrating the data of the sub-draft file to a main draft file; and combining the first text material fragment and the second set of resources in the main draft file to generate a first video draft, wherein the first text material fragment is recorded in the first video draft in the form of a sub-draft; and synthesizing the first video based on the first video draft and the target audio.

[0089] In some embodiments, the process of generating the first text material fragment based on the first set of resources and the process of acquiring the second set of resources are performed in parallel.

[0090] In some embodiments, the first editing effect is determined based on the following process: obtaining a processing file for the subtitle text sent by the target device, the processing file indicating the first editing effect for the subtitle text.

[0091] In some embodiments, the target audio includes music content, and the subtitle text is determined based on the lyrics corresponding to the music content.

[0092] Example devices and equipment

[0093] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example device 400 for generating video according to certain embodiments of the present disclosure is shown. Device 400 may be implemented as or included in electronic device 110. Various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0094] like Figure 4As shown, the device 400 includes a first acquisition module 410 configured to acquire a video generation request associated with a target audio; a second acquisition module 420 configured to acquire a first set of resources corresponding to subtitle text in a first video template in response to the video generation request; the subtitle text is text information associated with the target audio, and the first set of resources is used to indicate a first editing effect of the subtitle text; a first generation module 430 configured to generate a first text material fragment corresponding to the subtitle text based on the first set of resources, the first text material fragment being used to present the effect of dynamically switching subtitle text fragments according to the time of appearance of the subtitle text fragment in the target audio; and a second generation module 440 configured to generate a first video corresponding to the target audio based on the first text material fragment, the target audio, and a second set of resources in the first video template, the second set of resources including at least one first image material, the at least one first image material including video material and / or image material, and the at least one first image material matching the target audio and / or subtitle text.

[0095] In some embodiments, the first video template includes: a preset video template; a video template determined based on at least one attribute of the target audio; or a video template selected by the user.

[0096] In some embodiments, the apparatus 400 further includes a first presentation module configured to: present a set of candidate video templates in response to receiving a template replacement request; a third acquisition module configured to: in response to receiving a selection instruction for a second video template in the set of candidate video templates, acquire a third set of resources in the second video template corresponding to the subtitle text based on the second video template, the third set of resources being used to indicate a second editing effect of the subtitle text; a third generation module configured to: generate a second text material fragment corresponding to the subtitle text based on the third set of resources; and a fourth generation module configured to: generate a second video corresponding to the target audio based on the second text material fragment, the target audio, and the fourth set of resources in the second video template, the fourth set of resources including at least one second image material, the at least one second image material including video material and / or image material, and the at least one second image material matching the target audio and / or subtitle text.

[0097] In some embodiments, the apparatus 400 further includes a second presentation module configured to: present a set of candidate materials in response to receiving a material replacement request; a replacement module configured to: replace a second set of resources corresponding to the first video based on a portion of the at least one material in response to receiving a selection instruction for at least one material in the set of candidate materials, so as to obtain a fifth set of resources after replacement; and a fifth generation module configured to: generate a third video corresponding to the target audio based on the first text material fragment, the target audio and the fifth set of resources.

[0098] In some embodiments, the replacement module is further configured to: in response to a first duration of at least one material exceeding a second duration of the target audio, cut out a first target material corresponding to the second duration from the at least one material; and use the first target material to replace the second set of resources.

[0099] In some embodiments, the replacement module is further configured to: in response to a first duration of at least one material being less than a second duration of the target audio, generate a second target material corresponding to the second duration based on at least one material; and use the second target material to replace the second set of resources.

[0100] In some embodiments, the second generation module 440 is further configured to: in response to acquiring the first set of resources, create a sub-draft file corresponding to the subtitle text based on the first set of resources and the subtitle text to generate a first text material fragment; and in response to acquiring the second set of resources, migrate the data of the sub-draft file to the main draft file; and combine the first text material fragment and the second set of resources in the main draft file to generate a first video draft, wherein the first text material fragment is recorded in the first video draft in the form of a sub-draft; and synthesize a first video based on the first video draft and the target audio.

[0101] In some embodiments, the process of generating the first text material fragment based on the first set of resources and the process of acquiring the second set of resources are performed in parallel.

[0102] In some embodiments, the first editing effect is determined based on the following process: obtaining a processing file for the subtitle text sent by the target device, the processing file indicating the first editing effect for the subtitle text.

[0103] In some embodiments, the target audio includes music content, and the subtitle text is determined based on the lyrics corresponding to the music content.

[0104] The units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0105] Figure 5A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 Electronic devices 110.

[0106] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0107] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0108] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0109] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0110] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0111] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0112] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0113] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0114] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0116] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating video, comprising: Obtain the video generation request associated with the target audio; In response to the video generation request, the first set of resources corresponding to the subtitle text in the first video template is obtained; The subtitle text is text information associated with the target audio, and the first set of resources is used to indicate the first editing effect of the subtitle text; Based on the first set of resources, a first text material fragment corresponding to the subtitle text is generated. The first text material fragment is used to present the effect of dynamically switching the subtitle text fragment according to the time when the subtitle text fragment appears in the target audio. as well as Based on the first text material fragment, the target audio, and the second set of resources in the first video template, a first video corresponding to the target audio is generated. The second set of resources includes at least one first image material, which includes video material and / or image material, and the at least one first image material matches the target audio and / or the subtitle text.

2. The method according to claim 1, wherein the first video template comprises: Preset video templates; A video template determined based on at least one attribute of the target audio; or The video template selected by the user.

3. The method according to claim 1, further comprising: Upon receiving a template change request, a set of candidate video templates is presented. In response to receiving a selection instruction for a second video template from the set of candidate video templates, based on the second video template, a third set of resources corresponding to the subtitle text in the second video template is obtained, the third set of resources being used to indicate a second editing effect for the subtitle text; Based on the third set of resources, a second text material fragment corresponding to the subtitle text is generated; as well as Based on the second text material fragment, the target audio, and the fourth set of resources in the second video template, a second video corresponding to the target audio is generated. The fourth set of resources includes at least one second image material, which includes video material and / or image material. The at least one second image material matches the target audio and / or the subtitle text.

4. The method according to claim 1, further comprising: In response to a received material replacement request, a set of candidate materials is presented; In response to receiving a selection instruction for at least one material in the set of candidate materials, the second set of resources corresponding to the first video is replaced based on a portion of the at least one material to obtain a fifth set of resources after replacement; as well as A third video corresponding to the target audio is generated based on the first text material fragment, the target audio, and the fifth group of resources.

5. The method of claim 4, wherein replacing the second group of resources corresponding to the first video based on a portion of the at least one material comprises: In response to the fact that the first duration of the at least one material exceeds the second duration of the target audio, a first target material corresponding to the second duration is cropped from the at least one material; as well as Replace the second set of resources with the first target material.

6. The method of claim 4, wherein replacing the second group of resources corresponding to the first video based on a portion of the at least one material comprises: In response to the fact that the first duration of the at least one material is less than the second duration of the target audio, a second target material corresponding to the second duration is generated based on the at least one material; as well as Replace the second set of resources with the second target material.

7. The method according to claim 1, wherein generating a first video corresponding to the target audio based on the first text material fragment, the target audio, and the second set of resources in the first video template comprises: In response to obtaining the first set of resources, a sub-draft file corresponding to the subtitle text is created based on the first set of resources and the subtitle text to generate the first text material fragment; In response to acquiring the second set of resources, the data of the sub-draft file is migrated to the main draft file; The first text material fragment and the second set of resources are combined in the main draft file to generate a first video draft, wherein the first text material fragment is recorded in the first video draft in the form of a sub-draft; as well as The first video is synthesized based on the first video draft and the target audio.

8. The method according to claim 1, wherein the process of generating the first text material fragment based on the first set of resources and the process of acquiring the second set of resources are performed in parallel.

9. The method of claim 1, wherein the first editing effect is determined based on the following process: Obtain a processing file for the subtitle text sent by the target device, the processing file indicating a first editing effect for the subtitle text.

10. The method of claim 1, wherein the target audio includes music content, and the subtitle text is determined based on the lyrics content corresponding to the music content.

11. An apparatus for generating video, comprising: The first acquisition module is configured to acquire a video generation request associated with the target audio. The second acquisition module is configured to, in response to the video generation request, acquire a first set of resources in the first video template corresponding to the subtitle text; the subtitle text is text information associated with the target audio, and the first set of resources is used to indicate a first editing effect of the subtitle text; The first generation module is configured to generate a first text material fragment corresponding to the subtitle text based on the first set of resources. The first text material fragment is used to present the effect of dynamically switching the subtitle text fragment according to the time when the subtitle text fragment appears in the target audio. as well as The second generation module is configured to generate a first video corresponding to the target audio based on a second set of resources in the first text material fragment, the target audio, and the first video template. The second set of resources includes at least one first image material, which includes video material and / or image material, and the at least one first image material matches the target audio and / or the subtitle text.

12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.