Method and apparatus for generating music video, device, and storage medium

By generating music and visual content for music videos by acquiring reference information, the problem of users being unable to configure videos and music independently is solved, and personalized music videos are generated efficiently.

WO2026157445A1PCT designated stage Publication Date: 2026-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-11-12
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing technology cannot meet users' needs to independently configure and generate video and music content during the video creation process. Users need to select preset templates to create videos, which cannot satisfy personalized creation.

Method used

By acquiring reference images, reference text, and reference audio information, the music and visual content of music videos are generated, including visual elements based on reference images and lyric visual elements based on music content, enabling users to customize and generate music videos.

Benefits of technology

It enables the generation of music videos based on user configurations, meeting users' needs for creating video and music content and improving generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025134364_30072026_PF_FP_ABST
    Figure CN2025134364_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and apparatus for generating a music video, a device, and a storage medium. The method proposed herein comprises: acquiring reference information, wherein the reference information comprises a reference image, a reference text, and reference audio information; and providing a music video generated on the basis of the reference information, wherein music content of the music video is generated on the basis of the reference text and the reference audio information, and visual content of the music video comprises a first set of visual elements generated on the basis of a reference object in the reference image and a second set of visual elements generated on the basis of lyrics of the music content. In this way, the embodiments of the present disclosure can improve the efficiency of generating music videos.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and storage media for generating music videos

[0001] This application claims priority to Chinese Patent Application No. 202510128701.1, filed on January 27, 2025, entitled "Method, Apparatus, Device and Storage Medium for Generating Music Videos", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating music videos. Background Technology

[0003] With the increasing maturity of internet technology, more and more users are utilizing online platforms to assist in generating video content. For example, online platforms can generate video content based on user-provided images. However, when users generate video content using existing online platforms, they need to select preset video templates for video production and cannot generate video content based on their own configurations. Furthermore, current technology cannot meet users' needs to create new music content during the video creation process. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for generating a music video is provided. The method includes: acquiring reference information, including a reference image, reference text, and reference audio information; and providing a music video generated based on the reference information, wherein the music content of the music video is generated based on the reference text and reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in the reference image and a second set of visual elements generated based on the lyrics of the music content.

[0005] In a second aspect of this disclosure, an apparatus for generating a music video is provided. The apparatus includes: an acquisition module configured to acquire reference information, the reference information including a reference image, reference text, and reference audio information; and a providing module configured to provide a music video generated based on the reference information, wherein the music content of the music video is generated based on the reference text and reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in the reference image and a second set of visual elements generated based on the lyrics of the music content.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figure 2 shows a flowchart of an example process for generating a music video according to some embodiments of the present disclosure;

[0013] Figures 3A to 3D illustrate example interfaces according to some embodiments of the present disclosure;

[0014] Figure 4 illustrates a flowchart of an example process for generating a music video according to some embodiments of the present disclosure;

[0015] Figure 5 shows a schematic structural block diagram of an example apparatus for generating music videos according to some embodiments of the present disclosure; and

[0016] Figure 6 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0022] As mentioned above, with the increasing maturity of internet technology, more and more users are using online platforms to assist in generating video content. For example, online platforms can generate video content based on user-provided images. However, when users generate video content using existing online platforms, they need to select preset video templates for video production and cannot generate video content based on their own configurations. Furthermore, current technology cannot meet users' needs to create new music content during the video creation process.

[0023] Embodiments of this disclosure propose a scheme for generating music videos. According to this scheme, reference information, including reference images, reference text, and reference audio information, can be obtained; and a music video generated based on the reference information can be provided, wherein the music content of the music video is generated based on the reference text and reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in the reference images and a second set of visual elements generated based on the lyrics of the music content.

[0024] In this way, embodiments of this disclosure can obtain reference information for generating music videos. Embodiments of this disclosure can generate music content for music videos based on reference text and reference audio information in the reference information. Furthermore, embodiments of this disclosure can generate visual content for music videos based on visual elements generated from reference objects in reference images and visual elements corresponding to the lyrics of the music content. Thus, embodiments of this disclosure can generate music videos based on user configurations, thereby satisfying users' needs for creating video content and corresponding music content, and improving the efficiency of users generating music videos.

[0025] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0026] Example Environment

[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.

[0028] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction, examples of which may include, but are not limited to, video editing applications, media applications, social applications, or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.

[0029] In environment 100 of Figure 1, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.

[0030] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0031] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support virtual scenarios in electronic devices 110.

[0032] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.

[0033] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0034] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0035] Example of generating music video

[0036] Figure 2 shows a flowchart of an example process 200 for generating a music video according to some embodiments of the present disclosure. Process 200 can be implemented at electronic device 110 and / or server 130. Process 200 will now be described with reference to Figure 1.

[0037] In box 205, electronic device 110 and / or server 130 can acquire user-provided images (e.g., photos and / or videos).

[0038] The following description, based on Figures 3A to 3D, illustrates the interactive process for generating music videos. It should be understood that Figures 3A to 3D and the corresponding embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure.

[0039] Figures 3A to 3D illustrate example interfaces 300A to 300D according to some embodiments of the present disclosure. Interfaces 300A to 300D may be provided, for example, by the electronic device 110 shown in Figure 1.

[0040] In some embodiments, as shown in FIG3A, the electronic device 110 may present an interactive interface 300A. As an example, the interactive interface 300A may include an interactive interface for a media editing application and / or an interactive interface for a video application, etc.

[0041] As an example, electronic device 110 can obtain reference information via interactive interface 300A. The reference information can be used to generate a music video. As an example, the music video may include music content (or audio content) and visual content (or video content). The reference information includes reference images, reference text, and reference audio information.

[0042] In some embodiments, the electronic device 110 may acquire a reference image via an interactive interface 300A. As an example, the electronic device 110 may provide an image entry 305 in the interactive interface 300A for acquiring the reference image. Further, the electronic device 110 may present an image selection interface in response to triggering the image entry 305.

[0043] In some embodiments, as shown in FIG3B, the electronic device 110 may present an image selection interface 300B. The electronic device 110 may present a set of candidate images in the image selection interface 300B. As an example, the set of candidate images may be image content from a media library associated with the user. For example, the media library associated with the user may include a local media library (e.g., a local photo album) and / or an online media library (e.g., a cloud-based photo album stored on a cloud server). As an example, the set of candidate images may include candidate image 306-1, candidate image 306-2, and candidate image 306-1, etc.

[0044] As an example, electronic device 110 may use the target image as a reference image in response to the selection of a target image (e.g., candidate image 306-2) from a set of candidate images.

[0045] Alternatively, the electronic device 110 may provide a shooting control 307 in the interactive interface 300B. As an example, the shooting control 307 may invoke a camera component associated with the electronic device 110 to acquire image content. Furthermore, the electronic device 110 may acquire image content captured by the user via the shooting control 307 as a reference image.

[0046] As an example, electronic device 110 and / or server 130 may, in response to a reference image not meeting a condition associated with a reference object, present an alert message indicating that condition. As an example, the reference object may indicate a human face and / or an animal face.

[0047] As an example, the conditions associated with the reference object can indicate whether the reference object is detected in the reference image. As an example, the electronic device 110 can present a corresponding reminder message (e.g., "Please replace with a clear front-facing photo") in response to the absence of a reference object (e.g., a human face and / or an animal face) in the reference image.

[0048] As an example, the conditions associated with the reference object can indicate the occlusion ratio of the reference object in the reference image. For example, the occlusion ratio of a human face or an animal face in the reference image. As an example, the electronic device 110 can respond to the condition that the occlusion ratio of the reference object in the reference image does not meet the condition (e.g., the occlusion ratio of a human face and / or an animal face is greater than a preset threshold (e.g., 1 / 5)) by presenting a corresponding reminder message (e.g., "Too much face occlusion, please replace with a clear photo").

[0049] Alternatively, the conditions associated with the reference object can indicate the display size of the reference object in the reference image. For example, the display size of a human face or an animal face in the reference image. As an example, the electronic device 110 may display a corresponding reminder message (e.g., "Face size is too small, please replace with a clear frontal photo") in response to the display size of the reference object in the reference image not meeting the conditions (e.g., the display size of a human face and / or an animal face is smaller than a preset size).

[0050] Alternatively, the conditions associated with the reference object can indicate the display angle of the reference object in the reference image. For example, the display angle of a human face or an animal face in the reference image. As an example, the electronic device 110 may display a corresponding reminder message (e.g., "Not a frontal image, please replace it with a clear frontal photo") in response to the display angle of the reference object in the reference image not meeting the conditions (e.g., a human face and / or an animal face is in profile).

[0051] This disclosure, based on the detection of reference objects in a reference image and the determination of the occlusion ratio, display size, and / or display angle of the reference objects in the reference image, can guarantee the display quality of the screen content associated with the reference objects in the generated music video.

[0052] In some embodiments, as shown in FIG3C, the electronic device 110 may present a reference image in the image entry 305 of the interactive interface 300C. Alternatively, the electronic device 110 may present a return control 308 in the interactive interface 300C. Further, the electronic device 110 may switch to the image selection interface to reacquire the reference image in response to a triggering of the return control 308 (e.g., a click operation).

[0053] As an example, continuing to refer to Figure 2, in box 260, electronic device 110 can obtain reference text (e.g., username and / or greetings, etc.).

[0054] As an example, continuing to refer to Figure 3C, the electronic device 110 can display preset candidate text 310 in the interactive interface 300C. As an example, the preset candidate text 310 may include a first candidate text 311 (e.g., username), a second candidate text 312 (e.g., recipient), and a third candidate text 313 (e.g., a blessing). The electronic device 110 can receive editing operations on the candidate text via the interactive interface to determine the reference text.

[0055] Additionally or alternatively, the preset text 310 may include a fixed text portion and an editable text portion. As an example, the fixed text portion cannot be modified. The editable text portion allows user modification. Take the first candidate text 311 as an example. The first candidate text 311 may include a fixed text portion 311-1 (e.g., "I am") and an editable text portion 311-2 (e.g., "XX"). As an example, the fixed text portion may be a preset fixed field. The preset content of the editable text portion may be randomly determined from a preset text library or determined based on user reference information. For example, the editable text portion 311-2 may be determined based on a user nickname provided by the user.

[0056] As an example, the electronic device 110 may also present instruction text 315 (e.g., “Tip: Click the dotted text to enter your own text”) in the interactive interface 300C to indicate the editable text portion (e.g., the text portion with a dotted bottom) in the candidate text 310.

[0057] As an example, continuing to refer to Figure 2, in box 265, electronic device 110 can determine whether to customize the timbre based on the user's selection to obtain reference audio information (e.g., reference timbre).

[0058] In some embodiments, continuing to refer to FIG3C, the electronic device 110 can obtain reference audio information in the interactive interface 300C. As an example, the electronic device 110 can provide a set of tone controls in the interactive interface 300C. As an example, the set of tone controls may include a first tone control 320-1, a second tone control 320-2, and a third tone control 320-3.

[0059] In some embodiments, continuing to refer to FIG2, in block 270, electronic device 110 and / or server 130 can perform tone cloning.

[0060] As an example, the electronic device 110 can obtain the audio content uploaded by the user based on the first tone control 320-1 (also known as the recording entry). As an example, the electronic device 110 can present a recording interface in response to the selection of the first tone control 320-1. As an example, the electronic device 110 can present preset lyrics in the recording interface.

[0061] Additionally, the electronic device 110 can acquire the recorded singing content via a recording interface to determine reference audio information. As an example, the electronic device 110 and / or server 130 can determine reference audio information associated with the user's timbre information based on the singing content recorded by the user on the recording interface. As an example, the singing content may correspond to preset lyrics presented in the recording interface.

[0062] As an example, the recorded singing content needs to meet a preset duration. The electronic device 110 can respond to situations where the singing content does not meet the preset duration (e.g., 15 seconds) by displaying a reminder message (e.g., "Recording duration must not be less than 15 seconds") on the recording interface to guide the user to record singing content that meets the preset duration. Thus, embodiments of this disclosure, based on the limitation of the recording duration of the singing content, can ensure that the acquired reference audio information better reflects the user's vocal characteristics.

[0063] In some embodiments, continuing to refer to FIG2, in block 270, electronic device 110 and / or server 130 may determine the timbre selected by the user (e.g., male or female voice).

[0064] As an example, continuing to refer to Figure 3C, the second timbre control 320-2 can correspond to a male voice timbre. As an example, the electronic device 110 can respond to the selection of the second timbre control 320-2 by using the timbre information corresponding to the second timbre control 320-2 (e.g., a preset male voice audio segment) as a reference timbre.

[0065] As an example, the third timbre control 320-3 may correspond to a female voice timbre. As an example, the electronic device 110 may, in response to the selection of the third timbre control 320-3, use the timbre information corresponding to the third timbre control 320-3 (e.g., a preset female voice audio segment) as a reference timbre.

[0066] As an example, electronic device 110 may provide a music video generated based on the reference information obtained from the interactive interface 300C in response to confirmation of the reference information. As an example, electronic device 110 may present a confirmation control 325 in the interactive interface 300C. Furthermore, electronic device 110 may respond to the selection of the confirmation control 325 (e.g., a click operation) as confirmation of the reference information obtained from the interactive interface 300C.

[0067] As an example, electronic device 110 and / or server 130 can generate music videos based on reference information.

[0068] In some embodiments, electronic device 110 and / or server 130 may determine a sub-image corresponding to a reference object based on a reference image.

[0069] In some embodiments, continuing to refer to FIG2, in block 210, electronic device 110 and / or server 130 can identify a subject (also referred to as a reference object) in a reference image. The reference object may, for example, include a human face and / or an animal face. Further, electronic device 110 and / or server 130 can determine a sub-image corresponding to the reference object based on the reference image. For example, the sub-image corresponding to the reference object is cropped from the reference image (also referred to as image matting).

[0070] As an example, electronic device 110 and / or server 130 can provide sub-images to the mapping model to generate intermediate images.

[0071] As an example, electronic device 110 and / or server 130 can determine the object type of the reference object. As an example, the object type of the reference object may include, for example, a person object type (e.g., adult, child, man, woman, etc.) and an animal object type (e.g., cat, dog, etc.).

[0072] For example, in box 215, electronic device 110 and / or server 130 may provide the image expansion model with an image expansion cue word 1 (also referred to as the first cue information) to generate an intermediate image (also referred to as the first intermediate image) based on the sub-image. As an example, the first intermediate image may include a portion of the image corresponding to the sub-image. For example, the first intermediate image may be a full-body image of a person or an animal. As an example, the image expansion model may be implemented as a generative model capable of image expansion based on text. This disclosure is not intended to limit the specific implementation and training process of such generative model.

[0073] As an example, the first cue message can correspond to the object type of the reference object. For example, different object types can correspond to different cue messages. For example, the cue message for an animal object type can be different from that for a human object type. As an example, the cue message associated with a human object type can indicate clothing features, body proportions, height, etc. As an example, the cue message associated with an animal object type can indicate fur color, tail length, etc. As an example, multiple cue messages corresponding to multiple different object types can include preset cue messages and / or cue messages generated based on a text generation model. As an example, the text generation model can be implemented as a generative model capable of generating corresponding cue messages based on object type (e.g., animal type, human type). This disclosure is not intended to limit the specific implementation and / or training process of such generative model.

[0074] Alternatively, in box 235, electronic device 110 and / or server 130 may provide expansion prompt 2 to the expansion model to generate an intermediate image (also referred to as a second intermediate image) based on the sub-image. As an example, the second intermediate image may include a portion of the image corresponding to the sub-image. Expansion prompt 2 may differ from expansion prompt 1, thus the second intermediate image may differ from the first intermediate image. For example, the clothing of a person in the second intermediate image may differ from that in the first intermediate image.

[0075] In some embodiments, electronic device 110 and / or server 130 may utilize intermediate images to generate intermediate video content.

[0076] As an example, electronic device 110 and / or server 130 may provide intermediate images and preset second prompt information to the video model to generate intermediate video content.

[0077] In box 220, electronic device 110 and / or server 130 may provide a first intermediate image and action cue word 1 (also known as second cue information) to the video model to generate intermediate video content (e.g., the first intermediate video).

[0078] As an example, the second cue message may indicate a set of actions associated with a reference object. For example, the second cue message may include: "dance while bowing," "dance while clapping," "dance," "dance while going crazy," etc. As an example, the video model may be implemented as a generative model capable of generating videos based on images and text. This disclosure is not intended to limit the specific implementation and training process of such a generative model.

[0079] Alternatively, in box 240, electronic device 110 and / or server 130 can provide a second intermediate image and action cue word 2 to the video model to generate a second intermediate video. Action cue word 2 may differ from action cue word 1; therefore, the second intermediate video may differ from the first intermediate video. For example, the action associated with the reference object in the second intermediate video may differ from that in the first intermediate video.

[0080] In some embodiments, the electronic device 110 and / or server 130 can determine a first set of visual elements corresponding to a reference object from intermediate video content. As an example, the first set of visual elements may correspond to a set of image frames in the intermediate video content that are associated with the reference object. For instance, the first image frame of the intermediate video content includes a person image and a background image corresponding to the reference object. The first set of visual elements may include the image portion of the person image corresponding to the reference object, but not the background image.

[0081] As an example, in box 225, electronic device 110 and / or server 130 can extract motion frames associated with a reference object from the first intermediate video. For example, the first image frame in the first intermediate video may include a person and background image corresponding to the reference object. Electronic device 110 and / or server 130 can extract the image portion corresponding to the person in the first image frame from the first image frame. Further, electronic device 110 and / or server 130 can extract a set of visual elements (e.g., a first set of visual elements) corresponding to the reference object from the first intermediate video. Further, in box 230, electronic device 110 and / or server 130 can use a set of visual elements as motion material 1.

[0082] Alternatively, in frame 245, electronic device 110 and / or server 130 can extract a set of visual elements corresponding to the reference object from the second intermediate video. Further, in frame 255, electronic device 110 and / or server 130 can use this set of visual elements as motion data 2.

[0083] In box 280, electronic device 110 and / or server 130 can provide reference audio information and musical cue words (e.g., reference text) to the music model to generate musical material. The musical material may include, for example, audio content and lyrics. As an example, the music model can be implemented as a generative model capable of generating musical content based on reference audio information and cue information. This disclosure is not intended to train a specific implementation or training process of such a generative model.

[0084] In box 285, electronic device 110 and / or server 130 can generate a music video based on motion clip 1, motion clip 2, and / or music clips. As an example, electronic device 110 and / or server 130 can provide motion clip 1, motion clip 2, and / or music clips to a video model to generate a music video.

[0085] Alternatively, electronic device 110 and / or server 130 may determine a set of motion effects that match the rhythm information based on the rhythm information of the music content (e.g., music material).

[0086] As an example, a set of motion effects may include a first set of motion effects associated with a reference object. The first set of motion effects may correspond to a first set of visual elements. As an example, the actions presented by the first set of visual elements may match rhythmic information (e.g., the actions may follow the rhythm of music).

[0087] As an example, one set of animations may include a second set of animations corresponding to the lyrics. The second set of animations may correspond to a second set of visual elements (e.g., lyric text elements) generated based on the lyrics of the music content. As an example, the presentation of the lyric text elements presented by the second set of visual elements may be matched with rhythm information (e.g., the text content of the lyric text elements may correspond to the music content, and the presentation of the lyric text elements may follow the rhythm of the music).

[0088] Alternatively, electronic device 110 and / or server 130 may also provide a third set of visual elements associated with the video model and a preset video template to generate a music video. As an example, the third set of visual elements may include at least one image element, such as a sticker, background image, etc.

[0089] For example, as shown in FIG3D, the electronic device 110 can present a playback interface 300D. The electronic device 110 can play the generated music video 330 in the playback interface 300D.

[0090] As an example, music video 330 may include music content (e.g., music clips) generated based on reference text and reference audio information. The visual content of music video 330 may include a first set of visual elements (e.g., motion clip 1 and / or motion clip 2) generated based on reference objects in reference images. For example, electronic device 110 may display visual elements 335 corresponding to the reference objects in playback interface 300D.

[0091] Additionally, the music video 330 may also include a second set of visual elements (e.g., lyric text element 340) generated based on the lyrics of the music content. As an example, the presentation of the lyric text element 340 may match the rhythm information of the music video. For example, if the rhythm information of the music video is upbeat drumbeats, the presentation of the lyric text element 340 may include jumping to the rhythm of the drumbeats.

[0092] Alternatively, the music video may also include a third set of visual elements associated with a preset video template. For example, the third set of visual elements may include element 345 (e.g., "butterflies in flight"), etc.

[0093] As an example, electronic device 110 may display the cover image of a music video in response to receiving a sharing request. Furthermore, electronic device 110 may adjust the cover image in response to receiving an editing operation on the cover image. For example, electronic device 110 may adjust the image content in the cover image to the image currently selected by the user in response to an update of the image content in the cover image. For example, electronic device 110 may adjust the text content in the cover image in response to editing of the text content in the cover image.

[0094] In some embodiments, the electronic device 110 may present a sharing window in response to receiving a sharing request for a music video. The electronic device 110 may present multiple applications in the sharing window. Further, the electronic device 110 may publish the music video in one of the multiple applications in response to receiving a trigger on one of the applications.

[0095] Additionally, the electronic device 110 may present a save local option in the sharing window. Furthermore, in response to the triggering of the save local option, the electronic device 110 may save the music video to its local storage unit.

[0096] Alternatively, electronic device 110 may, in response to receiving a request to share a music video, create a web playback page associated with the music video. As an example, the web playback page may include a web playback page (e.g., an H5 playback page). The web playback page can be used to play the music video.

[0097] As an example, electronic device 110 can share a web playback page associated with a music video with at least one user based on the user's selection.

[0098] Alternatively, electronic devices associated with other users can display a web playback page corresponding to the music video. The web playback page may include interactive controls. These interactive controls can be used to receive user interaction content. As an example, the interactive content can be text content (e.g., bullet comments). The interactive content can be configured to be overlaid on the music video.

[0099] As an example, when the current user is playing music or video on the web playback page, they can also see interactive content sent by other users. Additionally, the current user can also delete interactive content on the web playback page. Alternatively, the current user can also set the display status of interactive content on the web playback page (e.g., show or hide).

[0100] As an example, the web playback page can also include a creation portal. Furthermore, in response to other users triggering the creation portal, a generation interface for creating music videos is presented on those users' electronic devices to guide them through the creation process.

[0101] In this way, embodiments of this disclosure can obtain reference information for generating music videos. Embodiments of this disclosure can generate music content for music videos based on reference text and reference audio information in the reference information. Furthermore, embodiments of this disclosure can generate visual content for music videos based on visual elements generated from reference objects in reference images and visual elements corresponding to the lyrics of the music content. Thus, embodiments of this disclosure can generate music videos based on user configurations, thereby satisfying users' needs for creating video content and corresponding music content, and improving the efficiency of users generating music videos.

[0102] Example process

[0103] Figure 4 shows a flowchart of an example process 400 for generating a music video according to some embodiments of the present disclosure. Process 400 can be implemented at an electronic device 110. Process 400 will now be described with reference to Figure 1.

[0104] As shown in Figure 4, in box 410, electronic device 110 acquires reference information, which includes reference images, reference text, and reference audio information.

[0105] In box 420, electronic device 110 provides a music video generated based on reference information, wherein the music content of the music video is generated based on reference text and reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in a reference image and a second set of visual elements generated based on the lyrics of the music content.

[0106] In some embodiments, the first set of visual elements is generated based on the following process: determining a sub-image corresponding to a reference object based on a reference image; providing the sub-image to an expansion model to generate an intermediate image; generating intermediate video content using the intermediate image; and determining the first set of visual elements corresponding to the reference object from the intermediate video content.

[0107] In some embodiments, providing a sub-image to the mapping model to generate an intermediate image includes: determining the object type of a reference object; and providing the mapping model with first cue information corresponding to the sub-image and the object type to generate the intermediate image.

[0108] In some embodiments, generating intermediate video content using an intermediate image includes: providing an intermediate image and a preset second cue message to a video model to generate intermediate video content, wherein the second cue message indicates a set of actions associated with a reference object.

[0109] In some embodiments, the music video is generated based on the following process: determining a set of motion effects that match the rhythm information based on the rhythm information of the music content, the set of motion effects including a first set of motion effects associated with a reference object and / or a second set of motion effects corresponding to the lyrics, the first set of motion effects corresponding to a first set of visual elements and the second set of motion effects corresponding to a second set of visual elements; and generating the music video based on the music content and the set of motion effects.

[0110] In some embodiments, obtaining reference information via an interactive interface includes: presenting a recording entry point in the interactive interface; presenting a recording interface in response to the selection of the recording entry point; and obtaining the recorded singing content via the recording interface to determine reference audio information.

[0111] In some embodiments, the singing content corresponds to the preset lyrics presented in the recording interface.

[0112] In some embodiments, obtaining reference information includes: presenting preset candidate text; and determining reference text based on editing operations on the candidate text.

[0113] In some embodiments, process 400 further includes: acquiring a reference image, including an uploaded image or a captured image; and, in response to the reference image not meeting the conditions associated with a reference object, presenting a reminder message indicating the conditions.

[0114] In some embodiments, the condition indication is related to at least one of the following: whether a reference object is detected in the reference image; the occlusion ratio of the reference object in the reference image; the display size of the reference object in the reference image; and the display angle of the reference object in the reference image.

[0115] In some embodiments, the music video includes a third set of visual elements associated with a preset video template.

[0116] Example devices and equipment

[0117] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows a schematic structural block diagram of an example apparatus 500 for generating music videos according to certain embodiments of this disclosure. Apparatus 500 may be implemented as or included in electronic device 110. The various modules / components in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0118] As shown in Figure 5, the device 500 includes: an acquisition module 510 configured to acquire reference information, the reference information including reference images, reference text, and reference audio information; and a providing module 520 configured to provide a music video generated based on the reference information, wherein the music content of the music video is generated based on the reference text and reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in the reference images and a second set of visual elements generated based on the lyrics of the music content.

[0119] In some embodiments, the first set of visual elements is generated based on the following process: determining a sub-image corresponding to a reference object based on a reference image; providing the sub-image to an expansion model to generate an intermediate image; generating intermediate video content using the intermediate image; and determining the first set of visual elements corresponding to the reference object from the intermediate video content.

[0120] In some embodiments, providing a sub-image to the mapping model to generate an intermediate image includes: determining the object type of a reference object; and providing the mapping model with first cue information corresponding to the sub-image and the object type to generate the intermediate image.

[0121] In some embodiments, generating intermediate video content using an intermediate image includes: providing an intermediate image and a preset second cue message to a video model to generate intermediate video content, wherein the second cue message indicates a set of actions associated with a reference object.

[0122] In some embodiments, the music video is generated based on the following process: determining a set of motion effects that match the rhythm information based on the rhythm information of the music content, the set of motion effects including a first set of motion effects associated with a reference object and / or a second set of motion effects corresponding to the lyrics, the first set of motion effects corresponding to a first set of visual elements and the second set of motion effects corresponding to a second set of visual elements; and generating the music video based on the music content and the set of motion effects.

[0123] In some embodiments, the acquisition module 510 is further configured to: present a recording entry point in an interactive interface; present a recording interface in response to the selection of the recording entry point; and acquire the recorded singing content via the recording interface to determine reference audio information.

[0124] In some embodiments, the singing content corresponds to the preset lyrics presented in the recording interface.

[0125] In some embodiments, the acquisition module 510 is further configured to: present preset candidate text; and determine reference text based on editing operations on the candidate text.

[0126] In some embodiments, the device 500 further includes an alert module configured to: acquire a reference image, including an uploaded image or a captured image; and, in response to the reference image not meeting the conditions associated with a reference object, present an alert message indicating the conditions.

[0127] In some embodiments, the condition indication is related to at least one of the following: whether a reference object is detected in the reference image; the occlusion ratio of the reference object in the reference image; the display size of the reference object in the reference image; and the display angle of the reference object in the reference image.

[0128] In some embodiments, the music video includes a third set of visual elements associated with a preset video template.

[0129] The units included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 500 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0130] Figure 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in Figure 6 can be used to implement the electronic device 110 of Figure 1.

[0131] As shown in Figure 6, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processor 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.

[0132] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.

[0133] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0134] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0135] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0136] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0137] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0138] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0139] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0141] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating a music video, comprising: Obtain reference information, which includes reference images, reference text, and reference audio information; as well as Provide a music video generated based on the reference information, wherein the music content of the music video is generated based on the reference text and the reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in the reference image and a second set of visual elements generated based on the lyrics of the music content.

2. The method of claim 1, wherein the first set of visual elements is generated based on the following process: Based on the reference image, determine the sub-image corresponding to the reference object; The sub-images are provided to the image expansion model to generate intermediate images; Using the intermediate images, intermediate video content is generated; as well as The first set of visual elements corresponding to the reference object are determined from the intermediate video content.

3. The method of claim 2, wherein providing the sub-image to the mapping model to generate the intermediate image comprises: Determine the object type of the reference object; as well as The image expansion model is provided with first prompt information corresponding to the sub-image and the object type to generate the intermediate image.

4. The method according to claim 2 or 3, wherein generating intermediate video content using the intermediate image comprises: The intermediate image and a preset second cue message are provided to the video model to generate the intermediate video content, wherein the second cue message indicates a set of actions associated with the reference object.

5. The method according to any one of claims 1 to 4, wherein the music video is generated based on the following process: Based on the rhythm information of the music content, a set of motion effects matching the rhythm information is determined. This set of motion effects includes a first set of motion effects associated with the reference object and / or a second set of motion effects corresponding to the lyrics. The first set of motion effects corresponds to the first set of visual elements, and the second set of motion effects corresponds to the second set of visual elements. The music video is generated based on the music content and the set of animation effects.

6. The method according to any one of claims 1 to 5, wherein obtaining reference information via an interactive interface includes: The recording entry point is presented in the interactive interface; In response to the selection of the recording entry point, a recording interface is displayed; as well as The recorded singing content is obtained through the recording interface to determine the reference audio information.

7. The method according to claim 6, wherein the singing content corresponds to the preset lyrics presented in the recording interface.

8. The method according to any one of claims 1 to 7, wherein obtaining the reference information comprises: Presents preset candidate text; as well as The reference text is determined based on the editing operations performed on the candidate text.

9. The method according to any one of claims 1 to 8, further comprising: The reference image is acquired, including an uploaded image or a captured image; as well as In response to the reference image not meeting the conditions associated with the reference object, a reminder message indicating the conditions is presented.

10. The method of claim 9, wherein the condition indication is related to at least one of the following: Whether the reference object is detected in the reference image; The occlusion ratio of the reference object in the reference image; The display size of the reference object in the reference image; The display angle of the reference object in the reference image.

11. The method according to any one of claims 1 to 10, wherein the music video includes a third set of visual elements associated with a preset video template.

12. An apparatus for generating music videos, comprising: The acquisition module is configured to acquire reference information, which includes reference images, reference text, and reference audio information. as well as A providing module is configured to provide a music video generated based on the reference information, wherein the music content of the music video is generated based on the reference text and the reference audio information, and the visual content of the music video includes a first set of visual elements generated based on reference objects in the reference image and a second set of visual elements generated based on the lyrics of the music content.

13. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processor.

14. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 11.

15. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 11.