Music generation method and device, equipment and storage medium

By obtaining music content and using machine learning models to generate adapted lyrics and music, the problem of difficult to improve the efficiency and quality of music generation in the prior art is solved, and efficient lyrics and music generation is achieved.

CN120343355APending Publication Date: 2025-07-18BYTEDANCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510537436.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The efficiency and quality of music generation in the prior art are difficult to improve simultaneously, especially when adapting lyrics and music content, there is a lack of efficient methods.

Method used

By obtaining the first music content, presenting the first lyric content and generating the second lyric content in response to the generation request, and then generating the second music content based on the third lyric content and the first music content, the machine learning model and preset constraints are used to improve the lyric adaptation efficiency and music generation quality.

Benefits of technology

It realizes the rapid generation and adaptation of lyrics, and improves the efficiency of music generation on the basis of ensuring quality, enhancing user participation and interactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343355A_ABST
    Figure CN120343355A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a music generation method and device, equipment and a storage medium. The method comprises the following steps: acquiring first music content; in response to receiving the first generation request, presenting second lyric content generated based on the first lyric content of the first music content; and in response to receiving the second generation request, providing second music content generated based on third lyric content and the first music content, the third lyric content being determined based on the second lyric content. Based on the mode, the embodiment of the invention can quickly generate other second lyric content based on the first lyric content of the first music content in response to the generation request, so that the efficiency of lyric reorganization is improved. Besides, the embodiment of the invention can also determine the third lyric content based on the second lyric content, and quickly generate the new music content according to the third lyric content and the first music content, thereby not only ensuring the quality of music content generation, but also improving the efficiency of music generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to methods, apparatuses, devices, and computer-readable storage media for generating music. Background Art

[0002] With the rapid development of the Internet, music adaptation, as an important type in the field of audio creation, is widely used in various fields. For example, music adaptation includes adaptations of changing lyrics without changing the melody to generate new music. How to improve the efficiency and quality of music generation is a focus issue of concern. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for generating music is provided. The method includes: obtaining first music content; presenting second lyric content generated based on first lyric content of the first music content in response to receiving a first generation request; and providing second music content generated based on third lyric content and the first music content in response to receiving a second generation request, where the third lyric content is determined based on the second lyric content.

[0004] In a second aspect of the present disclosure, an apparatus for generating music is provided. The apparatus includes: an obtaining module configured to obtain first music content; a presenting module configured to present second lyric content generated based on first lyric content of the first music content in response to receiving a first generation request; and a providing module configured to present second lyric content generated based on first lyric content of the first music content in response to receiving a first generation request.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium and can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figures 2A - 2F An example interface diagram showing some embodiments of the present disclosure;

[0011] Figure 3 A flowchart showing the process of generating music according to some embodiments of the present disclosure;

[0012] Figure 4 A schematic structural block diagram showing an apparatus for generating music according to certain embodiments of the present disclosure;

[0013] Figure 5 A block diagram showing an electronic device capable of implementing multiple embodiments of the present disclosure. Detailed Embodiments

[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0015] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0017] Embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. All these aspects comply with the corresponding laws, regulations and related provisions. In the embodiments of the present disclosure, the collection, acquisition, processing, processing, forwarding, use, etc. of all data are carried out on the premise that the user is aware and confirms. Accordingly, when implementing the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained through appropriate means according to relevant laws and regulations. The specific notification and / or authorization methods may vary according to the actual situation and application scenarios, and the scope of the present disclosure is not limited in this regard.

[0018] In the solutions described in this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.

[0019] Embodiments of the present disclosure propose a solution for generating music. According to this solution, a first music content is obtained; in response to receiving a first generation request, a second lyric content generated based on the first lyric content of the first music content is presented; and in response to receiving a second generation request, a second music content generated based on a third lyric content and the first music content is provided, where the third lyric content is determined based on the second lyric content.

[0020] Based on such a method, the embodiments of the present disclosure can, in response to a generation request, quickly generate other second lyric contents based on the first lyric content of the first music content, improving the efficiency of lyric adaptation. In addition, the embodiments of the present disclosure can also determine the third lyric content based on the second lyric content, and quickly generate new music content according to the third lyric content and the first music content, which not only ensures the quality of music content generation but also improves the efficiency of music generation.

[0021] Example environment

[0022] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the example environment 100 may include an electronic device 110.

[0023] In this example environment 100, the electronic device 110 may run an application 120 that supports interface interaction. The application 120 may be any suitable type of application for interface interaction, and its examples may include, but are not limited to: a music application or other suitable applications.

[0024] InFigure 1 In an environment 100, if an application 120 is active, the application 120 can provide a user 140 with a presentation interaction interface 150.

[0025] In some embodiments, the interaction interface can also be provided by a browser of the electronic device 110, for example.

[0026] In some embodiments, the electronic device 110 communicates with a server 130 to implement the supply of services for the application 120. The electronic device 110 can be any type of device with a display device, such as a mobile terminal, a fixed terminal, or a portable terminal, etc., including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a Personal Communication System (PCS) device, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for a target user (such as a "wearable" circuit, etc.).

[0027] The server 130 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server 130 can include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 can provide background services for the application 120 that supports interface interaction in the electronic device 110.

[0028] A communication connection can be established between the server 130 and the electronic device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this regard. In the embodiments of the present disclosure, the server 130 and the electronic device 110 can implement signaling interaction through the communication connection therebetween.

[0029] It should be understood that the structures and functions of the various elements in environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.

[0030] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0031] Example process

[0032] The following will refer to Figures 2A - 2F to describe an example interaction process according to an embodiment of the present disclosure. Figures 2A through 2F Example interaction interfaces 200A to 200F according to some embodiments of the present disclosure are shown. Interaction interfaces 200A to 200F can be provided, for example, by Figure 1 the electronic device 110 shown. In some embodiments, interaction interfaces 200A to 200F can be provided, for example, by application 120 or a browser of electronic device 110. For example, interaction interfaces 200A to 200F can be web pages providing music generation services, etc.

[0033] In some embodiments, electronic device 110 can present interface 200A as shown in Figure 2A . In some embodiments, interface 200A can be any suitable interface, such as a music playback interface or a music search interface, etc. This interface 200A can be configured to present the currently played music content (such as the currently playing song A presented in interface 200A) or pause the played music content, etc. In some embodiments, the music content can be any suitable content including lyric content and tune information. The lyric content is used to express the theme or emotion of the lyrics in text form. The tune information indicates the basic rhythm and pitch of the music content, etc. As an example, the music content can be a song.

[0034] In some embodiments, electronic device 110 can provide a navigation bar associated with the music content in interface 200A. Electronic device 110 can provide some function components (function labels) associated with the music content in the navigation bar, and these function components can correspond to different functions associated with the music content. These function components can be a music collection component, a playback history component, etc., which will not be elaborated here.

[0035] Taking Figure 2A as an example, electronic device 110 can provide component 201 in the navigation bar. This component 201 can be used to support the secondary creation of music content (music remix or music adaptation), where the secondary creation of music means modifying, processing, and innovating on the basis of the original music content (the first music content) in any suitable way to create new music content.

[0036] Further, the electronic device 110 may present a music configuration panel 202 in the interface 200A in response to receiving a selection of the component 201. This music configuration panel 202 may be configured to present configuration entries for a plurality of music contents, so as to support the user to import the first music content to be created based on any one of these configuration entries. This music configuration panel 202 may be presented on the interface 200A in any suitable form, such as in the form of a floating window, an embedded form, etc.

[0037] Taking Figure 2A as an example, the electronic device 110 may present a configuration entry 203 and a configuration entry 204 in the music configuration panel 202. The configuration entry 203 may support the user to select a music content from a plurality of music contents provided by the application 140 as the first music content. For example, the electronic device 110 may present a default music content at the configuration entry 203 and present a selection control associated with the default music content (such as Figure 2A the selection control corresponding to "directly use" shown). Further, the electronic device 110 may determine to use the default music content as the first music content for secondary creation in response to a selection of this selection control. The default music content may be any suitable music content recommended for the user to perform secondary creation, such as a song in the playlist of user A in the application 140.

[0038] In some embodiments, the configuration entry 204 may support the user to upload or specify a music content at any path as the first music content. Specifically, this music content may be the music content locally saved on the electronic device 110, or the music content corresponding to any other suitable path. Specifically, the user may upload the first music content based on the drag-and-drop upload or file selection upload method. For example, the user may drag the music content to be adapted to the configuration entry 204 in a drag-and-drop manner to use the dragged music content as the first music content. Alternatively, the user may click on the configuration entry 204, and at this time, the electronic device 110 may present a set of candidate music contents to support the user to select any music content from this set of candidate music contents as the first music content.

[0039] In some embodiments, the uploaded music content may correspond to any suitable file format, such as MP3, WAV, M4A, etc., which will not be elaborated here.

[0040] In some embodiments, before determining the music uploaded or selected by the user as the first music content, the electronic device 110 may first perform a predetermined security check on the music content selected by the user for uploading to determine whether the music content uploaded by the user meets the security requirements. This predetermined security check can be used to detect any appropriate level, such as but not limited to detecting whether the file type corresponding to the uploaded file content meets the requirements, whether the file size meets the requirements, whether the music content complies with the security regulations, and so on. Further, the electronic device 110 may present a prompt message regarding the upload result on the interface based on the result of the security check. This upload result may include various status results such as uploading, upload success, upload failure (with reasons), and so on. For example, if the security check result fails, the electronic device 110 may present a prompt message of upload failure on the interface. For another example, if the security check result passes, the electronic device 110 may present a prompt message of upload success on the interface. For yet another example, if the security check has not ended, the electronic device 110 may present a prompt message of uploading on the interface. Specifically, the prompt message of uploading may further indicate the specific upload progress information.

[0041] For Figure 2B example, in response to the user uploading song C based on the configuration entry 204, the electronic device 110 may present the upload progress of song C as 50% in the interface 200B to indicate that song C is still in the uploading state.

[0042] In some embodiments, the electronic device 110 may also present the detailed information of the uploaded music content in the interface 200B. The detailed information may be any appropriate information, such as music duration information, music upload path information, and so on. For Figure 2B example, the electronic device 110 may present the duration corresponding to song C in the interface 200B.

[0043] In some embodiments, the electronic device 110 may also present the upload history and creation history, etc. in the interface 200B. The upload history represents the music content that the user has uploaded historically and used for secondary creation. The creation history represents the historical music content obtained after the user's historical secondary creation.

[0044] In some embodiments, the electronic device 110 may use the music content uploaded by the user as the first music content for secondary creation. In other embodiments, the electronic device 110 may also respond to receiving a cropping request from the user for the uploaded music content, crop the uploaded music content based on the cropping request, and use the obtained music content segment after cropping as the first music content. At this time, the uploaded music content may be referred to as the reference music content, and the obtained music content segment after cropping is the first music content.

[0045] Taking Figure 2C as an example, the electronic device 110 may present the interface 200C (cropping interface) as shown in Figure 2C response to receiving any appropriate operation on the song C. As an example, the electronic device 110 may present the interface 200C in response to receiving a click operation on the song C in the interface 200B. As another example, the electronic device 110 may also present a cropping entry (not shown in the figure) in the interface 200B. Further, the electronic device 110 may present the interface 200C in response to receiving a selection of the cropping entry.

[0046] As Figure 2C shown, the electronic device 110 may present an editing control 210 (first editing control) and an editing control 220 (second editing control) in the interface 200C. As an example, the editing control 210 and the editing control 220 may be presented in any appropriate area of the interface 200C. For example, the electronic device 110 may present the editing control 210 in the upper area of the interface 200C and present the editing control 220 in the lower area of the interface 200C, etc.

[0047] In some embodiments, the electronic device 110 may present the lyrics 210-1 corresponding to the song C in the editing control 210. In some embodiments, the user may perform a first operation via the editing control 210 to select a part of the lyrics 210-1, so that the electronic device 110 may determine the selected part of the lyrics 210-1 via the editing control 210. The first operation may be any appropriate operation. As an example, the electronic device 110 may present a first input box (not shown in the figure) in the interface 200C to indicate that the user may input a first start point and a first end point (such as two lines of lyrics) based on the first input box. Further, the electronic device 110 may determine the selected part of the lyrics 210-1 in response to receiving the first start point and the first end point. As another example, the electronic device 110 may present a cropping handle 210-2 in the interface 200C to indicate that the user may drag both ends of the cropping handle to select a first start point and a first end point from the lyrics 210-1. Further, the electronic device 110 may determine the selected part of the lyrics 210-1 in response to receiving the selection of the first start point and the first end point.

[0048] Further, the electronic device 110 may crop the song C based on the selected part of the lyrics 210-1 (such as the user clicks on the control 230 shown in Figure 2C to obtain the first music content. Specifically, the electronic device 110 may determine the song segment corresponding to the selected part of the lyrics 210-1 in the song C as the first music content.

[0049] In some embodiments, the editing control 220 may be configured to present a timeline 220-1 associated with song C. The timeline 220-1 is a horizontal bar area with marks indicating the playback progress of the music. The user can preview any part of the music by dragging the playback head on the timeline 220-1.

[0050] In some embodiments, the user can perform a second operation via the editing control 220 to select a selected segment of the timeline, so that the electronic device 110 can determine the selected segment of the timeline 220-1 via the editing control 220. The second operation can be any appropriate operation. As an example, the electronic device 110 may present a second input box (not shown in the figure) in the interface 200C to indicate that the user can input a second starting point and a second ending point (such as a start time and an end time) based on the second input box. Further, the electronic device 110 can determine the selected segment of the timeline 220-1 in response to receiving the second starting point and the second ending point. As another example, the electronic device 110 may present cropping handles 220-2 in the interface 200C to indicate that the user can drag both ends of the cropping handle 220-2 to select the second starting point and the second ending point of the timeline 220-1. Further, the electronic device 110 can determine the selected time period of the timeline 220-1 in response to receiving the selection of the second starting point and the second ending point.

[0051] Further, the electronic device 110 can crop song C based on the selected time period of the timeline 220-1 to obtain the first music content. Specifically, the electronic device 110 can determine the song segment corresponding to the selected time period of the timeline 220-1 in song C as the first music content.

[0052] To ensure the consistency of the content cropped based on the editing control 220 and the editing control 210, the electronic device 110 can coordinately adjust the selected time period in the second editing control based on the first operation on the editing control 210. For example, when the user drags the cropping handle 210-2 to select the lyric segment corresponding to the third line to the eighth line of the lyrics, the electronic device 110 can adjust the position of the cropping handle 220-2 so that the selected time period by the cropping handle 220-2 is the time period corresponding to the third line to the eighth line of the lyrics in the timeline.

[0053] In some embodiments, the electronic device 110 may coordinately adjust a selected portion in the first editing control based on a second operation on the editing control 220. For example, when the user drags the cropping handle 220-2 to select the time period corresponding to the third to eighth lines of lyrics, the electronic device 110 may adjust the position of the cropping handle 210-2 such that the portion selected by the cropping handle 210-2 is the portion corresponding to the third to eighth lines of lyrics in the lyrics.

[0054] For Figure 2D example, the electronic device 110 may present the interface 200D (the first interface) as shown in response to obtaining the first music content. In some embodiments, the electronic device 110 may present the first lyric content 240 corresponding to the first music content in the interface 200D. As an example, the electronic device 110 may identify the first lyric content included in the first music content based on Automatic Speech Recognition (ASR) technology. Figure 2D

[0055] To improve efficiency, the electronic device 110 may stream-recognize and present each lyric segment in the first lyric content. For the convenience of prompting, the electronic device 110 may also present the progress information on the recognition or presentation of the lyrics included in the first music content in the interface 200D. For Figure 2D example, the electronic device 110 may present the progress information 211 in the interface 200D to prompt that 48% of the lyrics included in song C have been recognized and displayed.

[0056] In some embodiments, the electronic device 110 may receive a first generation request based on the input information received via the interface 200D. The first generation request is used to indicate generating a second lyric content based on the first lyric content corresponding to the first music content. In some embodiments, the input information may indicate at least one attribute of the to-be-generated second lyric content. This at least one attribute may correspond to any appropriate attribute, such as may correspond to the theme, style, language, rhyme, the word count limit corresponding to each line of lyrics, etc. The input information may be, for example, the text content input via the interface 200D or the selection of at least one candidate item in the interface 200D, etc., which will not be elaborated here.

[0057] For Figure 2D ​As an example, the electronic device 110 can present an input text box in the interface 200D to indicate that the user can input text content via this interface 200D. The text content can be any appropriate content, such as the style, rhyme, language, theme, etc. corresponding to the second lyric content to be generated. As an example, the electronic device 110 can present a theme input box 213 in the interface 200D to indicate that the user can input the music theme corresponding to the second lyric content based on this theme input box 213 (such as "sweet", "inspiring", "funny", etc.).

[0058] In some embodiments, the electronic device 110 can also present at least one candidate item (not shown in the figure) in the interface 200D. This at least one candidate item can be any appropriate candidate item that can assist in the generation of the second lyric content. For example, the electronic device 110 can present a set of theme candidate items in the interface 200D to indicate that the user can determine the theme target item from this set of theme candidate items.

[0059] Furthermore, the electronic device 110 can receive this first generation request in response to receiving the text content input via the interface 200D or the selection of at least one candidate item and the generation control 212 is selected. Further, the electronic device 110 can present the second lyric content generated based on the first lyric content of the first music content in response to receiving the first generation request.

[0060] In some embodiments, the second lyric content can be lyric content associated with the first lyric content, which is the second lyric content obtained by adapting on the basis of the first lyric content. As an example, the second lyric content is not completely the same as the corresponding lyrics in the second lyric content, but can have similar attributes.

[0061] Taking the server to generate the second lyric content as an example, the generation process of the second lyric content is described. It should be noted that the second lyric content can be generated not only by the server, but also by other devices (such as the electronic device 110), which will not be elaborated here.

[0062] In some embodiments, the server can provide the first lyric content to the model to generate the second song content.

[0063] In some other embodiments, the server may provide the first lyric content and prompt information to the model to generate the second lyric content. The model may be any suitable machine learning model, such as a generative model. The prompt information is guiding information for guiding the model on how to generate the second lyric content based on the first lyric content. In some embodiments, the prompt information may indicate at least one lyric constraint. This at least one lyric constraint may be used to constrain the relevance between the first lyric content and the second lyric content, and may also constrain the attribute information of the second lyric content, etc. For example, the lyric constraint may constrain the number of all statements (a statement may also be referred to as a lyric sentence) included in the generated second lyric content, the length of each lyric sentence in the generated second lyric content (length constraint), where the length may be the number of characters included in each lyric, the keywords included in each statement in the second lyric content should be the same as the keywords included in the first lyric content, the second lyric content retains the blank spaces or antithesis features in the first lyric content, punctuation marks and special characters are prohibited in the second lyric content, the second lyric content may not include sensitive content, the second lyric content meets the palatability detection (ensuring that the lyric content is catchy), etc.

[0064] As an example, this at least one lyric constraint may indicate that the number of statements included in the first lyric content is the same as the number of statements included in the second lyric content. For example, if song C includes 18 lyric sentences, then 18 adapted lyric sentences are also generated based on these 18 lyric sentences in song C.

[0065] As an example, this at least one lyric constraint may indicate a length constraint, and this length constraint is determined based on the lengths of multiple lyric sentences in the first lyric content. As an example, for each lyric sentence in the first lyric content, this at least one lyric constraint may indicate that the difference between the length of this lyric in the first lyric content and the length of the corresponding lyric in the second lyric content is less than a threshold. For example, the first lyric sentence in song C includes 4 characters, and the threshold is 2, then the number of characters included in the corresponding adapted lyric sentence of this first lyric sentence may be from 2 to 6 characters.

[0066] As an example, for each lyric sentence in the first lyric content, this at least one lyric constraint may indicate that the keywords included in this lyric in the first lyric content should be kept as the same as possible as the keywords included in the corresponding lyric in the second lyric content. For example, if the first lyric sentence in song C is "XXXYYZZZ", then the corresponding adapted lyric sentence of this first lyric sentence may be "XXXCCZZZ".

[0067] In some embodiments, the electronic device 110 may obtain the media content generated by the model. Further, the electronic device 110 may directly determine the media content generated by the model as the second lyric content.

[0068] In some other embodiments, to improve the quality of the generated second lyric content, the electronic device 110 may obtain intermediate media content generated by the model. Further, the electronic device 110 may determine whether the intermediate lyric content includes a lyric part that does not meet a preset constraint. Further, in response to the intermediate lyric content including a lyric part that does not meet the preset constraint, the electronic device 110 may adjust the lyric part of the intermediate lyric content based on the first lyric content to determine the second lyric content.

[0069] As an example, if the lyric content of the intermediate media content generated by the model corresponds to a predetermined language, which is different from the language corresponding to the lyrics in the first music content, the electronic device 110 may map the lyric content corresponding to the predetermined language one by one to the positions and lyric texts of the first lyric content to determine the second lyric content, thereby ensuring that the subsequent newly generated music content maintains the same rhythm and tempo as the first music content.

[0070] In some embodiments, to more clearly present the difference between the first lyric content and the second lyric content, the electronic device 110 may present at least one hint element associated with the second lyric content. The at least one hint element indicates the difference between the first lyric content and the second lyric content. The at least one hint element may be any suitable element, such as text, image, sticker, etc., which will not be elaborated here.

[0071] Specifically, the at least one hint element includes a first hint element, and the electronic device 110 may determine whether a first length difference between a first statement in the first lyric content indicated by the first hint element and a second statement in the second lyric content is greater than a first threshold. The first statement may be any line of lyrics in the first lyric content, and the second statement may be any line of lyrics in the second lyric content, but it is necessary to ensure that the first statement and the second statement have a predetermined association relationship, that is, the statement obtained after adapting the first statement is the second statement. The first threshold may be any suitable threshold, such as 3, etc., which will not be elaborated here.

[0072] It should be noted that when there is a first length difference between the first statement and the second statement, but the first length difference is less than or equal to the first threshold, the electronic device 110 may also present a third indication element regarding the first length difference. The third indication element is used to indicate the difference between the first statement and the second statement. For the sake of distinction, the electronic device 110 may set the first indication element and the third indication element to different display forms. For example, when the first length difference is greater than the first threshold, the first indication element may be set to red. When the first length difference is less than or equal to the first threshold, the first indication element may be set to gray.

[0073] Taking Figure 2E as an example, in order to highlight the differences between Lyric 2 and Adaptation 2, the electronic device 110 may present an indication element 250 (the first indication element) in the interface 200E. The indication element 250 may indicate that the difference (the first length difference) between the number of characters included in Lyric 2 (the first sentence in the first lyric content) and the number of characters included in Adaptation 2 (the second sentence in the second lyric content) is 2, where Lyric 2 corresponds to Adaptation 2, that is, Adaptation 2 is the lyric obtained after adaptation based on Lyric 2.

[0074] Taking Figure 2E as an example, in order to highlight the relevance and differences between the first lyric content and the second lyric content, the electronic device 110 may present the first lyric content and the second lyric content in a predetermined presentation style. The predetermined presentation style may indicate that the first display position of the first sentence in the first lyric content is associated with the second display position of the second sentence in the second lyric content. For example, for each lyric in the first lyric content, the electronic device 110 may present the lyric and the adapted lyric corresponding to the lyric in the second lyric content in the same column or the same row in the interface 200E, etc. Taking Figure 2E as an example, the electronic device 110 may display Lyric 1 and Adaptation 1 (the adapted lyric corresponding to Lyric 1) in the same row in the interface 200E, and display Lyric 2 and Adaptation 2 (the adapted lyric corresponding to Lyric 2) in the same row in the interface 200E, etc.

[0075] In some embodiments, this at least one hint element may further include a second hint element, and this second hint element indicates that the difference between the number of the first sentences in the first lyric content and the number of the second sentences in the second lyric content is greater than a second threshold. The second threshold may be any appropriate threshold, such as 0. For example, if Song C includes 20 lyrics, and the second lyric content adapted based on each lyric of Song C includes 19 lyrics, then the electronic device 110 may present this second hint element (not shown in the figure) in the interface 200E. Specifically, this second indication element may indicate the difference in number, and may also indicate the target lyric included in the second lyric content, where there is no corresponding lyric in the first lyric content for the target lyric, etc.

[0076] In some embodiments, the electronic device 110 may, in response to receiving a predetermined operation on the target hint element in this at least one hint element, adjust at least one sentence in the second lyric content corresponding to the target hint element. The predetermined operation may be any appropriate operation, such as a click operation, a swipe operation, etc. Taking Figure 2EAs an example, in response to receiving a selection of the indication element 250, the electronic device 110 may adjust the length corresponding to the adaptation 2 to be the same as the length corresponding to the lyrics 2.

[0077] In some embodiments, the electronic device 110 may also present lyric editing controls for the second lyric content in the interface 200E. Specifically, this lyric editing control may support the user in independently editing each statement in the second lyric content. Further, the electronic device 110 may receive at least one editing operation related to the second lyric content via the lyric editing control to determine the third lyric content. As an example, the user may use the editing control 260 to edit the adaptation 4. For example, the user may click on the lyric corresponding to the adaptation 4 so that the electronic device 110 can present a text input box 251 for the adaptation 4. Further, the electronic device 110 may receive the edited content of the adaptation 4 via the text input box 251. At this time, the electronic device 110 may present the interface 200F as Figure 2F shown. In some embodiments, the electronic device 110 may present the updated adaptation 4 in the interface 200F, where the lyric content composed of the updated adaptation 4 and other adaptations 1, adaptation 2, etc. is the third lyric content.

[0078] As an example, the electronic device 110 may present a control 253 in the interface 200F. The control 253 may support the user in inputting a request (second generation request) for generating new music content. Further, the electronic device 110 may receive the second generation request in response to receiving an operation indicating that the control 253 is selected.

[0079] Further, the electronic device 110 may provide second music content generated based on the third lyric content and the first music content in response to receiving the second generation request. As an example, if the third lyric content meets the at least one lyric constraint, the electronic device 110 may directly generate the second music content based on the third lyric content and the first music content. As another example, if the third lyric content does not meet the at least one lyric constraint, the electronic device 110 may not be able to generate the second music content based on the third lyric content and the first music content.

[0080] For Figure 2F example, if the length difference between the adaptation 4 in the third lyric content and the corresponding lyrics 4 is greater than a predetermined threshold, after the electronic device 110 receives an operation indicating that the control 253 is selected, it may highlight the adaptation 4 in the interface 200F, such as highlighting the adaptation 4.

[0081] In some embodiments, if the length difference between the adaptation 4 in the third lyric content and the corresponding length of lyric 4 is greater than a predetermined threshold, after the electronic device 110 receives the operation of selecting the control 253, instead of generating new music content, it can present an indication element corresponding to the adaptation 4 in the interface 200F. This indication element can be used to indicate modification suggestions regarding the adaptation 4 to assist the user in modifying the adaptation 4. For example, Figure 2F As an example, the electronic device 110 can present a modification suggestion 260 regarding the adaptation 4 to indicate that the user can reduce 2 characters based on the current adaptation 4.

[0082] As an example, the electronic device 110 can also, in response to receiving a predetermined operation corresponding to the modification suggestion 260, adjust the content of the adaptation 4 (such as reducing two characters) to achieve quick and automatic modification of the adaptation 4.

[0083] It should be noted that when there is a difference between the first statement in the first lyric content of the first music content and the third statement in the third lyric content, but the difference is not greater than the threshold, the electronic device 110 can provide a window (not shown in the figure) in the interface 200F. The electronic device 110 can present the reason for the inability to generate new music content and a solution to the problem of not being able to generate new music content in the window. For example, the electronic device 110 can present a prompt message such as "It is detected that the number of characters in the adapted lyrics is too different from the original lyrics, which may affect the generation effect of the music. It is recommended to adjust the number of characters in the lyrics according to the prompt" in the window. Further, the electronic device 110 can present a first control (such as a selection control corresponding to "Continue to Generate") and a second control (such as a selection control corresponding to "Adjust Lyrics") in the window to support the user in choosing whether to continue generating new music content based on the currently adapted lyric content. As an example, the electronic device 110 can, in response to receiving the selection of the first control, generate a second music content based on the third lyric content and the first music content. As an example, the electronic device 110 can, in response to receiving the selection of the second control, present the interface 200F as shown in Figure 2F to support the user in continuing to edit the adapted lyric content.

[0084] In some embodiments, the electronic device 110 can use a predetermined model to generate a second music content based on the third lyric content and the first music content. The predetermined model can be any suitable machine learning model, such as a generative model.

[0085] As an example, the electronic device 110 can provide third lyric content, first music content, and a predetermined prompt item to the predetermined model to generate second music content. The predetermined prompt item can be used to indicate relevant constraints regarding the generation of music content. For example, it can be used to constrain the tune information corresponding to the to-be-generated second music content to be the same as the tune information of the first music content, etc.

[0086] In some embodiments, the second music content has tune information corresponding to the first music content. The tune information indicates the melodic features and structural information in the music content, such as melody, rhythm, pitch, harmony, etc.

[0087] Based on such a manner, the embodiments of the present disclosure can, in response to a generation request, quickly generate other second lyric content based on the first lyric content of the first music content, improving the efficiency of lyric adaptation. Additionally, the embodiments of the present disclosure can also determine third lyric content based on the second lyric content, and quickly generate new music content according to the third lyric content and the first music content, ensuring both the quality of music content generation and improving the efficiency of music generation.

[0088] Example process

[0089] Figure 3 The flowchart of a process 300 for generating music according to some embodiments of the present disclosure is shown. The process 300 can be implemented at the electronic device 110. The following refers to Figure 1 Describe the process 300.

[0090] In block 310, the electronic device 110 obtains the first music content.

[0091] In block 320, in response to receiving a first generation request, the electronic device 110 presents second lyric content generated based on the first lyric content of the first music content.

[0092] In block 330, in response to receiving a second generation request, the electronic device 110 provides second music content generated based on the third lyric content and the first music content, where the third lyric content is determined based on the second lyric content.

[0093] Based on such a manner, the embodiments of the present disclosure can, in response to a generation request, quickly generate other second lyric content based on the first lyric content of the first music content, improving the efficiency of lyric adaptation. Additionally, the embodiments of the present disclosure can also determine third lyric content based on the second lyric content, and quickly generate new music content according to the third lyric content and the first music content, ensuring both the quality of music content generation and improving the efficiency of music generation.

[0094] In some embodiments, process 300 further includes: presenting a first interface in response to obtaining first music content; and receiving a first generation request based on input information received via the first interface.

[0095] In this way, embodiments of the present disclosure can receive input information on the first interface, improving user engagement and interactivity, enabling users to more directly participate in the secondary creation process of music content, and enhancing the interaction efficiency and experience.

[0096] In some embodiments, the input information indicates at least one of the following: text content input via the first interface; selection of at least one candidate item in the first interface.

[0097] In this way, embodiments of the present disclosure can ensure that users can select the required text content via the first interface or make a selection from candidate items, simplifying the submission process of the generation request and improving the interaction efficiency.

[0098] In some embodiments, the input information indicates at least one attribute of the second lyric content to be generated.

[0099] In this way, embodiments of the present disclosure can set the attributes of the second lyric content to be generated, enhancing the flexibility and pertinence of lyric generation.

[0100] In some embodiments, the second lyric content is generated based on the following process: providing first lyric content and prompt information to a model to generate the second lyric content, where the prompt information indicates at least one lyric constraint.

[0101] In this way, embodiments of the present disclosure can generate the second lyric content that meets specific constraint conditions based on the first lyric content and prompt information, improving the accuracy and applicability of lyric generation.

[0102] In some embodiments, at least one lyric constraint includes a length constraint, where the length constraint indicates the length of each sentence of the second lyric content to be generated.

[0103] In this way, embodiments of the present disclosure can improve the generation quality of the second lyric content by constraining the length of each sentence of the second lyric content.

[0104] In some embodiments, the length constraint is determined based on the lengths of multiple sentences of the first lyric content.

[0105] In this way, embodiments of the present disclosure can ensure that the generated second lyric content matches the first lyric content in sentence length through the length constraint, helping to maintain the rhythm and structural consistency of the music work.

[0106] In some embodiments, the second lyric content is generated based on the following process: obtaining intermediate lyric content generated by a model; and in response to the intermediate lyric content including a lyric portion that does not meet a preset constraint, adjusting the lyric portion of the intermediate lyric content based on the first lyric content to determine the second lyric content.

[0107] In this way, the embodiments of the present disclosure can improve the quality and compliance of lyric generation by obtaining intermediate lyric content and making adjustments according to preset constraints.

[0108] In some embodiments, process 300 further includes: presenting at least one hint element associated with the second lyric content, the at least one hint element indicating the difference between the first lyric content and the second lyric content.

[0109] In this way, the embodiments of the present disclosure can intuitively display the difference between the first lyric content and the second lyric content through the hint element, helping the user quickly identify and understand the parts that need to be adjusted, thereby improving the efficiency and accuracy of editing and enhancing the information transmission efficiency.

[0110] In some embodiments, the at least one hint element includes a first hint element, the first hint element indicating a first length difference between a first statement in the first lyric content and a second statement in the second lyric content, the first statement corresponding to the second statement.

[0111] In this way, the embodiments of the present disclosure can, when detecting that the lyric length difference exceeds a preset threshold, remind the user to adjust the lyric in a timely manner through the hint element, ensuring the consistency of the generated lyric content in length and improving the information transmission efficiency.

[0112] In some embodiments, presenting at least one hint element associated with the second lyric content includes: in response to the first length difference being greater than a first threshold, presenting the first hint element associated with the first statement and the second statement.

[0113] In this way, the embodiments of the present disclosure can present the hint element when the length difference between the first statement and the second statement is greater than the threshold, avoiding the redundant display of unnecessary information and improving the information transmission efficiency.

[0114] In some embodiments, the first display position of the first statement is related to the second display position of the second statement.

[0115] In this way, the embodiments of the present disclosure can display the positions of the first statement and the second statement in an associated manner, making it more convenient for the user to view and understand the relevance and difference between the first statement and the second statement, and improving the information transmission efficiency.

[0116] In some embodiments, at least one hint element includes a second hint element that indicates that the difference between the number of first statements in the first lyric content and the number of second statements in the second lyric content is greater than a second threshold.

[0117] In this way, embodiments of the present disclosure can indicate the difference in the number of statements through the hint element, helping the user identify problems that may affect the music rhythm and structure, so as to make corresponding adjustments during the editing process.

[0118] In some embodiments, process 300 further includes: in response to receiving a predetermined operation on a target hint element in at least one hint element, adjusting at least one statement in the second lyric content corresponding to the target hint element.

[0119] In this way, embodiments of the present disclosure can adjust the corresponding statements in the second lyric content by performing a predetermined operation on the target hint element in the hint element, improving the flexibility of lyric editing and the efficiency of lyric editing.

[0120] In some embodiments, process 300 further includes: providing a lyric editing control associated with the second lyric content; and receiving, via the lyric editing control, at least one editing operation on the second lyric content to determine a third lyric content.

[0121] In this way, embodiments of the present disclosure can perform an editing operation on the second lyric content based on the lyric editing control to determine the third lyric content, effectively improving the editing efficiency, and supporting the editing of the lyric content can effectively improve the accuracy of the lyric content.

[0122] In some embodiments, obtaining the first music content includes: obtaining a reference music content; presenting a cropping interface, the cropping interface including a first editing control and a second editing control, the first editing control presenting the reference lyric content of the reference music content, and the second editing control presenting a timeline associated with the reference music content; and in response to determining a selected part of the reference lyric content via the first editing control or a selected time period of the timeline via the second editing control, cropping the reference music to obtain the first music content.

[0123] In this way, embodiments of the present disclosure can accurately crop the reference music through the first editing control or the second editing control to obtain the required first music content, improving the accuracy of music editing. In addition, embodiments of the present disclosure provide two types of editing controls corresponding to different types, and crop the parameter music through different editing methods, improving the flexibility of editing.

[0124] In some embodiments, process 300 further includes: coordinately adjusting a selected time period in a second editing control based on a first operation on a first editing control; or coordinately adjusting a selected portion in the first editing control based on a second operation on the second editing control.

[0125] In this way, embodiments of the present disclosure can avoid the problem that the music content cropped based on different editing controls is different by coordinately adjusting the selected portions and time periods in the first editing control and the second editing control, and can effectively improve the accuracy of cropping music content.

[0126] Example apparatus and equipment

[0127] Embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 FIG. shows a schematic structural block diagram of a device 400 for touch interaction according to certain embodiments of the present disclosure. Device 400 may be implemented as or included in an electronic device 110 as discussed above. Each module / component in device 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0128] As Figure 4 shown, device 400 includes an acquisition module 410 configured to acquire first music content; a presentation module 420 configured to present second lyric content generated based on first lyric content of the first music content in response to receiving a first generation request; and a provision module 430 configured to provide second music content generated based on third lyric content and the first music content in response to receiving a second generation request, where the third lyric content is determined based on the second lyric content.

[0129] In some embodiments, device 400 further includes an interface presentation module configured to: present a first interface in response to acquiring the first music content; and a first reception module configured to: receive a first generation request based on input information received via the first interface.

[0130] In some embodiments, the input information indicates at least one of the following: text content input via the first interface; selection of at least one candidate item in the first interface.

[0131] In some embodiments, the input information indicates at least one attribute of the second lyric content to be generated.

[0132] In some embodiments, the second lyric content is generated based on the following process: providing the first lyric content and prompt information to a model to generate the second lyric content, where the prompt information indicates at least one lyric constraint.

[0133] In some embodiments, at least one lyric constraint includes a length constraint, where the length constraint indicates the length of each sentence of the second lyric content to be generated.

[0134] In some embodiments, the length constraint is determined based on the lengths of multiple sentences in the first lyric content.

[0135] In some embodiments, the second lyric content is generated based on the following process: obtaining intermediate lyric content generated by a model; and in response to the intermediate lyric content including a lyric part that does not meet a preset constraint, adjusting the lyric part of the intermediate lyric content based on the first lyric content to determine the second lyric content.

[0136] In some embodiments, the apparatus 400 further includes an element presentation module configured to: present at least one prompt element associated with the second lyric content, where the at least one prompt element indicates the difference between the first lyric content and the second lyric content.

[0137] In some embodiments, the at least one prompt element includes a first prompt element that indicates a first length difference between a first sentence in the first lyric content and a second sentence in the second lyric content, and the first sentence corresponds to the second sentence.

[0138] In some embodiments, the element presentation module is further configured to present the first prompt element associated with the first sentence and the second sentence in response to the first length difference being greater than a first threshold.

[0139] In some embodiments, the first display position of the first sentence is related to the second display position of the second sentence.

[0140] In some embodiments, the at least one prompt element includes a second prompt element that indicates that the difference between the number of first sentences in the first lyric content and the number of second sentences in the second lyric content is greater than a second threshold.

[0141] In some embodiments, the apparatus 400 further includes a first adjustment module configured to: in response to receiving a predetermined operation on a target prompt element in the at least one prompt element, adjust at least one sentence in the second lyric content corresponding to the target prompt element.

[0142] In some embodiments, the apparatus 400 further includes a control providing module configured to: provide lyric editing controls associated with the second lyric content; and a second receiving module configured to receive at least one editing operation on the second lyric content via the lyric editing controls to determine a third lyric content.

[0143] In some embodiments, the acquisition module 410 is further configured to: acquire reference music content; present a cropping interface, the cropping interface including a first editing control and a second editing control, the first editing control presenting reference lyric content of the reference music content, and the second editing control presenting a timeline associated with the reference music content; and in response to determining a selected portion of the reference lyric content via the first editing control or a selected time period of the timeline via the second editing control, crop the reference music to obtain first music content.

[0144] In some embodiments, the apparatus 400 further includes a second adjustment module configured to: coordinately adjust the selected time period in the second editing control based on a first operation on the first editing control; or coordinately adjust the selected portion in the first editing control based on a second operation on the second editing control.

[0145] The units included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in the apparatus 400 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0146] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 can be used to implement Figure 1 the illustrated electronic device 110.

[0147] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be an actual or virtual processor and can execute various processes according to programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0148] The electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that the electronic device 500 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be capable of storing information and / or data (such as training data for training) and can be accessed within the electronic device 500.

[0149] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules that are configured to execute various methods or actions of various embodiments of the present disclosure.

[0150] The communication unit 540 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that can communicate through a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0151] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed through the communication unit 540, such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (such as a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0152] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0153] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0154] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0155] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, so that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0157] The implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art to understand the various implementation manners disclosed herein.

Claims

1. A method for generating music, comprising: Obtaining first music content; In response to receiving a first generation request, presenting second lyric content generated based on the first lyric content of the first music content; And In response to receiving a second generation request, providing second music content generated based on third lyric content and the first music content, where the third lyric content is determined based on the second lyric content.

2. The method according to claim 1, further comprising: In response to obtaining the first music content, presenting a first interface; And Based on input information received via the first interface, receiving the first generation request.

3. The method according to claim 2, wherein the input information indicates at least one of the following: Text content input via the first interface; Selection of at least one candidate item in the first interface.

4. The method according to claim 2, wherein the input information indicates at least one attribute of the second lyric content to be generated.

5. The method according to claim 1, wherein the second lyric content is generated based on the following process: Providing the first lyric content and prompt information to a model to generate the second lyric content, where the prompt information indicates at least one lyric constraint.

6. The method according to claim 5, wherein at least one lyric constraint includes a length constraint, and the length constraint indicates the length of each sentence of the second lyric content to be generated.

7. The method according to claim 6, wherein the length constraint is determined based on the lengths of multiple sentences of the first lyric content.

8. The method according to claim 5, wherein the second lyric content is generated based on the following process: Obtaining intermediate lyric content generated by the model; and In response to the intermediate lyric content including a lyric part that does not meet a preset constraint, adjusting the lyric part of the intermediate lyric content based on the first lyric content to determine the second lyric content.

9. The method according to claim 1, further comprising: Associating with the second lyric content, presenting at least one prompt element, where the at least one prompt element indicates the difference between the first lyric content and the second lyric content.

10. The method according to claim 9, wherein the at least one prompt element includes a first prompt element, and the first prompt element indicates a first length difference between a first sentence in the first lyric content and a second sentence in the second lyric content, and the first sentence corresponds to the second sentence.

11. The method according to claim 10, wherein associating with the second lyric content and presenting at least one prompt element includes: In response to the first length difference being greater than a first threshold, presenting the first prompt element associated with the first sentence and the second sentence.

12. The method according to claim 10, wherein the first display position of the first sentence is related to the second display position of the second sentence.

13. The method according to claim 9, wherein the at least one hint element includes a second hint element, and the second hint element indicates that the difference between the number of first statements in the first lyric content and the number of second statements in the second lyric content is greater than a second threshold value.

14. The method according to claim 9, further comprising: In response to receiving a predetermined operation on a target hint element in the at least one hint element, adjusting at least one statement in the second lyric content corresponding to the target hint element.

15. The method according to claim 1, further comprising: Providing a lyric editing control associated with the second lyric content; And Receiving, via the lyric editing control, at least one editing operation on the second lyric content to determine the third lyric content.

16. The method according to claim 1, wherein obtaining the first music content includes: Obtaining a reference music content; Presenting a cropping interface, the cropping interface including a first editing control and a second editing control, the first editing control presenting reference lyric content of the reference music content, and the second editing control presenting a timeline associated with the reference music content; And In response to determining a selected portion of the reference lyric content via the first editing control or a selected time period of the timeline via the second editing control, cropping the reference music to obtain the first music content.

17. The method according to claim 16, further comprising: Based on a first operation on the first editing control, cooperatively adjusting the selected time period in the second editing control; Or Based on a second operation on the second editing control, cooperatively adjusting the selected portion in the first editing control.

18. An apparatus for generating music, comprising: An obtaining module configured to obtain a first music content; A presenting module configured to present, in response to receiving a first generation request, a second lyric content generated based on the first lyric content of the first music content; And A providing module configured to provide, in response to receiving a second generation request, a second music content generated based on a third lyric content and the first music content, the third lyric content being determined based on the second lyric content.

19. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 17.

20. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 17.