A music generation method and apparatus

By generating timbre and rhythm encoding vectors and inputting them into a large music generation model, the problem of timbre and rhythm decoupling in existing technologies is solved, enabling timbre replication and personalized music generation, and improving the flexibility of music generation.

CN119360810BActive Publication Date: 2026-04-21SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI XIYU JIZHI TECH CO LTD
Filing Date
2024-10-21
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing music generation technologies cannot decouple timbre from rhythm, resulting in the inability to replicate individual timbre and personalize songs, thus limiting flexibility.

Method used

By acquiring timbre and rhythm information, timbre encoding vectors and rhythm encoding vectors are generated, and then input into a pre-built large music generation model to generate target music.

Benefits of technology

It achieves the decoupling of timbre and rhythm, allowing any timbre to be replicated in any rhythm, and supports the free generation of music with replicated timbre, improving the flexibility and personalization of generated music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360810B_ABST
    Figure CN119360810B_ABST
Patent Text Reader

Abstract

This invention discloses a music generation method and apparatus. The method includes: in response to a music generation event being triggered, acquiring timbre information from first audio data and generating at least one timbre encoding vector based on the timbre information; acquiring prosody information from second audio data and generating at least one prosody encoding vector based on the prosody information; acquiring target lyrics, performing word segmentation on the target lyrics, and converting each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector; inputting at least one timbre encoding vector, at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed large-scale music generation model, and generating target music based on the output of the large-scale music generation model. This solution not only achieves decoupling of timbre and prosody, thereby enabling the replication of any timbre in any prosody, but also enables the free generation of music using the replicated timbre in any prosody.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio technology, and more particularly to a method and apparatus for generating music. Background Technology

[0002] Current song generation technologies typically process vocals and accompaniment separately. The accompaniment is usually extracted from a reference or template song, while the vocals are generally derived from voice input. Referencing the original singer's vocal spectrum, the input spectrum is manipulated through operations such as extension, boosting, and speed adjustment to generate a personalized spectrum corresponding to the original singer, thus producing a cappella vocals. Finally, the accompaniment or a cappella vocals can be output separately, or the two tracks can be merged into a single output song. In other words, current music generation technologies generally only optimize vocals by adjusting their spectrum; they cannot replicate individual timbre, personalize songs, create songs from scratch, or generate songs freely, resulting in limited flexibility. Summary of the Invention

[0003] This invention provides a music generation method and apparatus that can not only decouple timbre from rhythm, thereby enabling the replication of any timbre in any rhythm, but also enable the free generation of music in any rhythm using the replicated timbre.

[0004] According to one aspect of the present invention, a music generation method is provided, comprising:

[0005] In response to a music generation event being triggered, timbre information is obtained from the first audio data, and at least one timbre encoding vector is generated based on the timbre information;

[0006] Obtain prosodic information from the second audio data, and generate at least one prosodic coding vector based on the prosodic information;

[0007] Obtain the target lyrics, perform word segmentation on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector;

[0008] The at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector are input into a pre-constructed large music generation model, and the target music is generated based on the output of the large music generation model.

[0009] According to another aspect of the present invention, a music generation apparatus is provided, comprising:

[0010] The timbre encoding vector generation module is used to respond to the music generation event being triggered, obtain timbre information from the first audio data, and generate at least one timbre encoding vector based on the timbre information;

[0011] A prosody coding vector generation module is used to acquire prosody information in the second audio data and generate at least one prosody coding vector based on the prosody information;

[0012] The word segmentation unit vector determination module is used to obtain target lyrics, perform word segmentation processing on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector;

[0013] The target music generation module is used to input the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed music generation model, and generate target music based on the output of the music generation model.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the music generation method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the music generation method according to any embodiment of the present invention.

[0019] The music generation scheme of this invention, in response to a music generation event being triggered, acquires timbre information from first audio data and generates at least one timbre encoding vector based on the timbre information; acquires prosody information from second audio data and generates at least one prosody encoding vector based on the prosody information; acquires target lyrics, performs word segmentation on the target lyrics, and converts each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector; inputs the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed large-scale music generation model, and generates target music based on the output of the large-scale music generation model. Through the technical solution provided by this invention, not only can timbre and prosody be decoupled, thereby enabling the replication of any timbre in any prosody, but also the free generation of music using the replicated timbre and any prosody can be achieved.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a music generation method provided in Embodiment 1 of the present invention;

[0023] Figure 2 This is a flowchart of a music generation method provided in Embodiment 2 of the present invention;

[0024] Figure 3 This is a flowchart of a music generation method provided in Embodiment 3 of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of a music generation device provided in Embodiment 4 of the present invention;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the music generation method of this invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a music generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to the generation of music. The method can be executed by a music generation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0031] S110. In response to the music generation event being triggered, obtain the timbre information in the first audio data, and generate at least one timbre encoding vector based on the timbre information.

[0032] For example, when a user inputs a music generation request, it is determined that a music generation event has been triggered. In response to the triggering of the music generation event, first audio data is acquired. The first audio data is the sound whose timbre needs to be replicated; it can be a human voice or an instrumental sound. The first audio data can be any content, such as a sentence spoken by the user, or the sound of preset content, such as the sound of the user reading preset content, or a piece of music specified by the user or selected from a music database. The first audio data can be sound information of any length, or sound information of a preset length, or sound information of a preset length (such as 6 seconds or 10 seconds) extracted from the user's input sound information of a minimum preset length. It should be noted that this embodiment of the invention does not limit the form, content, or length of the first audio data.

[0033] In this embodiment of the invention, timbre information is extracted from the first audio data. Timbre refers to the unique characteristics of different sounds in terms of waveform; different object vibrations produce different timbre characteristics. Timbre is also called tone color. Different vibrations can combine to create different sounds. For example, each musical instrument, different people's vocal cords, and all other vibrating objects can produce unique and distinctive sounds. When an object vibrates, it emits a fundamental tone, and its various parts also have composite vibrations. The sounds produced by the vibrations of these parts combine to form overtones. Besides a 'fundamental tone,' sound naturally includes many different 'frequencies' (the number of times a vibrating object vibrates per second) interwoven with overtones, which determines different timbres, allowing listeners to distinguish different sounds. Therefore, sound contains three aspects: 1 / the amplitude of the sound, i.e., the intensity or amplitude of the audio frequency; 2 / the frequency of the sound, i.e., the frequency of the audio frequency or the number of times it changes per second; 3 / the timbre of the sound, which is determined by overtones. Optionally, obtaining timbre information from the first audio data includes: normalizing the first audio data by speed and / or pitch, and extracting timbre information from the normalized first audio data. For example, the first audio data is normalized, wherein the normalization process includes at least one of speed normalization and pitch (fundamental tone) normalization, so that only timbre information is retained in the first audio data, thereby enabling the extraction of timbre information from the normalized first audio data. Speed ​​normalization is the normalization of the interval between two notes by stretching or compressing to improve timbre extraction and eliminate the influence of speech rate; pitch normalization is the normalization of the loudness and / or pitch (frequency) of the fundamental tone to better extract timbre features other than pitch ("overtones").

[0034] In this embodiment of the invention, at least one timbre encoding vector with at least one layer is generated based on timbre information. For example, the timbre information can be encoded as a whole into at least one global vector with at least one layer using an encoder, and this global vector can be used as the timbre encoding vector. Alternatively, a portion of the timbre of a preset duration can be extracted from the timbre information, and the extracted portion of the preset duration can be encoded into at least one timbre encoding vector with at least one layer using an encoder. The encoder may include a MERT encoder and a Mel encoder.

[0035] S120. Obtain prosodic information from the second audio data, and generate at least one prosodic encoding vector based on the prosodic information.

[0036] In this embodiment of the invention, second audio data is acquired, and prosodic information in the second audio data is determined. The second audio data is the sound from which timbre needs to be removed; it can be human voice or instrumental sound. For example, the second audio data can be user-inputted reference music (i.e., a song containing human voice and accompaniment), or it can be music extracted from a portion of the user-inputted reference music, such as the a cappella vocals or accompaniment in a song. Optionally, acquiring the prosodic information in the second audio data includes: adjusting the second audio data using at least one of the following adjustment methods: random pitch perturbation, formant adjustment, and equalizer adjustment, and extracting the prosodic information from the adjusted second audio data. For example, the second audio data is perturbed to acquire the prosodic information. The perturbation process includes at least one of the following adjustment methods: random pitch perturbation, formant adjustment, and equalizer adjustment. Random pitch perturbation includes random perturbation of amplitude (loudness) and / or frequency (pitch); formant adjustment includes the position of the formants in the spectrum; and equalizer adjustment includes the proportion of audio energy at different frequencies. After adjusting the second audio data using the above-mentioned random perturbation and / or preset adjustment methods, the timbre information in the second audio data can be removed, while the rhythm information is retained.

[0037] In this embodiment of the invention, at least one prosodic encoding vector of at least one layer is generated based on the prosodic information in the second audio data. For example, the prosodic information can be encoded as a whole into at least one global vector of at least one layer using an encoder, and this global vector can be used as the prosodic encoding vector. Alternatively, a portion of the prosodic information of a preset duration can be extracted, and the extracted portion of the prosodic information of the preset duration can be encoded into at least one prosodic encoding vector of at least one layer using an encoder. The encoder may include a MERT encoder and a Mel encoder.

[0038] S130. Obtain the target lyrics, perform word segmentation on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector.

[0039] In this embodiment of the invention, target lyrics are obtained. These target lyrics can be lyrics input by the user (text), lyrics input by the user (voice), or lyrics composed of subtitles or voice information from an input video. The target lyrics can be any content and of any length; however, this embodiment does not limit the content or length of the target lyrics. The target lyrics are segmented using a preset word segmentation algorithm, dividing them into multiple segmentation units. Each segmentation unit can be a single character, a single word, or a single sentence. Each segmentation unit in the target lyrics is encoded using an encoder, generating a corresponding segmentation unit vector with at least one layer. The more layers of the segmentation unit vector, the richer the information in the target lyrics, such as emotional information.

[0040] S140. Input the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed music generation model, and generate the target music based on the output of the music generation model.

[0041] In this embodiment of the invention, at least one timbre encoding vector, at least one prosody encoding vector, and the word segmentation unit vector corresponding to each word segmentation unit are input into a pre-constructed large-scale music generation model. The large-scale music generation model then outputs a feature vector corresponding to the music based on the input information, and converts the feature vector into the target music using a music generation module. The large-scale music generation model is a pre-trained machine learning model; for example, it can be any open-source large-scale model, such as GPT.

[0042] The music generation method of this invention, in response to a music generation event being triggered, acquires timbre information from first audio data and generates at least one timbre encoding vector based on the timbre information; acquires prosody information from second audio data and generates at least one prosody encoding vector based on the prosody information; acquires target lyrics, performs word segmentation on the target lyrics, and converts each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector; inputs the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed large-scale music generation model, and generates target music based on the output of the large-scale music generation model. Through the technical solution provided by this invention, not only can timbre and prosody be decoupled, thereby enabling the replication of any timbre in any prosody, but also the free generation of music using the replicated timbre and any prosody can be achieved.

[0043] Example 2

[0044] Figure 2 This is a flowchart of a music generation method provided in Embodiment 2 of the present invention, as follows: Figure 2 As shown, the method includes:

[0045] S210. When a user input request for general music generation is received, it is determined that a music generation event has been triggered, and the timbre information in the first audio data is obtained. At least one timbre encoding vector is generated based on the timbre information.

[0046] S220. Obtain at least one target reference music, and use the at least one target reference music as second audio data, and extract rhythm information from the second audio data.

[0047] In this embodiment of the invention, when a user wants to generate music of a certain type with the same rhythm as the accompaniment and / or vocal parts, a general music generation request is input. This general music generation request can include any of the following: accompaniment music generation request, a cappella music generation request, and song music generation request. When the user inputs a general music generation request, a music generation event is determined to be triggered. In response to the music generation event being triggered, at least one target reference music of a preset duration (e.g., 5 seconds or 10 seconds) matching the general music generation request is obtained. The at least one target reference music is music containing at least one different element, feature, style, or dimension. For example, when the general music generation request is a cappella music generation request, the target reference music can be a cappella music containing only vocals; when the general music generation request is accompaniment music generation request, the target reference music can be accompaniment music containing only accompaniment; when the general music generation request is song music generation request, the target reference music can be song music containing both vocals and accompaniment. The at least one target reference music is used as second audio data, and rhythmic information is extracted from the second audio data.

[0048] S230. The prosodic information is encoded by an encoder to generate at least one prosodic encoding vector.

[0049] S240. Obtain the target lyrics, perform word segmentation on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector.

[0050] S250. Input the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed music generation model, and generate the target music based on the output of the music generation model.

[0051] In this embodiment of the invention, when the general music generation request is a song music generation request, the music generation big model is a song generation big model. This means that at least one timbre encoding vector, at least one rhythm encoding vector generated based on the rhythm information of at least one target reference music, and the word segmentation unit vector corresponding to the target lyrics are input into the song generation big model. Based on the song generation big model, a target song music containing the target lyrics and replicating the timbre of the first audio data is generated. The rhythm of the target song music is the same as or similar to the accompaniment rhythm and vocal rhythm of the target reference music. When the general music generation request is an a cappella music generation request, the music generation big model is an a cappella generation big model. This means that at least one timbre encoding vector, at least one rhythm encoding vector generated based on the rhythm information of at least one target reference music, and the word segmentation unit vector corresponding to the target lyrics are input into the a cappella generation big model. Based on the a cappella generation big model, a target a cappella music containing the target lyrics and replicating the timbre of the first audio data is generated. The rhythm of the target a cappella music is the same as or similar to the vocal rhythm of the target reference music. When the general music generation request is for accompaniment music generation, the large music generation model is an accompaniment generation model. This means that at least one timbre encoding vector, at least one prosodic encoding vector generated based on the prosodic information of at least one target reference music, and the word segmentation unit vector corresponding to the target lyrics are input into the accompaniment generation model. The model then generates a target accompaniment music with a timbre that replicates the first audio data. This target accompaniment music has the same or similar prosodic rhythm as the target reference music. It is important to note that the song generation model, a cappella generation model, and accompaniment generation model mentioned above are the same model. By controlling the content of the encoding vectors input to the model, the model can output encoding vectors of different data types.

[0052] The technical solution provided by the embodiments of the present invention can not only decouple timbre and rhythm, thereby enabling the replication of any timbre in any rhythm, but also enable the free generation of music in any rhythm using the replicated timbre.

[0053] Example 3

[0054] Figure 3 This is a flowchart of a music generation method provided in Embodiment 3 of the present invention, as follows: Figure 3 As shown, the method includes:

[0055] S310. When a personalized music generation request is received from the user, it is determined that a music generation event has been triggered, and the timbre information in the first audio data is obtained. At least one timbre encoding vector is generated based on the timbre information.

[0056] For example, a personalized music generation request may include any one or more of the following: a music generation request containing a cappella rhythm, a music generation request containing accompaniment rhythm, or a music generation request containing song rhythm. Optionally, the personalized music generation request may also be a music generation request containing more refined and personalized rhythms specified by the user, such as a music generation request containing male (bass, baritone, tenor) a cappella rhythm, a music generation request containing female (contralto, mezzo-soprano, soprano) a cappella rhythm, a music generation request containing a cappella rhythms of one or more roles, a music generation request containing a cappella rhythms of various languages ​​(such as Chinese, dialects, English, etc.), a music generation request containing accompaniment rhythms of one or more instruments (such as piano), a music generation request containing solo singing / solo rhythms, and a music generation request containing choral multi-instrument ensemble rhythms. When a personalized music generation request input by the user is received, it is determined that a music generation event has been triggered, and the timbre information in the first audio data is obtained, and at least one timbre encoding vector is generated based on the timbre information.

[0057] S320. Obtain at least one target reference music, and extract at least one dimension of the target reference music that matches the personalized music generation request from the at least one target reference music, and use the at least one dimension of the target reference music as the second audio data.

[0058] When a personalized music generation request is received from a user, at least one target reference music is acquired, and at least one dimension of the target reference music matching the personalized music generation request is extracted from the at least one target reference music, using this at least one dimension of the target reference music as second audio data. For example, the at least one target reference music is analyzed to extract two dimensions of target reference music: the accompaniment and the vocals. Various methods can be used to extract the audio data of the vocals and the accompaniment from the audio of the target reference music. For instance, a Fourier transform can be performed on the target reference music to obtain the mixed amplitude spectrum and mixed phase spectrum of the accompaniment and vocal signals. Then, a trained separation model (which can be a model built based on a deep neural network) can be used to separate the amplitude spectra of the vocals and the accompaniment. An inverse Fourier transform is then performed on the separated amplitude spectra and mixed phase spectra of the vocals and the accompaniment to obtain the audio data of the vocals and the accompaniment. Alternatively, audio filtering can be used to extract the audio data of the vocals and the accompaniment from the target reference music. As another example, the audio of the target reference music can be copied to two tracks, and the high-frequency vocal signal and the low-frequency accompaniment signal can be filtered out respectively to achieve the separation of the vocal signal and the accompaniment signal.

[0059] In this embodiment of the invention, if the personalized music generation request is a music generation request containing a male a cappella rhythm, then the target reference music is analyzed, and the male vocal part is extracted from the target reference music; if the personalized music generation request is a music generation request containing accompaniment rhythms of multiple instruments, then the target reference music is analyzed, and the accompaniment part of each instrument is extracted from the target reference music. It should be noted that this embodiment of the invention does not limit the number of dimensions of the target reference music part matching the personalized music generation request.

[0060] S330. For each dimension of the target reference music portion in the second audio data, extract rhythmic information from the target reference music portion, and encode the rhythm using an encoder to generate at least one rhythmic encoding vector.

[0061] In this embodiment of the invention, rhythmic information is extracted from the target reference music part in each dimension, and the rhythmic information extracted from the target reference music part in each dimension is encoded by an encoder to generate at least one corresponding rhythmic encoding vector.

[0062] S340. Obtain the target lyrics, perform word segmentation on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector.

[0063] S350. Input the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed music generation model, and generate the target music based on the output of the music generation model.

[0064] The technical solution provided by the embodiments of the present invention can not only decouple timbre and rhythm, thereby enabling the replication of any timbre in any rhythm, but also enable the free generation of music with any rhythm using the replicated timbre. This high degree of flexibility greatly improves the diversity of the generated target music and meets the user's needs for personalized music generation.

[0065] In some embodiments, before inputting the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into a pre-constructed large-scale music generation model, the method further includes: obtaining a music training sample set; wherein the music training sample set contains at least two sets of music training samples, each set of music training samples including fitted music, at least one reference timbre, at least one reference prosody, and at least one sample lyric; encoding the fitted music, the reference timbre, the reference prosody, and the sample lyric in each set of music training samples respectively to generate at least one corresponding sample encoding vector; and training the pre-constructed large-scale model based on the sample encoding vector to generate a large-scale music generation model.

[0066] In this embodiment of the invention, a music training sample set is obtained, wherein the music training sample set contains at least two sets of music training samples, each set of music training samples including fitted music, at least one reference timbre, at least one reference rhythm, and at least one sample lyric. For example, a certain mature piece of music can be used as the fitted music in a certain set of music training samples, the timbre information extracted from the mature music can be used as the reference timbre, the rhythm information extracted from the mature music can be used as the reference rhythm, and any lyrics from the mature music can be used as the sample lyric in the set of music training samples. At least one corresponding sample encoding vector is generated based on the fitted music, reference timbre, reference rhythm, and sample lyric in each set of music training samples, that is, at least one corresponding sample encoding vector is generated based on the fitted music, at least one corresponding sample encoding vector is generated based on the reference timbre, at least one corresponding sample encoding vector is generated based on the reference rhythm, and at least one corresponding sample encoding vector is generated based on the sample lyric in each set of music training samples. The sample encoding vector is an encoding vector with at least one layer. The music generation model is generated by training a pre-defined large model based on the sample encoding vector. The pre-defined large model can be any open-source large model, such as GPT.

[0067] Optionally, the method for obtaining the reference timbre and reference rhythm in each set of music training samples includes: obtaining a first sample reference music and obtaining a second sample reference music of the same type or singer as the first sample reference music; determining the reference rhythm based on the first sample reference music and determining the reference timbre based on the second sample reference music. For example, a certain mature piece of music can be used as the first sample reference music, obtaining the reference rhythm of at least a portion of the music in the first sample reference music, and using this reference rhythm as the reference rhythm in a certain set of music training samples. A second sample reference music of the same type or singer as the first sample reference music is obtained, and the timbre information in the second sample reference music is used as the reference timbre. For example, the voice data of the singer of the first sample reference music in other contexts can be used as timbre reference data, or the singing data of the singer of the first sample reference music outside the first sample reference music can be used as timbre reference data, or audio of the same instrument played in other contexts can be used as timbre reference data.

[0068] Optionally, determining the reference rhythm based on the first sample reference music includes: extracting at least one dimension of the first sample reference music from the first sample reference music; and extracting the reference rhythm from the first sample reference music portion of all dimensions or a first part of the dimensions. For example, the first sample reference music can be split into a preset number of dimensions, such as splitting it into at least one dimension of the first sample reference music portion (e.g., sample reference music portions of both accompaniment and vocal parts). Extracting the reference rhythm from the first sample reference music portion of all dimensions can be done, for example, extracting the reference rhythm from the accompaniment part and extracting the reference rhythm from the vocal part of the first sample reference music. Optionally, the reference rhythm can also be extracted from the first sample reference music portion of a first part of the dimensions, for example, extracting the reference rhythm only from the accompaniment part of the first sample reference music, or extracting the reference rhythm only from the vocal part of the first sample reference music.

[0069] Optionally, the method for obtaining the reference timbre and reference rhythm in each set of music training samples includes: obtaining a third sample reference music and extracting at least one dimension of the second sample reference music from the third sample reference music; extracting reference rhythm from the second sample reference music from all dimensions or the first part of the dimensions, and extracting reference timbre from the second sample reference music from all dimensions or the second part of the dimensions. For example, a mature piece of music can be used as the third sample reference music, and the third sample reference music can be split into a preset number of dimensions, such as splitting at least one dimension of the second sample reference music from the third sample reference music (e.g., sample reference music from two dimensions: accompaniment and vocals). Reference rhythm is extracted from the second sample reference music from all dimensions, for example, extracting reference rhythm from the accompaniment and vocals of the third sample reference music respectively. Optionally, reference rhythm can also be extracted from the third sample reference music from the first part of the dimensions, for example, extracting reference rhythm only from the accompaniment or only from the vocals of the third sample reference music. Reference timbres are extracted from the second sample reference music portion across all dimensions. For example, reference timbres can be extracted from the accompaniment portion and the vocal portion of the third sample reference music, respectively. Optionally, reference timbres can also be extracted from the third sample reference music portion of the second dimension. For example, reference timbres can be extracted only from the accompaniment portion of the third sample reference music, or only from the vocal portion of the third sample reference music.

[0070] Example 4

[0071] Figure 4This is a schematic diagram of a music generation device provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes:

[0072] The timbre encoding vector generation module 410 is used to, in response to the music generation event being triggered, acquire timbre information in the first audio data, and generate at least one timbre encoding vector based on the timbre information;

[0073] The prosody encoding vector generation module 420 is used to acquire prosody information in the second audio data and generate at least one prosody encoding vector based on the prosody information.

[0074] The word segmentation unit vector determination module 430 is used to obtain target lyrics, perform word segmentation processing on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector;

[0075] The target music generation module 440 is used to input the at least one timbre encoding vector, the at least one prosody encoding vector and the word segmentation unit vector into a pre-constructed music generation model, and generate target music based on the output of the music generation model.

[0076] Optionally, in response to a music generation event being triggered, the following include:

[0077] When a user inputs a general music generation request, it is determined that a music generation event has been triggered;

[0078] Correspondingly, the prosody encoding vector generation module is used for:

[0079] Acquire at least one target reference music, and use at least one target reference music as second audio data, and extract prosodic information from the second audio data;

[0080] The prosodic information is encoded by an encoder to generate at least one prosodic encoding vector.

[0081] Optionally, in response to a music generation event being triggered, the following include:

[0082] When a personalized music generation request is received from the user, it is determined that the music generation event has been triggered;

[0083] Correspondingly, the prosody encoding vector generation module is used for:

[0084] Obtain at least one target reference music, and extract at least one dimension of the target reference music that matches the personalized music generation request from at least one target reference music, and use the at least one dimension of the target reference music as the second audio data;

[0085] For each dimension of the target reference music portion in the second audio data, rhythmic information is extracted from the target reference music portion, and the rhythm is encoded by an encoder to generate at least one rhythmic encoding vector.

[0086] Optional, a timbre encoding vector generation module, used for:

[0087] The first audio data is normalized in terms of speed and / or pitch, and timbre information is extracted from the normalized first audio data.

[0088] Optional, a prosodic encoding vector generation module, used for:

[0089] The second audio data is adjusted using at least one of the following methods: random pitch perturbation, formant adjustment, and equalizer adjustment. Prosodic information is then extracted from the adjusted second audio data.

[0090] Optionally, the device further includes:

[0091] The music training sample set acquisition module is used to acquire a music training sample set before inputting the at least one timbre encoding vector, the at least one prosody encoding vector and the word segmentation unit vector into the pre-constructed music generation large model; wherein, the music training sample set contains at least two sets of music training samples, each set of music training samples includes fitted music, at least one reference timbre, at least one reference prosody and at least one sample lyric.

[0092] The sample encoding vector generation module is used to encode the fitted music, the reference timbre, the reference rhythm and the sample lyrics in each group of music training samples, and generate at least one corresponding sample encoding vector.

[0093] The music generation large model generation module is used to train a preset large model based on the sample encoding vector to generate a music generation large model.

[0094] Optionally, the method for obtaining the reference timbre and reference rhythm in each set of music training samples includes:

[0095] Obtain a first sample reference music, and obtain a second sample reference music that is the same type or singer as the first sample reference music;

[0096] The reference rhythm is determined based on the first sample reference music, and the reference timbre is determined based on the second sample reference music.

[0097] Optionally, determining a reference rhythm based on the first sample reference music includes:

[0098] Extract at least one dimension of the first sample reference music portion from the first sample reference music;

[0099] Extract reference rhythm from the first sample reference music portion of the full dimension or the first part dimension.

[0100] Optionally, the method for obtaining the reference timbre and reference rhythm in each set of music training samples includes:

[0101] Obtain a third sample reference music, and extract at least one dimension of the second sample reference music portion from the third sample reference music;

[0102] Extract reference prosody from the second sample reference music portion of all dimensions or the first part of the dimensions, and extract reference timbre from the second sample reference music portion of all dimensions or the second part of the dimensions.

[0103] The music generation apparatus provided in the embodiments of the present invention can execute the music generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0104] Example 5

[0105] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0106] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0107] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0108] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as music generation methods.

[0109] In some embodiments, the music generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the music generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the music generation method by any other suitable means (e.g., by means of firmware).

[0110] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0111] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0112] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0114] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0115] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0116] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.

[0117] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for generating music, characterized in that, include: In response to a music generation event being triggered, timbre information is obtained from the first audio data, and at least one timbre encoding vector with at least one layer is generated based on the timbre information; wherein, a portion of the timbre for a first preset duration is extracted from the timbre information, and the extracted portion of the timbre for the first preset duration is encoded into at least one timbre encoding vector with at least one layer by a timbre encoder. Obtain prosodic information from the second audio data, and generate at least two prosodic encoding vectors with at least one layer based on the prosodic information; wherein, the second audio data is a sound from which timbre needs to be removed; extract a portion of the prosodic information for a second preset duration, and encode the extracted portion of the prosodic information for the second preset duration into at least one prosodic encoding vector with at least one layer through a prosodic encoder; The target lyrics are obtained, and the target lyrics are segmented into words. Each segmentation unit in the target lyrics is converted into a corresponding segmentation unit vector with at least one layer. The more layers of the segmentation unit vector, the richer the information in the target lyrics can be represented. The at least one timbre encoding vector, the at least two prosody encoding vectors, and the word segmentation unit vector are input into a pre-constructed music generation model, and the target music is generated based on the output of the music generation model; wherein, the model type of the music generation model is controlled according to the music generation request type, and the encoding vector content input to the music generation model is controlled so that the music generation model outputs encoding vectors of different data types; The process of acquiring prosodic information from the second audio data and generating at least two prosodic coding vectors with at least one layer based on the prosodic information includes: Extract at least two-dimensional target reference music portions from at least one target reference music, use the at least two-dimensional target reference music portions as second audio data, and generate at least two prosodic coding vectors with at least one layer. The step of obtaining prosodic information from the second audio data includes: The second audio data is adjusted using at least one of the following methods: random pitch perturbation, formant adjustment, and equalizer adjustment; and prosodic information is extracted from the adjusted second audio data. Among them, the response to the music generation event being triggered includes: When a personalized music generation request is received from the user, it is determined that a music generation event has been triggered; wherein, the personalized music generation request includes multiple music generation requests that include a cappella rhythm, music generation requests that include accompaniment rhythm, music generation requests that include song rhythm, or music generation requests that include more detailed and personalized rhythms specified by the user.

2. The method according to claim 1, characterized in that, In response to a music generation event being triggered, including: When a user inputs a general music generation request, it is determined that a music generation event has been triggered; Accordingly, obtaining prosodic information from the second audio data and generating at least one prosodic coding vector based on the prosodic information includes: Acquire at least one target reference music, use the at least one target reference music as second audio data, and extract prosodic information from the second audio data; The prosodic information is encoded by an encoder to generate at least one prosodic encoding vector.

3. The method according to claim 1, characterized in that, In response to a music generation event being triggered, including: When a personalized music generation request is received from the user, it is determined that the music generation event has been triggered; Accordingly, obtaining prosodic information from the second audio data and generating at least one prosodic coding vector based on the prosodic information includes: Obtain at least one target reference music, and extract at least one dimension of the target reference music that matches the personalized music generation request from at least one target reference music, and use the at least one dimension of the target reference music as the second audio data; For each dimension of the target reference music portion in the second audio data, rhythmic information is extracted from the target reference music portion, and the rhythm is encoded by an encoder to generate at least one rhythmic encoding vector.

4. The method according to claim 1, characterized in that, The step of obtaining the timbre information from the first audio data includes: The first audio data is normalized in terms of speed and / or pitch, and timbre information is extracted from the normalized first audio data.

5. The method according to claim 1, characterized in that, Before inputting the at least one timbre encoding vector, the at least one prosody encoding vector, and the word segmentation unit vector into the pre-constructed large-scale music generation model, the method further includes: Obtain a music training sample set; wherein the music training sample set contains at least two sets of music training samples, each set of music training samples includes fitted music, at least one reference timbre, at least one reference rhythm and at least one sample lyric; Encode the fitted music, reference timbre, reference rhythm, and sample lyrics in each group of music training samples to generate at least one corresponding sample encoding vector; The preset large model is trained based on the sample encoding vector to generate a large music generation model.

6. The method according to claim 5, characterized in that, The methods for obtaining the reference timbre and reference rhythm in each set of music training samples include: Obtain a first sample reference music, and obtain a second sample reference music that is the same type or singer as the first sample reference music; The reference rhythm is determined based on the first sample reference music, and the reference timbre is determined based on the second sample reference music.

7. The method according to claim 6, characterized in that, Determining a reference rhythm based on the first sample reference music includes: Extract at least one dimension of the first sample reference music portion from the first sample reference music; Extract reference rhythm from the first sample reference music portion of the full dimension or the first part dimension.

8. The method according to claim 5, characterized in that, The methods for obtaining the reference timbre and reference rhythm in each set of music training samples include: Obtain a third sample reference music, and extract at least one dimension of the second sample reference music portion from the third sample reference music; Extract reference prosody from the second sample reference music portion of all dimensions or the first part of the dimensions, and extract reference timbre from the second sample reference music portion of all dimensions or the second part of the dimensions.

9. A music generation device, characterized in that, include: The timbre encoding vector generation module is used to respond to the music generation event being triggered, obtain timbre information from the first audio data, and generate at least one timbre encoding vector with at least one layer based on the timbre information; wherein, a portion of the timbre for a first preset duration is extracted from the timbre information, and the extracted portion of the timbre for the first preset duration is encoded into at least one timbre encoding vector with at least one layer by a timbre encoder. A prosody encoding vector generation module is used to obtain prosody information from second audio data and generate at least two prosody encoding vectors with at least one layer based on the prosody information; wherein, the second audio data is a sound from which timbre needs to be removed; a portion of the prosody of a second preset duration is extracted from the prosody information, and the extracted portion of the prosody of the second preset duration is encoded into at least one prosody encoding vector with at least one layer by a prosody encoder. The word segmentation unit vector determination module is used to acquire target lyrics, perform word segmentation on the target lyrics, and convert each word segmentation unit in the target lyrics into a corresponding word segmentation unit vector with at least one layer; wherein, the more layers of the word segmentation unit vector, the more rich the information in the target lyrics can be represented; The target music generation module is used to input the at least one timbre encoding vector, the at least two prosody encoding vectors, and the word segmentation unit vector into a pre-constructed music generation model, and generate target music based on the output of the music generation model; wherein, the model type of the music generation model is controlled according to the music generation request type, and the music generation model outputs encoding vectors of different data types by controlling the encoding vector content input to the music generation model; The prosody encoding vector generation module is used for: Extract at least two-dimensional target reference music portions from at least one target reference music, use the at least two-dimensional target reference music portions as second audio data, and generate at least two prosodic coding vectors with at least one layer. The step of obtaining prosodic information from the second audio data includes: The second audio data is adjusted using at least one of the following methods: random pitch perturbation, formant adjustment, and equalizer adjustment; and prosodic information is extracted from the adjusted second audio data. Among them, the response to the music generation event being triggered includes: When a personalized music generation request is received from the user, it is determined that a music generation event has been triggered; wherein, the personalized music generation request includes multiple music generation requests that include a cappella rhythm, music generation requests that include accompaniment rhythm, music generation requests that include song rhythm, or music generation requests that include more detailed and personalized rhythms specified by the user.

Citation Information

Patent Citations

  • Voice conversion method, device and equipment and computer readable medium

    CN117612545A

  • Speech synthesis method, device and equipment and computer readable storage medium

    CN118155602A

  • Waveform generating method and appts. thereof

    CN1383129A