Generating target music using machine learning model
By introducing multi-source classifier-free guidance and causal bias into the music audio generation model, the problem that existing models cannot be directly used in iterative synthesis methods is solved, and higher quality and diverse music generation is achieved.
Patent Information
- Application Number
- CN202411513493.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-10-28
- Publication Date
- 2025-06-06
AI Technical Summary
Existing music audio generation models cannot be used directly in iterative synthesis methods, and cannot effectively utilize text descriptions and style category information.
The non-autoregressive language model based on the transformer backbone is adopted, combining the causal deviations during multi-source classifier-free boot and iterative decoding to generate end-to-end music audio.
A generative model that can listen to music backgrounds and generate appropriate responses is implemented, supporting existing music production workflows, and the quality and diversity of generated music are improved.
Smart Images

Figure CN120108362A_ABST
Abstract
Description
Background Art
[0001] Machine learning models are increasingly being used in a variety of industries to perform a variety of different tasks. Such tasks may include audio-related tasks. Improved techniques for utilizing machine learning models to perform audio-related tasks are desired. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The following detailed description may be better understood when read in conjunction with the accompanying drawings.For purposes of illustration, example embodiments of various aspects of the disclosure are shown in the drawings; however, the invention is not limited to the specific methods and instrumentalities disclosed.
[0003] Figure 1 An example system for generating target music using a machine learning model according to the present disclosure is shown.
[0004] Figure 2 An example system for generating continuous embeddings according to the present disclosure is shown.
[0005] Figure 3 An example system for training a machine learning model to generate target music according to the present disclosure is shown.
[0006] Figure 4 An example system for generating training data pairs according to the present disclosure is shown.
[0007] Figure 5 An example process for generating target music using a machine learning model according to the present disclosure is shown.
[0008] Figure 6 An example process for generating target music using a machine learning model according to the present disclosure is shown.
[0009] Figure 7 An example process for generating target music using a machine learning model according to the present disclosure is shown.
[0010] Figure 8 An example process for generating target music using a machine learning model according to the present disclosure is shown.
[0011] Fig. 9 An example process for generating a context-text embedding according to the present disclosure is shown.
[0012] Fig.10 An example process for generating a target-text embedding according to the present disclosure is shown.
[0013] Fig.11 An example process for generating a continuous embedding according to the present disclosure is shown.
[0014] Fig.12 An example process for training a machine learning model to generate target music according to the present disclosure is shown.
[0015] Fig.13 An example process of a process for generating training data pairs according to the present disclosure is shown.
[0016] Fig.14 An example process for training a machine learning model to generate target music according to the present disclosure is shown.
[0017] Fig.15A , Fig. 15B , Fig.16A , Fig. 16B Results of evaluating the performance of a machine learning model according to the present disclosure are shown.
[0018] Fig.17 An example computing device is shown that can be used to perform any of the techniques disclosed herein. DETAILED DESCRIPTION
[0019] End-to-end musical audio generation using deep learning techniques is a fairly new discipline. Recent work has shown substantial improvements in the quality and diversity of generated music by borrowing techniques from the fields of image and language processing. Some methods use techniques from the Large Language Model (LLM) literature to operate on audio representations of tokens, while others use score matching techniques to generate audio directly or encode into continuous latent representations.
[0020] Music can generally be thought of as the sum of several independent but closely related individual parts, conceived with different hierarchies and text layouts, which come together to produce a complete piece of music. A part can correspond to the performance of a musician, a specific instrument, or something more abstract (such as the output of a sampler or synthesizer). These parts are often colloquially referred to as stems. Music production can be thought of as an iterative process, the goal of which is to add and refine these stems to suit the aesthetic preferences of the producer or musician. For example, in a first iteration, a stem corresponding to a piano performance can be added to a background audio clip. The background audio can be, for example, a recording of a vocal or a guitar performance. Then, in a second iteration, a stem corresponding to a drum performance can be added to the piano background audio mix. Additional stems can be added during additional iterations to produce the final piece of music.
[0021] In order for the final music piece to sound coherent, each of the added tracks needs to be constructed to fit the context of the existing piece both musically and textually. Therefore, a generative model that can perform this task is an ideal tool for music makers. Such a generative model can support existing music production workflows rather than replace them. However, most existing generative models for musical audio are conditioned on relatively abstract information, ranging from textual descriptions to stylistic categories. Therefore, these existing generative models cannot be directly used in this iterative synthesis approach.
[0022] This paper describes an end-to-end generative model that is able to listen to the musical context and generate appropriate responses. The architecture of the end-to-end generative model described in this paper is based on a non-autoregressive language model with a transformer backbone. The end-to-end generative model described in this paper exploits two novel improvements to audio generation: multi-source classifier-free guidance (CFG) and causal bias during iterative decoding. The end-to-end generative model is trained on two datasets based on tracked music audio and evaluated using standard objective metrics, a novel evaluation method that combines Music Information Retrieval (MIR) metrics and listening tests.
[0023] Figure 1 An example system 100 according to the present disclosure is illustrated. System 100 may include a machine learning model 101. Machine learning model 101 may be configured to generate target audio. Audio 102 may be input into machine learning model 101 (e.g., received by the machine learning model). Audio 102 may indicate the background of the target music to be generated. For example, audio 102 may include audio of a guitar performance (or any other instrumental performance or vocal performance). Text 104 may be input into machine learning model 101 (e.g., received by the machine learning model). Text 104 may specify one or more instruments of the target music to be generated. For example, text 104 may specify that the target music should include audio of a piano performance (or any other instrumental performance or vocal performance). Additionally or alternatively, text 104 may specify one or more of the genre, tempo, style, melody, rhythm, pitch, or any other characteristics of the desired target music.
[0024] The machine learning model 101 can utilize the input audio 102 and the input text 104 to generate target music that features one or more instruments specified by the input text 104 and is aligned with (e.g., corresponds to) the input audio 102. For example, if the input audio 102 includes audio of a guitar performance, and the input text 104 specifies that the target music should include audio of a piano performance, the machine learning model 101 can utilize the input audio 102 and the input text 104 to generate target music that features a piano performance and is aligned with (e.g., corresponds to) a guitar performance.
[0025] The machine learning model 101 can generate a representation of the input audio 102. The representation of the input audio 102 can include a token 104. The token 104 can be, for example, a string of integer numbers. The representation of the input audio 102 can be generated using a codec 103. The codec 103 can be configured to generate a representation of the input audio 102 by compressing the input audio 102 into the token 104. The machine learning model 101 can generate a representation of the input text 104. The representation of the input text 104 can include a token 106. The token 106 can be, for example, a string of integer numbers. The representation of the input text 104 can be generated using an encoder 105. The encoder 105 can be configured to generate a representation of the input text 104 by compressing the input text 104 into the token 106.
[0026] The machine learning model can generate target music that features one or more instruments specified by the input text 104 and is aligned with the input audio 102 (e.g., corresponding to the input audio) using an iterative process. In an example, for a first iteration, a representation of the input audio 102 and a representation of the input text 104 can be input into a symbol combiner 108 (e.g., received by the symbol combiner). The symbol combiner 108 can be configured to combine information from multiple channels (e.g., a representation of the input audio 102 and a representation of the input text 104) to generate a continuous embedding (e.g., a set of floating point numbers). Figure 2 The symbol combiner 108 is discussed in more detail. The continuous embedding is input into (e.g., received by) the transformer submodel 110. The transformer submodel 110 can receive the continuous embedding from the symbol combiner 108. The transformer submodel 110 can generate a set of initial logical values (logits) 112 based on (e.g., using) the continuous embedding. The logical values 112 can include predictions of what the machine learning model 101 determined at each point in the target music being generated.
[0027] The set of initial logic values 112 can be sampled by adopting a sampling mechanism 114. The sampling mechanism 114 can be, for example, a multi-source CFG mechanism. CFG is a mechanism that allows amplification of the effect (e.g., influence) of certain external condition information on the network output. However, unlike the existing CFG mechanism that only allows amplification of the effect (e.g., influence) of one type of external condition information on the network output, the multi-source CFG mechanism for sampling the logic value 112 allows amplification of the effect (e.g., influence) of multiple types of external condition information on the network output. For example, the multi-source CFG mechanism can be configured to weight the impact of the input audio 102 and the input text 104 on the target music being generated, respectively. The following formula can be used to independently weight the guidance of the input audio 102 and the input text 104:
[0028]
[0029] Among them, c i is a separate conditional source, and λ i is the bootstrap coefficient for each conditional source. This derivation follows the derivation of single-source classifier-free bootstrapping via application of the Bayesian criterion. To support this multi-source classifier-free bootstrapping, a model with independent dropout needs to be trained for each conditional source.
[0030] A multi-source CFG can be applied to two conditional sources simultaneously using the following formula:
[0031] log(p(t)p(c|t)λ≈λlog(p(t|c))+(1-λ)logp(t))
[0032] where c is the combined conditioning source including both the background mix and any other conditioning, and λ is the bootstrap coefficient.
[0033] Depending on the individual weights assigned to the input audio 102 and the input text 104, the strength of the influence of the input audio 102 or the strength of the influence of the input text 104 can be used to dominate the generation of the target music. For example, if the weight assigned to the input audio 102 is greater than the weight assigned to the input text 104, the input audio 102 can dominate the generation of the target music. If the input audio 102 dominates the generation of the target music, the target music will be closely aligned with the input audio 102, but may or may not feature the instruments specified in the input text 104. Conversely, if the weight assigned to the input text 104 is greater than the weight assigned to the input audio 102, the input text 104 can dominate the generation of the target music. If the input text 104 dominates the generation of the target music, the target music will feature the instruments specified in the input text 104, but may or may not be well aligned with the input audio 102.
[0034] The sampled logic values may be sorted using a sorting submodel 116. The sorting submodel 116 may sort the sampled logic values using causal bias during iterative decoding. The sorting submodel 116 may sort the sampled logic values using one or more criteria. The sorted and sampled logic values may be used to generate an initial representation of the target music. The initial representation of the target music may include a symbol 118. The symbol 118 may be, for example, a string of integer numbers.
[0035] The criteria for sorting the sampled logical values have a great impact on the quality of the generated output. The first criterion can sort the sampled logical values by the confidence of the machine learning model 101, as indicated by the confidence value of each sampled logical value. The second criterion can randomly select the sampled logical values, such as by adding Gumbel noise to the confidence ranking with defined weights. However, the generation biased towards confidence will result in monotonous and boring output, while over-reliance on random selection will result in poor output transients and unnatural amplitude fluctuations. Therefore, a third criterion that encourages sampling of earlier sequence elements first can be used to sort the sampled logical values. Using the third criterion to sort the sampled logical values can strengthen a fuzzy causal relationship. The third criterion can be used alone or in combination with the first criterion and / or the second criterion to sort the sampled logical values. For example, the following sorting function combining the first criterion, the second criterion and the third criterion can be used:
[0036] ρ(x n )=w c ·c(x n )+w s (1-n / N)+wr X
[0037] Among them, x n is the sampled logical value at sequence index n, N is the total sequence length, c(x n ) is the confidence of the model on the sampled logical value calculated by applying softmax to the logical value, X~U(0,1) is a uniformly distributed random variable, and w c 、w s and w r is a scalar weight.
[0038] For a second iteration, a representation of the input audio 102, a representation of the input text 104, and an initial representation of the target music may be input into a symbol combiner 108 (e.g., received by the symbol combiner). The symbol combiner 108 may be configured to combine multiple channels of audio information (e.g., a representation of the input audio 102, a representation of the input text 104, and an initial representation of the target music) to generate a continuous embedding (e.g., a set of floating point numbers). The continuous embedding is input into a transformer submodel 110 (e.g., received by the transformer submodel). The transformer submodel 110 may receive the continuous embedding from the symbol combiner 108. The transformer submodel 110 may generate an updated set of (e.g., second) logical values 112 based on (e.g., using) the continuous embedding. The set of second logical values 112 may include predictions of the symbols determined by the machine learning model 101 at each point of the target music being generated.
[0039] The set of second logic values 112 can be sampled by adopting a sampling mechanism 114. As described above, the sampling mechanism 114 can be, for example, a multi-source CFG mechanism. The multi-source CFG mechanism can be configured to weight the impact of the input audio 102 and the input text 104 on the target music being generated, respectively. The sampled logic values from the set of second logic values 112 can be sorted. The sampled logic values from the set of second logic values 112 can be sorted using causal deviations during iterative decoding. For example, sorting the sampled logic values from the set of second logic values 112 using causal deviations during iterative decoding can include sorting the sampled logic values from the set of second logic values 112 using one or more criteria described above (e.g., confidence, random selection, and / or earlier sequence elements). The sorted and sampled logic values can be used to generate an updated (e.g., second) representation of the target music. For example, the second representation of the target music can include an updated symbol 118.
[0040] During a third iteration, a representation of the input audio 102, a representation of the input text 104, and a second representation of the target music may be input into a symbol combiner 108 (e.g., received by the symbol combiner). The symbol combiner 108 may be configured to combine multiple channels of audio information (e.g., a representation of the input audio 102, a representation of the input text 104, and a second representation of the target music) to generate a continuous embedding (e.g., a set of floating point numbers). The continuous embedding is input into a transformer submodel 110 (e.g., received by the transformer submodel). The transformer submodel 110 may receive the continuous embedding from the symbol combiner 108. The transformer submodel 110 may generate a further updated set of (e.g., third) logical values 112 based on (e.g., using) the continuous embedding. The set of third logical values 112 may include predictions of the symbols considered by the machine learning model 101 at each point of the target music being generated.
[0041] The set of the third logic value 112 can be sampled by adopting a sampling mechanism 114. As described above, the sampling mechanism 114 can be, for example, a multi-source CFG mechanism. The multi-source CFG mechanism can be configured to weight the impact of the input audio 102 and the input text 104 on the target music being generated, respectively. The sampled logic values from the set of the third logic value 112 can be sorted. The sampled logic values from the set of the third logic value 112 can be sorted using causal deviations during iterative decoding. For example, sorting the sampled logic values from the set of the third logic value 112 using causal deviations during iterative decoding can include sorting the sampled logic values from the set of the third logic value 112 using one or more criteria described above (e.g., confidence, random selection and / or earlier sequence elements). The sorted and sampled logic values can be used to generate a further updated (e.g., third) representation of the target music. For example, the third representation of the target music can include further updated symbols 118.
[0042] Any number of additional iterations may be performed. For example, five, ten, twenty, etc. iterations may be performed until a final representation of the target music (e.g., a set of final symbols 118) is generated. The target music may be constructed (e.g., synthesized, generated, etc.) based on the final representation of the target music. For example, the target music may be constructed by inputting the set of final symbols 118 into a decoder that receives the set of final symbols 118 and uses the set of final symbols to construct an audio waveform associated with the target music.
[0043] Figure 2The symbol combiner 108 is shown in more detail. As described above, the symbol combiner 108 can be configured to combine the audio information of multiple channels to generate a continuous embedding (e.g., a set of floating point numbers). The symbol combiner 108 can receive the symbol 104 associated with the audio 102 as input. The symbol combiner 108 can receive the symbol 106 associated with the text 104 as input. The symbol combiner 108 can receive the symbol 118 associated with the target music as input. The symbol combiner 108 may include at least one symbol layout component. The at least one symbol layout component can convert the symbol 104 into a first embedding 202. The first embedding 202 can be an embedding representing the audio input. The at least one symbol layout component can convert the symbol 118 into a second embedding 204. The second embedding 204 can be an embedding representing the target audio. The symbol combiner 108 may include at least one text embedding layer 206. The at least one symbol layout component can convert the symbol 106 into a third embedding 206. The third embedding 206 can be an embedding representing the text 104 input.
[0044] The symbol combiner 108 can be configured to add together (e.g., add together) the first embedding 202 and the third embedding 206. The sum of the first embedding 202 and the third embedding 206 can be an audio-text embedding 208. The symbol combiner 108 can be configured to add together (e.g., add together) the second embedding 204 and the third embedding 206. The sum of the second embedding 204 and the third embedding 206 can be a target-text embedding 210. The symbol combiner 108 can include a splice 112. The splice 112 can be configured to splice or append the audio-text embedding 208 and the target-text embedding 210 together to generate a continuous embedding 214.
[0045] Figure 3 An example system 300 for training a machine learning model 101 to generate target music according to the present disclosure is shown. The machine learning model 101 can be trained using a masking process on training data pairs. Each training data pair can include background audio and ground truth target audio. The masking process can include masking a portion of the ground truth target audio in any particular training data pair. The machine learning model 101 can be trained to generate a masked portion of the target audio based on the background audio in the same particular training data pair.
[0046] In an embodiment, a particular training data pair may include background audio 302 and ground truth target audio 315. Background audio 302 and ground truth target audio 315 may be fed into at least one codec 103. Background audio 302 and ground truth target audio 315 may be fed into the same codec or different codecs. (Multiple) codecs may generate a representation of background audio 302. The representation of background audio 302 may include symbol 304. Symbol 304 may be, for example, a string of integer numbers. (Multiple) codecs may generate a representation of ground truth target audio 315. The representation of ground truth target audio 315 may include symbol 318. Symbol 318 may be, for example, a string of integer numbers. Symbol 318 may be input into mask sub-model 324. Mask sub-model 324 may generate a masked target symbol by masking or removing a portion of symbol 318.
[0047] The machine learning model 101 can generate a representation of text 304. The text 304 can specify one or more instruments of the ground truth target audio. Additionally or alternatively, the text 304 can specify one or more of the genre, rhythm, style, melody, rhythm, pitch, or any other characteristics of the ground truth target audio. The representation of text 304 may include symbols 306. Symbols 306 may be, for example, a string of integer numbers. The representation of the input text 304 can be generated using encoder 105. Encoder 105 can be configured to generate a representation of text 304 by compressing text 304 into symbols 306.
[0048] The symbol combiner 108 may receive as input the symbol element 304, the masked target symbol, and the symbol 306. The symbol combiner 108 may be configured to combine the symbol 304, the masked target symbol, and the symbol 306 to generate a continuous embedding (e.g., a set of floating point numbers). The continuous embedding is input into the transformer submodel 110 (e.g., received by the transformer submodel). The transformer submodel 110 may receive the continuous embedding from the symbol combiner 108. The transformer submodel 110 may generate a logical value 312 based on (e.g., using) the continuous embedding. The logical value 312 may include a prediction of the symbol considered by the machine learning model 101 for each point in the masked portion of the symbol 318. The logical value 312 may be sent to a masked language model (MLM) loss submodel 322. The masked language model (MLM) loss submodel 322 may compare the logical value 312 with the ground truth target audio to determine the loss associated with the logical value 312. The loss may indicate how bad (eg, how different from the ground truth) the prediction of the logical value 312 is. This process may be repeated for multiple iterations until the loss is below a desired threshold.
[0049] Figure 4An example system 400 for generating training data pairs is shown, on which the machine learning model 101 is trained. As described above, each training data pair includes background audio and target audio. The training data pairs can be generated using multiple music clips 402. Each music clip 402 in the multiple music clips 402 can be divided into M sub-tracks, where M represents a positive integer. Each sub-track in the M sub-tracks can correspond to the performance of a musician, a specific instrument, or something more abstract (such as the output of a sampler or synthesizer).
[0050] The background audio in each training data pair can be generated by randomly selecting N tracks from the M tracks of one of the multiple music clips 402. N represents a positive integer less than M. The target audio in the same training data pair can be generated by randomly selecting the remaining tracks from the M tracks of the same music clip. For example, the generation system 405 can use the first music clip to generate a first training data pair. The first music clip can be divided into five tracks. The generation system 405 can randomly select a mix 404 of one, two, three or four tracks from the five tracks. The mix 404 of the tracks can be the background audio in the first training data pair. The generation system 405 can randomly select one or more remaining tracks 406 (for example, one or more tracks that are not randomly selected to generate background audio). The (multiple) remaining tracks 406 can be the target audio in the first training data pair. This process can be repeated multiple times to generate multiple training data pairs. The machine learning model 101 can be trained on multiple training data pairs, such as using a reference. Figure 3 Describe the masking process.
[0051] Figure 5 An example process 500 for generating target music is shown. Figure 5 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0052] At 502, audio and text can be input to a machine learning model. The input audio can indicate the context of the target music to be generated. For example, the input audio can include audio of a guitar performance (or any other instrumental performance or vocal performance). The input text can specify one or more instruments of the target music to be generated. For example, the input text can specify that the target music should include audio of a piano performance (or any other instrumental performance or vocal performance). Additionally or alternatively, the input text can specify one or more of the genre, rhythm, style, melody, rhythm, pitch, or any other characteristics of the desired target music.
[0053] The machine learning model can utilize the input audio and the input text to generate target music characterized by one or more instruments specified by the input text and aligned with the input audio (e.g., corresponding to the input audio). At 504, a representation of the target music can be generated. The representation of the target music can be generated by the machine learning model through an iterative process. The representation of the target music can include symbols corresponding to the target music (e.g., a string of integer numbers). The representation of the target music can be used to generate (e.g., construct, synthesize) the target music. The generated target music can be characterized by one or more instruments specified by the input text and aligned with the input audio. For example, if the input audio includes audio of a guitar performance, and the input text specifies that the target music should include audio of a piano performance, the machine learning model can utilize the input audio and the input text to generate target music characterized by a piano performance and aligned with a guitar performance (e.g., corresponding to a guitar performance).
[0054] Figure 6 An example process 600 for generating target music is shown. Figure 6 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0055] For each of multiple iterations, the symbol combiner can be configured to combine audio information of multiple channels (e.g., one or more of a representation of input audio, a representation of input text, and a representation of target audio) to generate a continuous embedding (e.g., a set of floating point numbers). The continuous embedding is input into a transformer sub-model (e.g., received by the transformer sub-model). The transformer sub-model can receive the continuous embedding from the symbol combiner. The transformer sub-model can generate a logical value based on (e.g., using) the continuous embedding. The logical value can include a prediction of the symbol that the machine learning model believes at each point in the target music being generated.
[0056] At 602, a multi-source classifier-free guidance mechanism can be used to sample the logic value output from each iteration. The multi-source classifier-free guidance mechanism can be configured to weight the impact of the input audio and the input text respectively. At 604, the sampled logic values can be sorted. The sampled logic values can be sorted at least in part based on the timeline of the target music. For example, one or more criteria can be used to sort the sampled logic values. The first criterion can sort the sampled logic values by the confidence of the machine learning model, as indicated by the confidence value of each sampled logic value. The second criterion can randomly select the sampled logic values, such as by adding Gumbel noise to the confidence ranking with a defined weight. The third criterion can sort the sampled logic values based on the timeline of the target music. For example, the third criterion can encourage sampling of earlier sequence elements first. Using the third criterion to sort the sampled logic values can strengthen a fuzzy causal relationship. The third criterion can be used alone or in combination with the first criterion and / or the second criterion to sort the sampled logic values.
[0057] The sorted and sampled logical values may be used to generate and / or update a representation of the target music. The representation of the target music may include symbols (e.g., a string of integer numbers). At 606, the representation of the target music may be updated based on the sorted sampled logical values. The representation of the target music may be used to generate (e.g., construct, synthesize) the target music. The generated target music may feature one or more instruments specified by the input text and be aligned with the input audio.
[0058] Figure 7 An example process 700 for generating target music is shown. Figure 7 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0059] At 702, audio and text can be input to a machine learning model. The input audio can indicate the context of the target music to be generated. For example, the input audio can include audio of a guitar performance (or any other instrumental performance or vocal performance). The input text can specify one or more instruments of the target music to be generated. For example, the input text can specify that the target music should include audio of a piano performance (or any other instrumental performance or vocal performance). Additionally or alternatively, the input text can specify one or more of the genre, rhythm, style, melody, rhythm, pitch, or any other characteristics of the desired target music.
[0060] The machine learning model can utilize the input audio and the input text to generate target music featuring one or more instruments specified by the input text and aligned with the input audio (e.g., corresponding to the input audio). At 704, a representation of the target music can be generated. The representation of the target music can be generated by the machine learning model through an iterative process. The representation of the target music can include symbols corresponding to the target music (e.g., a string of integer numbers).
[0061] The representation of the target music can be used to generate (e.g., construct, synthesize) the target music. At 706, the target music can be constructed. The target music can be constructed based on the representation of the target music. The target music can be characterized by one or more instruments specified by the input text and can be aligned with the input audio. For example, if the input audio includes audio of a guitar performance, and the input text specifies that the target music should include audio of a piano performance, the machine learning model can utilize the input audio and the input text to generate target music characterized by a piano performance and aligned with (e.g., corresponding to) a guitar performance.
[0062] Figure 8 An example process 800 for generating target music is shown. Figure 8 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0063] A representation of the input audio and a representation of the input text can be input into a symbol combiner (e.g., received by the symbol combiner). The symbol combiner can be configured to combine the representation of the input audio and the representation of the input text to generate a continuous embedding (e.g., a set of floating point numbers). The continuous embedding can be input into a transformer submodel (e.g., received by the transformer submodel). At 802, an initial set of logical values can be generated. The initial set of logical values can be generated by the transformer submodel. The initial set of logical values can be generated based on the representation of the input audio and the representation of the input text. For example, the initial set of logical values can be generated based on the continuous embedding. The initial set of logical values can include predictions of symbols in each point in the target music being generated.
[0064] At 804, an initial representation of the target music can be generated. The initial representation of the target music can be generated based on sampling and sorting the initial logic value set. The initial logic value set can be sampled by adopting a multi-source CFG mechanism. Unlike the existing CFG mechanism that only allows amplification of the effect (e.g., influence) of one type of external condition information on the network output, the multi-source CFG mechanism for sampling the initial logic value set allows amplification of the effect (e.g., influence) of multiple types of external condition information on the network output. For example, the multi-source CFG mechanism can be configured to weight the impact of the input audio and the input text on the target music being generated, respectively. The multi-source CFG can be applied to two condition sources simultaneously.
[0065] One or more criteria can be used to sort the sampled logical values. The first criterion can sort the sampled logical values by the confidence of the machine learning model, as indicated by the confidence value of each sampled logical value. The second criterion can randomly select the sampled logical values, such as by adding Gumbel noise to the confidence ranking with defined weights. The third criterion can sort the sampled logical values based on the timeline of the target music. For example, the third criterion can encourage sampling of earlier sequence elements first. Using the third criterion to sort the sampled logical values can strengthen a fuzzy causal relationship. The third criterion can be used alone or in combination with the first criterion and / or the second criterion to sort the sampled logical values. The sorted and sampled logical values can be used to generate an initial representation of the target music. The initial representation of the target music can include symbols. The symbol can be, for example, a string of integer numbers.
[0066] The representation of the input audio, the representation of the input text and the initial representation of the target music can be input into the symbol combiner. The symbol combiner can be configured to combine the representation of the input audio, the representation of the input text and the initial representation of the target music to generate a continuous embedding (e.g., a set of floating point numbers). The continuous embedding can be input into the transformer submodel 110 (e.g., received by the transformer submodel). The transformer submodel 110 can receive the continuous embedding from the symbol combiner. At 806, a second set of logical values can be generated. The second set of logical values can be generated based on the representation of the input audio, the representation of the input text and the initial representation of the target music. For example, the transformer submodel can generate the second set of logical values based on (e.g., using) continuous embedding. The second set of logical values can include an updated prediction of the symbol considered by the machine learning model at each point in the target music being generated.
[0067] At 808, an updated representation of the target music can be generated. The updated representation of the target music can be generated based on sampling and sorting the second logic set values. The second logic value set can be sampled by adopting a multi-source CFG mechanism. The sampled logic values from the second logic value set can be sorted. The sampled logic values from the second logic value set can be sorted using one or more criteria described above (e.g., confidence, random selection, and / or earlier sequence elements). The sorted and sampled logic values can be used to generate an updated (e.g., second) representation of the target music. For example, the second representation of the target music may include updated symbols.
[0068] Fig. 9 An example process 900 for generating an audio-text embedding is illustrated. Fig. 9 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0069] The symbol combiner may be configured to combine the audio information of multiple channels to generate a continuous embedding (e.g., a set of floating point numbers). The symbol combiner may receive symbols associated with the audio of the input as input. At 902, an embedding representing the audio of the input may be generated. For example, the symbol combiner may include at least one symbol layout component. At least one symbol layout component may convert symbols associated with the audio of the input into an embedding representing the audio of the input. The symbol combiner may receive symbols associated with the text of the input as input. At 904, an embedding representing the text of the input may be generated. For example, at least one symbol layout component may convert symbols associated with the text of the input into an embedding representing the text of the input. The symbol combiner may be configured to add together (e.g., add together) the embedding representing the audio of the input and the embedding representing the text of the input. At 906, the embedding representing the audio of the input and the embedding representing the text of the input may be summed to generate an audio-text embedding.
[0070] Fig.10 An example process 1000 for generating a target-text embedding is illustrated. Fig.10 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0071] The symbol combiner may be configured to combine the audio information of multiple channels to generate a continuous embedding (e.g., a set of floating point numbers). The symbol combiner may receive symbols associated with the text of the input as input. At 1002, an embedding representing the text of the input may be generated. For example, the symbol combiner may include at least one symbol layout component, and the at least one symbol layout component may be configured to convert symbols associated with the text of the input into an embedding representing the text of the input. The symbol combiner may receive symbols associated with the target music as input. At 1002, an embedding representing the target music may be generated. For example, the symbol combiner may include at least one symbol layout component. At least one symbol layout component may convert symbols associated with the target music into an embedding representing the target music. The symbol combiner may be configured to add together (e.g., add together) the embedding representing the target music and the embedding representing the text of the input. At 1006, the embedding representing the target music and the embedding representing the text of the input may be summed to generate a target-text embedding.
[0072] Fig.11 An example process 1100 for generating continuous embeddings and generating logical values according to the present disclosure is illustrated. Fig.11 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0073] The symbol combiner can be configured to add together (e.g., add together) an embedding representing background audio and an embedding representing input text to generate a target-text embedding. The symbol combiner can be configured to add together (e.g., add together) an embedding representing target music and an embedding representing input text to generate a target-text embedding. At 1102, the audio-text embedding and the target-text embedding can be spliced. The audio-text embedding and the target-text embedding can be spliced to generate a continuous embedding. At 1104, the continuous embedding can be input into a sub-model of a machine learning model. At 1106, a logical value can be generated by a sub-model of the machine learning model. A logical value can be generated by the sub-model based on the continuous embedding. The logical value can include a prediction of the symbol determined by the machine learning model at each point in the target music being generated.
[0074] Fig.12 An example process 1200 for training a machine learning model to generate target music according to the present disclosure is illustrated. Fig.12 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0075] At 1202, training data pairs can be generated. The training data pairs can be generated using music clips. Each music clip can be divided into M tracks, where M represents a positive integer. Each training data pair can include background audio and target audio. For example, the background audio in each pair can include a randomly selected subset N of the M tracks associated with a particular music clip. The background audio in each pair can include at least one track randomly selected from the remaining MN tracks. At 1204, a machine learning model can be trained on the training data pairs. The machine learning model can be trained to generate target music based on input text and input background audio.
[0076] Fig.13 An example process 1300 for generating training data pairs is illustrated. Fig.13 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0077] As described above, each training data pair includes background audio and target audio. The training data pair can be generated using multiple music clips. Each music clip in the multiple music clips can be divided into M sub-tracks, where M represents a positive integer. Each sub-track in the M sub-tracks can correspond to the performance of a musician, a specific musical instrument, or something more abstract (such as the output of a sampler or synthesizer). At 1302, the background audio in each training data pair can be generated. The background audio in each training data pair can be generated by randomly selecting N sub-tracks from the M sub-tracks of the music clip, where N represents a positive integer less than M. At 1304, the target audio in each training data pair can be generated. The target audio in each training data pair can be generated by randomly selecting the remaining sub-tracks from the M sub-tracks of the same music clip.
[0078] Fig.14 An example process 1400 for training a machine learning model to generate target music is illustrated. Fig.14 Although depicted as a series of operations, those skilled in the art will appreciate that various embodiments may add, remove, reorder, or modify the depicted operations.
[0079] A machine learning model can be trained using a masking process on training data pairs to generate target music. Each training data pair can include background audio and ground truth target audio. A representation of the background audio in any particular training data pair can be generated. The representation of the background audio can include symbols (e.g., a string of integer numbers). A representation of the target audio in the same particular training data pair can be generated. The representation of the target audio can include symbols (e.g., a string of integer numbers). The representation of the target audio can be input into a mask sub-model. At 1402, a portion of the target audio can be masked. For example, the mask sub-model can generate a masked representation (e.g., a masked target symbol) by masking a portion of the target audio.
[0080] At 1404, a machine learning model can be trained to generate a masked portion of the target audio. The machine learning model can be trained to generate a masked portion of the target audio based at least in part on the background audio in the same specific training data pair. For example, a representation of text can be generated. The text can specify one or more instruments of the ground truth target audio. Additionally or alternatively, the text can specify one or more of the genre, rhythm, style, melody, rhythm, pitch, or any other characteristics of the ground truth target audio. A logical value can be generated using a representation of the text, a representation of the background audio, and a masked representation of the target audio. The logical value can include a prediction of the content determined by the machine learning model at each point of the masked portion of the target audio. The masked language model (MLM) loss submodel can compare the logical value with the ground truth target audio to determine the loss associated with the logical value. The loss can indicate how bad the prediction of the logical value is (e.g., how big the difference from the ground truth is). The process can be repeated for multiple iterations until the loss is below a desired threshold.
[0081] The machine learning model 101 was evaluated. The machine learning model 101 trained on a first dataset consisting of 145 hours of synthesized music audio divided into multiple tracks was evaluated. Additionally, the machine learning model 101 trained on a second dataset consisting of 500 hours of authorized human-performed music divided into multiple tracks was evaluated. For the evaluation, the model was trained according to the same Figure 4A similar process to the one shown for constructing training examples produces a set of 400 example outputs, but this set of example outputs is extracted from a separate set of test data that is not included in the training set. For each example, a background mix is generated and a target track category is randomly selected. The generated tracks produced by the model for the conditional set are placed in one group. The real tracks that match the target category are placed in another group. The two groups are then compared using a variety of metrics. This process ensures that no systematic errors are introduced due to an imbalance in the target categories between the two groups. FAD-type metrics are calculated between two isolated groups of tracks, while MIRDD is calculated on a modified group consisting of the track plus its original background mix. This allows the MIRDD metric to better penalize poor musical alignment and coherence.
[0082] To evaluate the performance of causal bias during iterative decoding, two ablations were performed. For these ablations, the model trained on the first dataset was used. First, by computing the causal bias with two conditional sources c a and c i The various guiding coefficients λ a and λ i The results can be found in Fig.15A Table 1500 shows the average objective index for different ranges of guidance coefficients. A guidance coefficient value of 1.0 is equivalent to no guidance with respect to the condition source. The exact optimal value of the guidance coefficient depends largely on the condition information, so for cases where the guidance coefficient is greater than 1.0, the optimal value of the guidance coefficient is obtained at different values of λ. i , a The average results of 1000 examples are shown below (up to a maximum value of 4.0). The results confirm that adding multi-source guidance is beneficial. a =λ i = 3.0 was determined as the general setting for evaluation.
[0083] Second, we test the impact of causal biases during iterative decoding by comparing the relative strengths of the causal biases. c and w r are set to 0.1 and 1.0 respectively. The results can be seen in Fig. 15B Table 1501 shows the causal bias weight w s Table 1501 shows that adding a small amount of causal deviation has a positive impact on the FAD and MIRDD indicators, indicating that the sound quality and music alignment are improved. s As w increases further, this effect weakens. s= 0.1 for further experiments. During iterative decoding, 128, 64, 32, and 32 steps were used for the four levels of the tokenizer, respectively.
[0084] The model trained on the first dataset and the model trained on the second dataset are evaluated. Both models are sampled using the sampling technique described in this article. In addition, each model is sampled using the original sampled parameter set, which is equivalent to eliminating the classifier-free bootstrapping and causal bias in decoding. Fig.16A The objective metrics for these output sets are shown in Table 1600 of . Table 1600 shows the objective evaluation metrics for both models with the optimal sampling parameters and the original sampling parameters. While a direct comparison cannot be made due to the different tasks, the FAD scores for the models trained on both the first and second datasets are comparable to those seen on state-of-the-art text-conditioned models. It is also clear that the second dataset is larger in scale and contains more human content, so the output quality is significantly improved.
[0085] A Mean Opinion Score (MOS) test was also performed by asking ten music-trained participants to verify the subjective quality of the generated model. Three output sets were constructed by mixing the generated tracks or the true tracks with their corresponding background mixes. The generated tracks were taken from the original and best output sets of the model trained on the second data set as evaluated in Table 1600. The true tracks were taken from the reference set used in the previous evaluation. A set of 60 mixes was collated using this technique (evenly divided into original, best and true), and the listeners were asked to rate the overall quality on a scale from very poor (1) to very good (5). The results are shown in Fig. 16B The results shown in Table 1601 confirm that the proposed model is able to create reasonable musical results.
[0086] Fig.17 The diagram shows a computing device that can be used in various aspects, such as Figures 1 to 4 Any of the services, networks, modules and / or devices depicted in any of the. Figures 1 to 4 , any or all components can be Fig.17 It is implemented by one or more instances of computing device 1700. Fig.17 The computer architecture shown in the diagram illustrates a conventional server computer, workstation, desktop computer, laptop computer, tablet computer, network appliance, PDA, e-reader, digital cellular telephone, or other computing node, and can be used to perform any aspects of the computer described herein, such as implementing the methods described herein.
[0087] The computing device 1700 may include a baseboard or "motherboard," which is a printed circuit board to which multiple components or devices may be connected via a system bus or other electrical communication paths. One or more central processing units (CPUs) 1704 may operate in conjunction with a chipset 1706. The CPU(s) 1704 may be a standard programmable processor that performs the arithmetic and logic operations required for the operation of the computing device 1700.
[0088] (Multiple) CPUs 1704 can perform necessary operations by manipulating switching elements from one discrete physical state to the next, which can distinguish and change these states. Switching elements can generally include electronic circuits that maintain one of two binary states (such as flip-flops), and electronic circuits that provide output states based on logical combinations of the states of one or more other switching elements (such as logic gates). These basic switching elements can be combined to create more complex logic circuits, including registers, adder-subtractors, arithmetic logic units, floating point units, etc.
[0089] CPU(s) 1704 may be enhanced or replaced with other processing units, such as GPU(s) 1705. GPU(s) 1705 may include processing units specialized for, but not necessarily limited to, highly parallel computing, such as graphics and other visualization-related processing.
[0090] Chipset 1706 may provide an interface between CPU(s) 1704 and the rest of the components and devices on the baseboard. Chipset 1706 may provide an interface to random access memory (RAM) 1708, which is used as main memory in computing device 1700. Chipset 1706 may further provide an interface to a computer-readable storage medium, such as read-only memory (ROM) 1720 or non-volatile RAM (NVRAM) (not shown), for storing basic routines that may help boot up computing device 1700 and transfer information between various components and devices. ROM 1720 or NVRAM may also store other software components required for operation of computing device 1700 according to various aspects described herein.
[0091] The computing device 1700 may operate in a networked environment using logical connections to remote computing nodes and computer systems through a local area network (LAN). Chipset 1706 may include functionality for providing network connectivity through a network interface controller (NIC) 1722, such as a Gigabit Ethernet adapter. NIC 1722 may be capable of connecting the computing device 1700 to other computing nodes through network 1716. It should be understood that multiple NICs 1722 may be present in the computing device 1700, thereby connecting the computing device to other types of networks and remote computer systems.
[0092] The computing device 1700 may be connected to a mass storage device 1728 that provides non-volatile storage for the computer. The mass storage device 1728 may store system programs, application programs, other program modules, and data, all of which have been described in more detail herein. The mass storage device 1728 may be connected to the computing device 1700 via a storage controller 1724 connected to the chipset 1706. The mass storage device 1728 may be comprised of one or more physical storage units. The mass storage device 1728 may include a management component 1710. The storage controller 1724 may interface with the physical storage unit via a serial attached SCSI (SAS) interface, a serial advanced technology attachment (SATA) interface, a fiber channel (FC) interface, or other types of interfaces for physically connecting and transferring data between a computer and the physical storage unit.
[0093] The computing device 1700 may store data on the mass storage device 1728 by transforming the physical state of the physical storage unit to reflect the stored information. The specific transformation of the physical state may depend on various factors and different implementations of this specification. Examples of such factors may include, but are not limited to, the technology used to implement the physical storage unit and whether the mass storage device 1728 is characterized as a primary storage device or an auxiliary storage device.
[0094] For example, the computing device 1700 can issue instructions through the storage controller 1724 to change the magnetic properties of specific locations in the disk drive unit, the reflective or refractive properties of specific locations in the optical storage unit, or the electrical properties of specific capacitors, transistors or other discrete components in the solid-state storage unit to store information in the mass storage device 1728. Other transformations of the physical medium are also possible without departing from the scope and spirit of the present specification, wherein the foregoing examples are provided only to facilitate the present specification. The computing device 1700 can further read information from the mass storage device 1728 by detecting the physical state or characteristics of one or more specific locations in the physical storage unit.
[0095] In addition to the mass storage device 1728 described above, the computing device 1700 may also access other computer-readable storage media to store and retrieve information, such as program modules, data structures, or other data. It should be understood by those skilled in the art that computer-readable storage media may be any available media that provides non-transitory data storage and can be accessed by the computing device 1700.
[0096] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile computer-readable storage media, transient computer-readable storage media and non-transient computer-readable storage media, and removable and non-removable media implemented in any method or technology. Computer-readable storage media include, but are not limited to, RAM, ROM, erasable programmable ROM ("EPROM"), electrically erasable programmable ROM ("EEPROM"), flash memory or other solid-state memory technology, compact disk ROM ("CD-ROM"), digital versatile disk ("DVD"), high-definition DVD ("HD-DVD"), BLU-RAY, or other optical storage devices, cassettes, magnetic tape, magnetic disk storage devices, other magnetic storage devices, or any other medium that can be used to store desired information in a non-transient manner.
[0097] Mass storage devices (such as Fig.17 The mass storage device 1728 (depicted in FIG. 17 ) may store an operating system for controlling the operation of the computing device 1700. The mass storage device 1728 may store other system or application programs and data used by the computing device 1700.
[0098] The mass storage device 1728 or other computer-readable storage medium may also be encoded with computer-executable instructions that, when loaded into the computing device 1700, transform the computing device from a general-purpose computing system into a special-purpose computer capable of implementing the various aspects described herein. These computer-executable instructions transform the computing device 1700 by specifying the manner in which the CPU(s) 1704 transition between states, as described above. The computing device 1700 may access a computer-readable storage medium storing computer-executable instructions that, when executed by the computing device 1700, may perform the methods described herein.
[0099] Computing devices (such as Fig.17The computing device 1700 depicted in FIG. 1 may also include an input / output controller 1732 for receiving and processing input from a plurality of input devices (e.g., a keyboard, a mouse, a touch pad, a touch screen, an electronic stylus) or other types of input devices. Similarly, the input / output controller 1732 may provide output to a display (e.g., a computer monitor, a flat panel display, a digital projector, a printer, a plotter) or other types of output devices. It should be understood that the computing device 1700 may not include Fig.17 All components shown may include Fig.17 Other components not explicitly shown in the figure or may be used in conjunction with Fig.17 A completely different architecture is shown in .
[0100] As described herein, a computing device may be a physical computing device, such as Fig.17 The computing device 1700 of the present invention may further include a virtual machine host process and one or more virtual machine instances. The computer executable instructions may be indirectly executed by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed in the context of the virtual machine.
[0101] It should be understood that the methods and systems are not limited to specific methods, specific components or specific implementations. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting.
[0102] Unless the context clearly dictates otherwise, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when a value is expressed as an approximation by use of the antecedent "about," it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each range are significant relative to the other endpoints, and independent of the other endpoints.
[0103] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not occur.
[0104] Throughout the description and claims of this specification, the word "comprise" and variations of the word (such as "comprising and comprises") mean "including but not limited to", and are not intended to exclude, for example, other components, integers or steps. "Exemplary" means "an example of..." and is not intended to convey an indication of a preferred or ideal embodiment. "Such as" is not used in a limiting sense, but for explanatory purposes.
[0105] Components that can be used to perform the described methods and systems are described. When combinations, subsets, interactions, groups, etc. of these components are described, it should be understood that, although specific references to each of the various individual and collective combinations and arrangements of these components may not be explicitly described, for all methods and systems, each component is specifically contemplated and described herein. This applies to all aspects of the application, including but not limited to the operations in the described methods. Therefore, if there are various additional operations that can be performed, it should be understood that each of these additional operations can be performed with any specific embodiment or combination of embodiments of the described methods.
[0106] The present method and system may be understood more readily by reference to the following detailed description of the preferred embodiments and examples included therein and to the accompanying drawings and their description.
[0107] As will be appreciated by those skilled in the art, the present method and system may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. In addition, the present method and system may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) embodied in the storage medium. More specifically, the present method and system may take the form of computer software implemented by a web. Any suitable computer-readable storage medium may be utilized, including a hard disk, a CD-ROM, an optical storage device, or a magnetic storage device.
[0108] Embodiments of the method and system are described below with reference to block diagrams and flowchart illustrations of methods, systems, devices, and computer program products. It should be understood that each block of the block diagrams and flowchart illustrations and combinations of blocks in the block diagrams and flowchart illustrations can be implemented by computer program instructions, respectively. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed on the computer or other programmable data processing device create means for implementing the functions specified in one or more blocks of the flowchart.
[0109] These computer program instructions may also be stored in a computer-readable memory, which may instruct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product, which includes computer-readable instructions for implementing the functions specified in one or more blocks of the flowchart. The computer program instructions may also be loaded onto a computer or other programmable data processing device to cause a series of operating steps to be performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more blocks of the flowchart.
[0110] The various features and processes described above can be used independently of each other, or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. In addition, in some embodiments, certain method or process blocks can be omitted. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated with the methods and processes can be performed in other appropriate orders. For example, the blocks or states described can be performed in an order different from that specifically described, or multiple blocks or states can be combined in a single block or state. The example blocks or states can be performed serially, in parallel, or in some other way. Blocks or states can be added to or removed from the described example embodiments. The example systems and components described herein can be configured differently from those described. For example, elements can be added, removed, or rearranged compared to the described example embodiments.
[0111] It should also be understood that the various items are illustrated as being stored in memory or storage when in use, and that for the purposes of memory management and data integrity, these items or portions thereof may be transferred between memory and other storage devices. Alternatively, in other embodiments, some or all of the software modules and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. In addition, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as at least partially implemented or provided in firmware and / or hardware, including but not limited to one or more application specific integrated circuits (“ASICs”), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field programmable gate arrays (“FPGAs”), complex programmable logic devices (“CPLDs”), etc. Some or all of the modules, systems, and data structures may also be stored (e.g., as software instructions or structured data) on a computer-readable medium, such as a hard disk, memory, network, or portable media product to be read by an appropriate device or via an appropriate connection. Systems, modules, and data structures may also be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) on various computer-readable transmission media (including wireless and wire / cable-based media), and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). In other embodiments, such computer program products may also take other forms. Therefore, the present invention may be practiced with other computer system configurations.
[0112] Although the method and system have been described in conjunction with preferred embodiments and specific examples, it is not intended to limit the scope to the specific embodiments set forth, as the embodiments herein are intended in all respects to be illustrative rather than restrictive.
[0113] Unless otherwise expressly stated, it is in no way intended that any method described herein be construed as requiring that its operations be performed in a particular order. Therefore, in the event that a method claim does not actually recite the order in which its operations are to be followed, or in the absence of other specific statements in the claims or specification that the operations are to be limited to a particular order, no order is intended to be inferred in any way. This applies to any possible non-express basis for interpretation, including: logical issues regarding the arrangement of steps or the flow of operations; direct meaning derived from grammatical organization or punctuation; and the number or type of embodiments described in the specification.
[0114] It will be apparent to those skilled in the art that various modifications and variations may be made without departing from the scope or spirit of the present disclosure. Other embodiments will be apparent to those skilled in the art in view of the specification and practice described herein. This specification and example drawings are to be considered exemplary only, with the true scope and spirit being indicated by the appended claims.
Claims
1. A method for generating target music using a machine learning model, comprising: inputting audio and text into the machine learning model, wherein the input audio indicates the background of the target music, and wherein the input text specifies one or more instruments of the target music; and A representation of the target music is generated by an iterative process, wherein the target music features the one or more instruments specified by the input text and is aligned with the input audio, and wherein the iterative process comprises a plurality of iterations, each of the plurality of iterations comprising the following: A multi-source classifier-free guidance mechanism is used to sample the logical value from the previous iteration output, wherein the multi-source classifier-free guidance mechanism is configured to weight the influence of the input audio and the input text respectively, ordering the sampled logic values based at least in part on a timeline of the target music, and The representation of the target music is updated based on the sorted, sampled logical values.
2. The method according to claim 1, further comprising: The target music is constructed based on the representation of the target music.
3. The method according to claim 1, further comprising: generating an initial set of logical values based on the input representation of the audio and the input representation of the text; as well as An initial representation of the target music is generated based on sampling and sorting the initial set of logical values.
4. The method according to claim 3, further comprising: A second set of logical values is generated based on the representation of the input audio, the representation of the input text, and the initial representation of the target music.
5. The method according to claim 4, further comprising: An updated representation of the target music is generated based on sampling and sequencing the second set of logical values.
6. The method according to claim 1, further comprising: generating an embedding representing the input audio; generating an embedding representing the text of the input; as well as The embedding representing the audio of the input and the embedding representing the text of the input are summed to generate an audio-text embedding.
7. The method according to claim 6, further comprising: generating an embedding corresponding to a representation of the target music; as well as The embedding corresponding to the representation of the target music and the embedding of the text representing the input are summed to generate a target-text embedding.
8. The method according to claim 7, further comprising: concatenating the audio-text embedding and the target-text embedding to generate a continuous embedding; Inputting the continuous embedding into a sub-model of the machine learning model; as well as The logical value is generated by the sub-model.
9. The method according to claim 1, wherein: The text specifies at least one of the genre, tempo, style, melody, rhythm, or pitch of the target music.
10. The method according to claim 1, further comprising: Training data pairs are generated using music clips, where each training data pair includes background audio and target audio, where each music clip is divided into M sub-tracks, and M represents a positive integer.
11. The method according to claim 10, further comprising: generating the background audio in each training data pair by randomly selecting N sub-tracks from the M sub-tracks of the music segment in the music segment, wherein N represents a positive integer less than M; and The target audio in each training data pair is generated by randomly selecting remaining tracks from the M tracks of the music segment in the music segment.
12. The method according to claim 10, further comprising: The machine learning model is trained using a masking process on the training data pairs, wherein the masking process comprises: masking a portion of the target audio in any particular training data pair, and The machine learning model is trained to generate the masked portion of the target audio based on the background audio in the same particular training data pair.
13. A system for generating target audio using a machine learning model, comprising: at least one processor; as well as at least one memory, the at least one memory being communicatively coupled to the at least one processor and comprising computer-readable instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: inputting audio and text into the machine learning model, wherein the input audio indicates the background of the target music, and wherein the input text specifies one or more instruments of the target music; and A representation of the target music is generated by an iterative process, wherein the target music features one or more instruments specified by the input text and is aligned with the input audio, and wherein the iterative process comprises a plurality of iterations, each of the plurality of iterations comprising the following: A multi-source classifier-free guidance mechanism is used to sample the logical value from the previous iteration output, wherein the multi-source classifier-free guidance mechanism is configured to weight the influence of the input audio and the input text respectively, ordering the sampled logic values based at least in part on a timeline of the target music, and A representation of the target music is updated based on the sorted, sampled logic values.
14. The system of claim 13, wherein the operations further comprise: generating an initial set of logical values based on the input representation of the audio and the input representation of the text; as well as generating an initial representation of the target music based on sampling and sorting the initial set of logical values; generating a second set of logical values based on the representation of the input audio, the representation of the input text, and the initial representation of the target music; as well as An updated representation of the target music is generated based on sampling and sequencing the second set of logical values.
15. The system of claim 13, the operations further comprising: Training data pairs are generated using music clips, where each training data pair includes background audio and target audio, where each music clip is divided into M sub-tracks, and M represents a positive integer.
16. The system according to claim 15, further comprising: generating the background audio in each training data pair by randomly selecting N sub-tracks from the M sub-tracks of the music segment in the music segment, wherein N represents a positive integer less than M; and The target audio in each training data pair is generated by randomly selecting remaining tracks from the M tracks of the music segment in the music segment.
17. A non-transitory computer-readable storage medium storing computer-readable instructions, the computer-readable instructions, when executed by a processor, causing the processor to perform operations comprising: inputting audio and text into a machine learning model, wherein the inputted audio indicates the background of target music, and wherein the inputted text specifies one or more instruments of the target music; and A representation of the target music is generated by an iterative process, wherein the target music features the one or more instruments specified by the input text and is aligned with the input audio, and wherein the iterative process comprises a plurality of iterations, each of the plurality of iterations comprising the following: A multi-source classifier-free guidance mechanism is used to sample the logical value from the previous iteration output, wherein the multi-source classifier-free guidance mechanism is configured to weight the influence of the input audio and the input text respectively, ordering the sampled logic values based at least in part on a timeline of the target music, and The representation of the target music is updated based on the sorted, sampled logical values.
18. The non-transitory computer readable storage medium of claim 17, the operations further comprising: generating an initial set of logical values based on the input representation of the audio and the input representation of the text; as well as generating an initial representation of the target music based on sampling and sorting the initial set of logical values; generating a second set of logical values based on the representation of the input audio, the representation of the input text, and the initial representation of the target music; as well as An updated representation of the target music is generated by sampling and sequencing based on the second set of logical values.
19. The non-transitory computer readable storage medium of claim 17, the operations further comprising: Training data pairs are generated using music clips, where each training data pair includes background audio and target audio, where each music clip is divided into M sub-tracks, and M represents a positive integer.
20. The non-transitory computer readable storage medium of claim 19, the operations further comprising: generating the background audio in each training data pair by randomly selecting N sub-tracks from the M sub-tracks of the music segment in the music segment, wherein N represents a positive integer less than M; and The target audio in each training data pair is generated by randomly selecting remaining tracks from the M tracks of the music segment in the music segment.