Generating music accompaniment

By tokenizing the input music data and generating music token sequences using machine learning algorithms, the computing resource-intensive and synchronization problems in the prior art are solved, and high-quality accompaniment synchronized with the original music is achieved on a personal device.

CN120418862APending Publication Date: 2025-08-01MACDOUGAL STREET TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088165.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-10
Filing Date
2023-12-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing machine learning methods are computationally resource-intensive when generating music, cannot run on personal devices, and the generated accompaniment is out of sync with the original music, resulting in inaccuracy and synchronization problems.

Method used

By tokenizing the input music data, using machine learning algorithms to generate music token sequences, combining feedback loops and prediction techniques, ensure that the generated music is synchronized with the original music, and the ML model is trained to accurately predict music events and time, generating accurate accompaniment.

Benefits of technology

Generating high-quality accompaniment synchronized with the original music on a computing resource-constrained device reduces computing requirements and improves the accuracy and synchronization of generated music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120418862A_ABST
    Figure CN120418862A_ABST
Patent Text Reader

Abstract

Techniques for generating accompaniment music data using machine learning (ML) techniques are described. In one implementation, music data in the MIDI format is tokenized into a sequence of music tokens. The trained ML model generates a sequence of output music tokens for a portion of the generated music signal based on the sequence of input music tokens for a portion of the original music signal. The portion of the generated music signal is temporally ahead of the portion of the original music signal used to generate the portion of the generated music signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing, and more particularly to generating music accompaniment. Background Art

[0002] The methods described in this section are methods that can be adopted, but not necessarily methods that have been previously envisioned or adopted. Therefore, unless otherwise stated, any method described in this section should not be assumed to be eligible as prior art merely because it is incorporated in this section.

[0003] Computer systems are widely used in audio processing. The acquisition, editing, encoding, storage, decoding, and reproduction of audio are key functions performed by computers today. Computer tools that perform these and other audio processing functions have greatly improved the quality of music production and consumption.

[0004] Although most audio processing capabilities have advanced significantly with the development of computer-related technologies, the conversion and generation capabilities of audio processing (especially music processing) are still limited in scope. Such capabilities mainly focus on improving the quality of existing music recordings or mixing existing music sources. Therefore, digital music processing lacks tools for creating new music or at least assisting in creating new music.

[0005] With the rise of artificial intelligence (AI) and the application of machine learning (ML) in many computer-related industries, methods for generating music using AI have been developed. However, current ML methods are computationally resource-intensive and inaccurate.

[0006] First, any machine learning method requires the use of existing large music datasets to train an ML model, and then only the trained ML model can be executed to generate music. According to one method, in order to provide a training dataset whose patterns can be used as a basis for music generation, audio can first be converted into a spectral representation, i.e., a spectrogram. However, when the spectrogram is fed into an ML algorithm to generate a trained ML model, the trained ML model may still produce inaccurate outputs due to the inherent inaccuracies of the spectrogram (e.g., errors due to noise and information loss in the conversion, especially time-related).

[0007] Second, large ML models cannot be generated on client computing devices with limited computational resources such as personal laptops or smartphones. Even if the generation of the model is transferred to cloud infrastructure that can actually have unlimited resources, audio models generated based on spectrograms or waveforms may still be computationally intensive and unable to run on client computing devices.

[0008] More importantly, for some music generation tasks, especially for music accompaniment generation, the latency between the input and output of the ML model must be minimized. Otherwise, the generated accompaniment will be out of sync (off-key) with the music it is generated for. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the drawings of certain implementations, the same reference numerals refer to corresponding parts throughout the drawings:

[0010] Figure 1 is a block diagram depicting the data flow of a music generation system in one implementation;

[0011] Figure 2 is a text image depicting an example of MIDI data;

[0012] Figure 3 is a block diagram depicting the tokenization of music data streams from different music sources in one implementation;

[0013] Figure 4 is a flowchart depicting the process of generating a new music token sequence in one implementation;

[0014] Figure 5A and Figure 5B is a block diagram depicting an example of a variant in one implementation;

[0015] Figure 6 is a block diagram of a basic software system in one or more implementations;

[0016] Figure 7 is a block diagram depicting a computer system on which an implementation of the present invention may be implemented. DETAILED DESCRIPTION

[0017] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it is apparent that the present invention may be practiced without these specific details. In other instances, structures and devices are shown in block diagram form to avoid unnecessarily obscuring the present invention. OVERALL OVERVIEW

[0018] The method of this article describes a music generation system that can continuously generate an audio signal based at least in part on an input music data stream. The system can play (or cause to play) the generated music (audio) signal in parallel with the original input signal on which the generated music signal is based. Thus, a user can compose music (e.g., play an instrument), while the system acquires / receives the original music signal / music data and generates new music data. When the new music data is played immediately after it is generated, the new music provides the user with what is commonly referred to as a "jamming" feeling for the original music.

[0019] The techniques described in this article include tokenizing the input music data and generating a new (next) sequence of music tokens based on the music tokens of the input music. Then, the generated sequence of music tokens can be converted into an audio signal that can be played by the system.

[0020] In one implementation, the tokenization technique is performed on data in MIDI format. The input music data is either received directly in MIDI (Musical Instrument Digital Interface) format or converted into music data in MIDI format, which is referred to as "MIDI data" in this article.

[0021] In one implementation, the input music digital representation (e.g., MIDI data) is converted into a sequence of music tokens. The term "music token" refers herein to a token that represents a specific feature of music. Music tokens with temporal correlation (e.g., representing events) are arranged in an order corresponding to their temporal occurrence within the signal.

[0022] Additionally or alternatively, the music tokens include separate tokens for time information. For example, a token can represent a specific time delay, indicating a time duration during which no additional events occur. Different from the MIDI format (where each event contains its time information), music tokens indicate the event itself (e.g., note on or note off), while separate tokens represent the time information (absolute time information or relative time information) of when the event occurs.

[0023] In one implementation, a training set of music token sequences is used to train a machine learning (ML) algorithm to generate an ML model, such as a large language model (LLM), such as A sequence of music tokens (whose next music token sequence in time is known) is provided to the ML model to generate the next sequence of music tokens. The generated sequence of music tokens is compared with the known next time sequence to determine the error. Based on the error, the parameters of the ML model are modified. These steps can be repeated until the most accurate token output sequence is produced, thereby generating a trained ML model.

[0024] For more details on ML algorithms and model techniques, see the following sections: "Generating New Music Tokens", "Other Machine Learning Techniques", "Machine Learning Algorithms and Domains", and "Hyperparameters, Cross-Validation, and Algorithm Selection".

[0025] At runtime, the system provides a sequence of music tokens as input to the trained ML model. The ML model generates a new sequence of music tokens, predicting music that the system may not have received yet. The trained ML model can make accurate predictions about music data because the tokens not only represent music events, but also the delay between individual tokens represents the delay between music events. Therefore, the trained ML model can not only accurately generate the notes to be played, but also accurately generate the timing for playing each corresponding note. The new sequence of music tokens can then be converted into an audio signal and played.

[0026] However, when generating a new sequence of music tokens, the music data for which the music token sequence is generated may already have been played. Therefore, the generated music data and the current music data may still be out of sync. To correct this problem, the previously generated music data (which is a prediction of the next music input data) is provided back to the ML model as a new input (partially). Such a music data loop can also increase the difference between the generated music and the input, and generate accompaniments of various degrees. The degree of difference depends on the input ratio of the number of music token sequences received by the system to the number of music token sequences previously generated by the system.

[0027] For example, in addition to the newly received sequence of music tokens (e.g., from received MIDI data or captured audio signals), the ML model also uses the newly generated sequence of music tokens as an input feature. Therefore, the trained ML model is used at least in part in a feedback loop where the newly generated sequence of music tokens is used again as an input together with the newly captured sequence of music tokens. Because in one implementation, the newly generated sequence of music tokens is at least partially based on a) a music token sequence that is different from the original and b) at least in part on its own prediction of the timing of the next token (future of the future), the music of the new token sequence should be in sync with the original music, but is also a derivative (accompaniment) of that music.

[0028] In one implementation, different ML models can be trained for different types of music tokens. In particular, separate ML models can be trained to predict metrics of the generated music, such as polyphony and / or intensity. The term "polyphony" refers to the maximum number of different notes that are simultaneously on in a sequence of music tokens. The on state is described by a note that has a note-on event token in the music token sequence and that does not have a corresponding note-off event token during the duration of the state. The term "intensity" refers to the maximum number of different notes that are on during a specific time period (e.g., a measure), but not necessarily simultaneously.

[0029] Additionally or alternatively, the intensity and / or polyphony of the generated music data can be configured by the user. The polyphony and intensity (determined by the ML model or configured by the user) can adjust the ML model to generate music tokens for the output signal. The ML model can be trained to take the polyphony and / or intensity as inputs and generate music tokens based on the polyphony and / or intensity. Thus, the same ML model with the same music token sequence can generate different music data based on different polyphony and / or intensity. Audio signal to digital representation

[0030] Figure 1 is a block diagram depicting the data flow for generating an audio system in one implementation. At block 110, a user-generated music signal is converted into a digital audio signal representation, such as MIDI data 115. Thus, the execution of block 110 generates real-time MIDI data 115 for the user-generated music signal. As a non-limiting example, the user-generated music signal can be generated by one or more musical instruments and / or vocal performances.

[0031] Data in MIDI format (such as MIDI data 115) fully describes music in the form of notes, digitizing the musical score of a piece of music. MIDI is a standard for digital music representation in computing systems and describes the communication protocol, digital interface, and electrical connectors that connect various electronic musical instruments, computers, and related audio devices to play, edit, and record music. The MIDI standard includes text symbols representing various events similar to playing musical score notes. Although the examples and implementations in this document refer to the MIDI format and its data as digital note representations, the exact input format used to digitally represent notes is not important for the techniques described herein.

[0032] Additionally or alternatively, MIDI data 115 can be received directly from the user device. For example, a user guitar can include a MIDI pickup device that converts the played music into MIDI data 115 in real time, or a synthesizer can directly generate MIDI data 115. Such user devices can be communicatively coupled to the system, and the system can receive MIDI data 115 without any conversion.

[0033] Figure 2 is a text image depicting an example of MIDI data. MIDI data includes messages, and each message can describe an event. A note start play event or a note stop play event (referred to as a MIDI note on event and a MIDI note off event, respectively) can describe a note identifier ("note"), a track identifier ("channel"), and time information ("time").

[0034] In one implementation, to accurately represent captured audio data in digital note representation, large amounts of audio data segments are processed. The larger the segment, the more accurate the conversion of the audio data to its frequency domain and the notes generated based on the frequency-transformed data. However, the larger the segment, the greater the introduced latency. Thus, an accurate transformation introduces such a large latency between the audio signal and the final digital representation that it is no longer considered real-time.

[0035] However, when a real-time audio signal is processed in smaller parts (referred to herein as "windows"), no significant latency is introduced. Although each window may not have a sufficient number of audio samples to accurately convert to digital notes, the middle part of the window is accurate when converted to digital note data. In one implementation, the next window is selected such that its filtered middle part is a time continuation of the filtered middle part of the previous window. Thus, a real-time audio signal can be processed in a set of overlapping frame windows. The term "frame" refers to a sequence of audio samples and its transformation. When processing the current window, the audio samples for the next or subsequent window are collected through real-time acquisition.

[0036] In one implementation, the sequence of audio samples for each frame is converted to a corresponding frequency-domain frame, generating a set of frequency-domain frames for the window. Each set of frequency-frame windows is converted to a corresponding set of note (abbreviated as "n") event probability values. The probability values for the filtered set of frames for each window are converted to n-on and n-off events. The term "n-on event" refers to the event of detecting the playing of an n that was not previously played in the audio signal. The term "n-off event" refers to the event of detecting the end of the playing of a previously played n. Based on the probabilities determined for the frame, the technique describes determining whether the frame contains an n-on event, an n-off event, or neither for each n. Music Tokenization

[0037] A large language model is trained and operated using tokens retrieved from input natural language. In the case of natural language, a token can be a string between letters and words, such as a syllable. Since MIDI events aggregate multiple pieces of information, using MIDI event representation creates inherent complexity and inaccuracy in the LLM output that requires tokenized input. In one implementation, to use the LLM to predict the next music event, MIDI events are converted to music tokens. Continuing reference Figure 1 , at block 120, an input stream of MIDI data is converted to a sequence of music tokens 125.

[0038] Each MIDI event type is converted into a corresponding musical token. For example, a note-on MIDI message for a specific note is converted into a note-on token with a specific note value, while a note-off MIDI message for a specific note is converted into a note-off token with a specific note value. The musical tokens of the events are arranged in the same order as the corresponding MIDI event messages.

[0039] The time information of MIDI events can also be tokenized. Generally, the tokens of an LLM for natural language are arranged in a sequential information stream in text format. This representation fails to consider the time information. Therefore, the time information is also tokenized and inserted into the appropriate position in the musical token sequence.

[0040] In one implementation, in addition to the event information (such as the notes being played or stopped), musical tokens indicating the time information are generated. Therefore, the LLM can not only predict whether a note is being played, but also predict when the note starts or stops and for how long.

[0041] Additionally or alternatively, the musical tokens are arranged in a time grid corresponding to a user-configurable "beat" (a period during which the music can repeat). The period based on the beat is referred to as a measure in this article (for example, 4 beats per measure). For example, a measure start token and / or a measure end token are inserted into the appropriate positions at the start and end of a measure in the musical token sequence.

[0042] In one implementation, each measure is divided into a grid of equal time period metrics, referred to as measure grid cells in this article. The time musical token can have a value of the number of measure grid cells of the time period between adjacent events (relative time information) or the number of measure grid cells from the start of the measure to the occurrence of the event (absolute time information).

[0043] In one implementation, the musical event tokens are adjusted to represent events occurring in the nearest grid cell of a measure. The MIDI data at a specified time is converted to the nearest grid unit of the measure to determine the time musical token to be generated for the musical event token. The "time" attribute in the MIDI event message specifies the amount of time ticks that have passed since the previous message (see Figure 2 )). The MIDI data specifies the exact tick resolution used for the measure. If the tick resolution is the same as the grid cell resolution, the value of the time attribute of the event in the MIDI data can be used for the time musical token. If the tick resolution is different from the measure grid unit resolution, the time attribute of the MIDI data is converted to the nearest measure grid unit. Then this value is assigned to the time musical token.

[0044] In one implementation, the input music data representation can have multiple tracks for different types of music sources (e.g., soprano, alt voice, and different musical instruments). The input MIDI data can contain separate MIDI event message streams for each different music source and can indicate the type of music source. The system can receive these separate message sequences in parallel from different sources. For example, different microphones can be connected to different instruments / people, and different electronic musical instruments can generate their corresponding different MIDI data streams.

[0045] Figure 3 is a block diagram depicting tokenization of music data streams from different sources in one implementation. The MIDI sample data 115 received by the system includes four tracks / channels. For each channel, separate MIDI data is received, including 8-bar piano MIDI track data 310, guitar MIDI track data 320, drum MIDI track data 330, and bass MIDI track data 340.

[0046] In this example, the system can convert the received MIDI sample data 115 into a configured number of token sequences of bars. If the system is configured to process MIDI samples 115 in 4-bar chunks at a time, the system will tokenize the first 4 bars of each track in the MIDI sample 115 and then tokenize the last 4 bars of the MIDI sample data 115.

[0047] In one implementation, when the system identifies a track in the MIDI data (e.g., after processing a configured number of bars in the current track and the system switches to the next track), the system generates a token indicating the start of that track. After the track token start is a token representing the track metadata.

[0048] In one implementation, the metadata of a track includes the type of music source of the track (e.g., instrument or voice type). The system maps the type of music source identified in the MIDI track to the music source type identifier of the instrument token. In one implementation, the system can map multiple types of music sources in the MIDI data to a single music source type in the music token sequence. For example, multiple types of guitars (e.g., acoustic nylon guitar, acoustic steel guitar, electric guitar, electric clean guitar, electric muted guitar, overdriven guitar, distorted guitar, and guitar harmonics in MIDI) can be mapped to the guitar type music source in the music token sequence.

[0049] Additionally or alternatively, the track metadata may indicate metric values of the music of the MIDI data. The metrics in the metadata may include the polyphony and / or intensity of the track. To determine the polyphony based on the MIDI data of the track, the system may calculate the average number of notes (or any other statistical aggregation function) played during a single MIDI tick to determine the polyphony. To calculate the intensity per bar (or any other time period), the system may calculate the total number of notes in the sequence within that time period divided by the polyphony. In one implementation, the system generates tokens for polyphony and / or intensity as track metadata and inserts them after the track start indication.

[0050] The following is an example of a sequence of music tokens generated for a bar in a track in sample MIDI data. This example is for a track with instrument indicator 30 (e.g., piano). After generating the TRACK_START token indicating the start of the piano track, the determined track metadata is inserted into the token sequence: "INT = 4" indicates an intensity value of 4, and "POL = 2" indicates a polyphony value of 2.

[0051] The "BAR_START" and "BAR_END" tokens indicate the start and end of the first bar of the track. In a multi-track MIDI sample, the number of bars of a track generated before generating tokens for the next track depends on the configuration of the system. If the system is configured to process 4 bars at a time, 4 piano tracks will be generated before the system moves on to the next guitar track to generate 4 bars. Similarly, for 8 bars, the system generates the complete 8 bars of the piano MIDI track 310, such as the piano token track 315, and then continues to generate 8 bars of the guitar MIDI track 320 to produce the guitar token track 325.

[0052] Between the "BAR_START" and "BAR_END" tokens, the system generates music event tokens "NOTE_ON" and "NOTE_OFF" and time token "TIME_DELTA" based on the MIDI event messages of the track for a particular bar. In the following example, notes #57 and #53 are played at the start of the bar, 20 bar grid units later, notes #41 and #48 are played, and 6 grid units later, note #57 stops being played. TRACK_START INT=4 POL=2 INSTR=30 BAR_START NOTE_ON=57 NOTE_ON=53 TIME_DELTA=20 NOTE_ON=41 NOTE_ON=48 TIME_DELTA=6 NOTE_OFF=57……BAR_END……

[0053] ContinueFigure 3 An 8-bar audio track example in Generate new music tokens

[0054] Continue Figure 1 , at block 140, the token sequence of the original music is used to generate a new music token sequence that represents new music that can be played as an accompaniment to the original music. In one implementation, machine learning techniques are used to generate the new music token sequence 145 based on the token sequence 125. Machine learning techniques include applying a machine learning algorithm to a training data set of music token sequences for which the next music token sequence result is known and using initialization parameters whose values are modified in each training iteration to more accurately produce the known result (referred to herein as a "label").

[0055] The term "machine learning algorithm" (or simply "algorithm") as used herein refers to a process or set of rules to be followed in computing, where the model artifact containing one or more computing parameters is unknown. The term "machine learning model" (or simply "model") as used herein refers to a process or set of rules to be followed in computing, where the model artifact containing one or more parameters is known and is derived by training the corresponding machine learning algorithm using one or more training data sets. Once training is complete, the input is applied to the machine learning model for prediction, which may also be referred to herein as a prediction result or output.

[0056] In one implementation, the system can be configured as a machine learning model to use a specific number of bars of a music token sequence to produce the next music token sequence of a specific number of bars. Thus, the machine learning algorithm can also be trained using music token sequences of the same length in each training iteration. For example, a 4-track 8-bar sample music token sequence is passed through the machine learning algorithm. The tracks can be randomly selected, and each track represents a different type of music source (e.g., a different instrument).

[0057] The machine learning algorithm predicts the probability of the next token based on a given set of parameter values. For example, in each iteration, the ML model is trained to generate an array containing the probability of each token in the token dictionary. The probability describes the likelihood that the corresponding token will be the next token given the current parameter values. The music token vocabulary can include all possible MIDI event types and can be sized from 300 - 500 tokens. In one implementation, the cross-entropy loss function is used to optimize the parameters of the ML model in each iteration.

[0058] For such applications based on a training dataset, the technique generates a machine learning model with known parameters. Thus, the machine learning model includes a model data representation or a model artifact. The model artifact contains parameter values that the machine learning algorithm applies to the input to generate a predicted output. Training the machine learning model requires determining the parameter values of the model artifact. The structure and organization of the parameter values depend on the machine learning algorithm. Expected music metrics

[0059] Continuing to refer Figure 1 , the system can determine expected music metrics, such as polyphony and / or intensity, based on the token 125 at block 130. The expected music metrics are used in the configuration 135 to adjust the ML model to generate new music tokens.

[0060] In one implementation, the system can include a trained machine learning model for determining the expected music metrics of new (not yet generated) music tokens. Such a machine learning model is also trained by a training set of sequences of music tokens with known music metric label values.

[0061] The system can receive a first sequence of music tokens (e.g., 8 bars of a multi-track) and use the trained metric machine learning model to determine the probability of the expected metrics for each track. The expected metric value with the highest probability is selected for each track.

[0062] To ensure that the generated music data conforms to the expected music metrics, the new token generation ML is configured with the determined metric values. Since the music token generation ML model has been trained using training set data that includes metric tokens, there is an inherent dependency between the metric values and the generation of new music event tokens. Thus, once the music token generation ML model is configured with the expected metrics, the generated new music tokens will conform to the expected music metrics.

[0063] For example, the expected music metric model can use the number of track bars of the first input to generate probabilities for polyphony values from 1 - 10 and intensity values from 1 - 5. Based on the highest probability values, the system selects the expected polyphony and expected intensity metric values. Then, the system configures the LLM with the expected polyphony and intensity metric values. When the first token sequence and subsequent token sequences are input into the configured LLM, the LLM generates a new sequence of music tokens indicating music events according to the polyphony and intensity levels.

[0064] In another implementation, the system can receive the expected metric values, which can be user-configurable (e.g., received from a user device). Thus, block 130 for determining the music metrics may not be executed. The music token generation machine learning model is configured by the configuration 135, which includes the expected metric values received from the user device. Generate a new music signal

[0065] Continue Figure 1 At block 140, the system generates a new music token sequence 145 based on the original music token sequence 125. At block 150, the new music token sequence 145 can be converted into a music audio signal. While continuously receiving MIDI data 115 (derived from continuous music audio capture), new music signals can be continuously generated (resulting in generation) based on the continuously generated new token sequences 145. The system can be configured to generate the next music token sequence for a specific time period (e.g., 4 bars or 8 bars at a time) at block 140. Additionally, the system can be configured based on a configuration 135 with intensity and polyphony configuration values to generate new music tokens for a specific density of note-on states (intensity) or parallelism (polyphony) in a plurality of simultaneous note-on states represented by corresponding tokens.

[0066] Figure 4 is a flowchart depicting the process of generating a new music token sequence in one implementation. In step 420, the system receives a first set of original music token sequences tokenized from MIDI data. In step 430, in the first iteration, no new music tokens have been generated. Thus, the system provides a configured number of music token sequences from the received original music token sequence to a music token generation ML model, which generates the probability of a new music token sequence. The configured number can be a specific number of bars of music tokens per track.

[0067] In step 440, the ML model generates the probability for each music token to be the next music token based on the provided music token sequence. The ML model can generate probabilities for each music token of each sequence corresponding to the input sequence (e.g., each bar). In one implementation, the ML model generates probabilities for every token / value combination representing a unique music event and all permutations of time values. The probabilities of non-event tokens (e.g., tokens for track / bar separators and metadata) can be omitted from the generation. Thus, the ML model can generate a set of probabilities that includes the probabilities of note-on and note-off event tokens for each supported note, as well as the probabilities of time tokens representing the delay from one grid cell to a full bar.

[0068] In step 450, tokens for the next music token sequence are selected based on the generated corresponding probabilities. The system may discard possible next tokens, which may cause the sequence to repeat itself. The system can sample the (remaining) tokens by picking some of the most likely continuations from the probability distribution and comparing the overall probabilities of the sequences that could be generated. If the ML model is configured to generate a single set of probabilities for multiple input bars, then in step 440, the system can determine the positions where the selected new token sequence needs to be split to generate a token sequence for each bar.

[0069] In step 460, the system can generate an audio signal based on the new music token sequence. Additionally or alternatively, the system can convert the music tokens into MIDI data and provide the output MIDI data to a communicatively coupled device for playback. Thus, the system generates or causes to be generated a musical accompaniment for the original music based on the original music signal itself. New music signal variation

[0070] The system can vary the new music data such that the new music signal has a different degree of variation from the original music signal. Continuing to refer Figure 1 to, the system can be configured to generate a music token sequence at block 140 based on a newly generated token sequence 145 (or a portion thereof) in addition to the original tokens 125. The configuration of the amount of variation or mixing between the original token sequence 125 and the newly generated token sequence 145 is referred to herein as "variance" or "variance level". The variance can be set as a percentage (relative) of the new token sequence that serves as the input to block 140, or as the number of measures (absolute) of the new token sequence that serves as the input to block 140.

[0071] Continuing Figure 4 to, in one implementation, after step 450, the process can return to step 420 to process the next original music token sequence. For the second and subsequent iterations, the new music token sequence has been generated. Thus, in step 410, the system can receive the music token sequence generated in the last iteration.

[0072] For the second and subsequent iterations, in step 430, the system can change the input by providing the newly generated music token sequence from the previous iteration along with the original music token sequence, thereby introducing variance in the output.

[0073] Figure 5A and Figure 5B are block diagrams depicting examples of variation in one implementation. In Figure 5A and Figure 5B 's example, the music token generator 550 (corresponding to Figure 1 operation block 140 in) is configured to take an 8-measure original music token sequence as input to produce an 8-measure new music token sequence. The system is configured to have a 3 / 8 variation, i.e., provide the previously generated new music token sequence for at least 3 of the 8 measures of this input.

[0074] Figure 5A depicts an example of the first iteration. Since no new music tokens have been generated yet, the system provides the original music tokens, i.e., original measures 501 (earliest) to 508 (latest), to the music token generator 550. The music token generator 550 generates the corresponding new music tokens, new measures 551 to new measure 558.

[0075] In one implementation, the system can randomly or by time (selecting the earliest or latest music token sequence) select one or more new music token sequences as input. Continuing Figure 5A with the example, the system selects the variance token sequence 531, which, according to the variance configuration, contains 3 new bar music token sequences 556 - 558 from among the 8 new bar music token sequences 551 - 558 for re - input. The selection is based on the latest new bar music token sequences 551 - 558 in terms of time.

[0076] Figure 5B Depicts an example of the next iteration. Not all of the input music token sequences are the original token sequences from the token data 125. The input contains Figure 5A the variance token sequence 531 from. Thus, the next original bar music token sequences 509 - 513 are provided to the music token generator 550 together with the new bars 556 - 558, which were generated by the music token generator 550 in the previous iteration.

[0077] The generated token sequence, i.e., the new bar sequence 559 - 566, has now changed by having 3 non - original music token sequences at the input of the music token generator 550. For the next iteration, the 3 non - original music token sequences are the variance token sequence 532.

[0078] Thus, in such an implementation, the new music data generated in each iteration is different from the original music data because (at least in part) non - original music token sequences are used. At the same time, the new music data continues to be based on the original music data. Therefore, the new music signal generated is an accompaniment to the original music signal. Additionally, when the input is based on the latest new music token sequences, the lag between the new music generated from the new music data and the original music should be small or non - existent. Other machine learning training techniques

[0079] In supervised training, the training data is used by a supervised training algorithm to train a machine learning model. The training data includes inputs and "known" output labels. In one implementation, the supervised training algorithm is an iterative process. In each iteration, the machine learning algorithm applies the model artifact and the input to generate a predicted output. The error or variance between the predicted output and the known output is calculated using an objective function. In fact, the output of the objective function indicates the accuracy of the machine learning model based on a particular state of the model artifact in the iteration. By applying an optimization algorithm based on the objective function, the parameter values of the model artifact are adjusted. The iteration can be repeated until the desired accuracy is achieved or some other criteria are met.

[0080] In one implementation, to iteratively train an algorithm to generate a training model, the training dataset can be arranged such that each row of the dataset is input into the machine learning algorithm, and the actual result and label value corresponding to that row are further stored. For example, each row of the Adult Income dataset represents a specific adult, and the result of that adult is known, such as whether the total income of that adult exceeds $500,000. Each column of the Adult training dataset contains a numerical representation of a specific adult characteristic (e.g., whether the adult has a college degree, the age of the adult...). Based on this, after training, the algorithm can accurately predict whether the total income of any adult (even an adult not described in the training dataset) exceeds $500,000.

[0081] The row values of the training dataset can be used as the input to the machine learning algorithm and can be modified according to one or more parameters of the algorithm to produce a prediction result. The prediction result of a row is compared with the label value, and an error value is calculated based on the difference. One or more error values of the batch rows are used in a statistical aggregation function to calculate the error value of the batch. The term "loss" refers to the error value of a batch of rows.

[0082] In each training iteration, based on one or more predicted values, the loss value corresponding to that iteration is calculated. For the next training iteration, one or more parameters will be modified to reduce the loss based on the current loss. The training dataset can be iterated any number of times to reduce the loss. The training iteration using the training dataset can be stopped when the change in loss between iterations is within a threshold range. In other words, the iteration stops when the losses of different iterations are basically the same.

[0083] After the training iteration, the generated machine learning model includes the machine learning algorithm of the model artifact that produces the minimum loss.

[0084] For example, the Support Vector Machine (SVM) algorithm can be used to iterate over the above Adult Income dataset to train an SVM-based Adult Income dataset model. Each row of the Adult dataset is provided as the input to the SVM algorithm, and the result (prediction result) of the SVM algorithm is compared with the actual result of that row to determine the loss. Based on the loss, the parameters of the SMV are modified. The next row will be provided to the SVM algorithm with the modified parameters to obtain the prediction result of the next row. This process can be repeated until the difference between the loss values of the previous iteration and the current iteration is below a predefined threshold, or in some implementations, until the difference between the achieved minimum loss value and the loss of the current iteration is below a predefined threshold.

[0085] Once the machine learning model of the machine learning algorithm is determined, a new dataset with unknown results can be used as the input to the model to calculate the prediction results of the new dataset.

[0086] In a software implementation scenario, when a machine learning model is said to receive inputs, perform, and / or generate outputs or predictions, the computer system process that executes the machine learning algorithm applies the model artifacts to the inputs to generate a predicted output. The computer system process executes the machine learning algorithm by running software configured to cause the algorithm to execute. Machine Learning Algorithms and Domains

[0087] A machine learning algorithm can be selected based on the problem domain and the type of expected result required for the problem. Non-limiting examples of algorithm result types can be discrete values for problems in the classification domain, continuous values for problems in the regression domain, or anomaly detection problems in the clustering domain.

[0088] However, even for a specific domain, there are many algorithms to choose from for selecting the most accurate algorithm to solve a given problem. As non-limiting examples, in the classification domain, Support Vector Machines (SVMs), Random Forests (RFs), Decision Trees (DTs), Bayesian Networks (BNs), stochastic algorithms (such as Genetic Algorithms (GAs)), or connectionist topologies (such as Artificial Neural Networks (ANNs)) can be used.

[0089] Implementations of machine learning may rely on matrices, symbolic models, and hierarchical and / or associative data structures. Parameterized (i.e., configurable) implementations of best-of-breed machine learning algorithms can be found in open-source libraries such as Google's Python and C++ TensorFlow or the Georgia Institute of Technology's C++ MLPack. Shogun is an open-source C++ ML library with adapters for multiple programming languages including C#, Ruby, Lua, Java, MatLab, R, and Python. Hyperparameters, Cross-Validation, and Algorithm Selection

[0090] A machine algorithm can have an infinite number of variations based on one or more hyperparameters. The term "hyperparameter" refers to a parameter in the model artifacts that is set before the training of the machine algorithm model and is not modified during the model training process. In other words, hyperparameters are constant values that affect (or control) the generated training model independent of the training dataset. A machine learning model with a model artifact containing only hyperparameter values is referred to in this document as a "variant of the machine learning algorithm" or simply a "variant". Thus, during the model training process, different hyperparameter values of the same type of machine learning algorithm can produce significantly different loss values on the same training dataset.

[0091] For example, the SVM machine learning algorithm contains two hyperparameters: "C" and "gamma". The hyperparameter "C" can be set to range from 10 -3 to 105 for any value, and the hyperparameter "gamma" can be set to range from 10 -5 to 10 3 for any value. Thus, there are an infinite number of permutations and combinations of the "C" and "gamma" parameters, which can produce different loss values when training the same adult income training dataset.

[0092] Therefore, to select an algorithm, or further select the best-performing variant among algorithms, various hyperparameter selection techniques are used to generate different sets of hyperparameter values. Non-limiting examples of hyperparameter value selection techniques include Bayesian optimization (such as Gaussian processes for hyperparameter value selection), random search, gradient-based search, grid search, manual tuning techniques, techniques based on Tree-structured Parzen Estimator (TPE).

[0093] By selecting different sets of hyperparameter values based on one or more of these techniques, each machine learning algorithm variant is trained on the training dataset. The test dataset is used as the input to the trained model to calculate the predicted result values. The predicted result values are compared with the corresponding label values to determine the performance score. The performance score can be calculated based on the error rate of the predicted results related to the corresponding labels. For example, in the classification domain, if among 10,000 inputs of the model, only 9,000 match the input labels, the performance score is calculated as 90%. In the non-classification domain, the performance score can be further based on the statistical aggregation of the differences between the label values and the predicted result values.

[0094] The term "trial" as used herein refers to training a machine learning algorithm using different sets of hyperparameter values and testing the machine learning algorithm using at least one test dataset. In one implementation, cross-validation techniques, such as k-fold cross-validation, are used to create multiple pairs of training and test datasets from the original training dataset. Each pair of datasets together contains the original training dataset, but these pairs divide the original dataset in different ways between the training dataset and the test dataset. For each pair of datasets, the training dataset is used to train the model based on the selected set of hyperparameters, and the corresponding test dataset is used to calculate the predicted result values with the trained model. Based on inputting the test dataset into the trained machine learning model, the performance score for that pair (or fold) is calculated. If there are more than one pair (i.e., folds), the performance scores will be statistically aggregated (such as average, mean, minimum, maximum) to obtain the final performance score of the machine learning algorithm variant.

[0095] Each trial is computationally very expensive because it includes multiple training iterations of the machine algorithm variant to generate performance scores for different sets of hyperparameter values of the machine learning algorithm. Therefore, reducing the number of trials can significantly reduce the computational resources (such as processor time and cycles) required for fine-tuning.

[0096] In addition, since the generated performance score is used to select the most accurate algorithm variant, the more precise the performance score itself is, the more precise the relative accuracy of the predictions of the generated model will be compared to other variants. In fact, once a machine learning algorithm and its variant based on hyperparameter values are selected, the algorithm variant can be applied to the complete training data set to train a machine model by leveraging the techniques discussed above. It is expected that the generated machine learning model can predict the results more accurately than the machine learning models of any other algorithm variant.

[0097] The precision of the performance score itself depends on the amount of computing resources spent on fine-tuning the hyperparameters of the algorithm. Computing resources may be wasted on testing sets of hyperparameter values that cannot achieve the required accuracy of the final model.

[0098] Similarly, for one algorithm, it may take less (or no) computing resources to fine-tune these hyperparameters, and this algorithm is likely to be less accurate than another algorithm. Therefore, the number of trials of the hyperparameters of the discount algorithm can be reduced or eliminated, thereby greatly improving the performance of the computer system. Software Overview

[0099] Figure 6 is a block diagram of a basic software system 600 that can be used to control Figure 7 the operation of a computing system 700. The software system 600 and its components (including their connections, relationships, and functions) are for illustration only and are not intended to limit the implementation of exemplary implementations. Other software systems suitable for implementing exemplary implementations may have different components, including components with different connections, relationships, and functions.

[0100] The software system 600 is provided to guide the operation of the computing system 700. The software system 600 can be stored in the system memory (RAM) 706 and the fixed storage (e.g., hard disk or flash memory) 710, including a kernel or operating system (OS) 610.

[0101] The OS 610 manages the underlying aspects of computer operations, including managing process execution, storage allocation, file input / output (I / O), and device I / O. One or more applications represented as 602A, 602B, 602C... and 602N can be "loaded" (e.g., transferred from the fixed storage 710 to the memory 706) for execution by the system 600. Applications or other software intended to be used on the computing system 700 can also be stored as a set of downloadable computer-executable instructions, e.g., for downloading and installing from an Internet location (e.g., a web server, an app store, or other online services).

[0102] The software system 600 includes a graphical user interface (GUI) 615 for receiving user commands and data in a graphical manner (e.g., "click" or "touch gesture"). In turn, the system 600 can act on these inputs according to instructions from the operating system 610 and / or application 602. The GUI 615 is also used to display the operation results of the OS 610 and application 602, and then the user can provide additional inputs or terminate the session (e.g., log off).

[0103] The OS 610 can execute directly on the bare hardware 620 (e.g., processor 704) of the computer system 700. Alternatively, a virtual machine manager or virtual machine monitor (VMM) 630 can be inserted between the bare hardware 620 and the OS 610. In this configuration, the VMM 630 acts as a software "buffer" or virtualization layer between the OS 610 and the bare machine 620 of the computer system 700.

[0104] The VMM 630 instantiates and runs one or more virtual machine instances ("guests"). Each guest includes: a "guest" operating system, such as the OS 610; and one or more applications, such as the application 602, which are designed to execute on the guest operating system. The VMM 630 provides a virtual operating platform for the guest operating system and manages the execution of the guest operating system.

[0105] In some cases, the VMM 630 can allow the guest operating system to run as if it were running directly on the bare machine 620 of the computer system 700. In these cases, the same version of the guest operating system configured to execute directly on the bare machine 620 can also execute on the VMM 630 without modification or reconfiguration. In other words, in some cases, the VMM 630 can provide full hardware and CPU virtualization for the guest operating system.

[0106] In other cases, the guest operating system can be specifically designed or configured to execute on the VMM 630 for increased efficiency. In these cases, the guest operating system "knows" that it is executing on the virtual machine monitor. In other words, in some cases, the VMM 630 can provide paravirtualization for the guest operating system.

[0107] The computer system processes include hardware processor time allocation and storage allocation (physical and / or virtual), the storage allocation is used to store the instructions executed by the hardware processor, to store the data generated by the hardware processor executing the instructions, and / or, when the computer system process is not running, to store the hardware processor state (e.g., register contents) between the hardware processor time allocations. The computer system processes run under the control of the operating system and can also run under the control of other programs executed on the computer system.

[0108] Multiple threads can run within a process. Each thread also includes hardware processing time allocation, but shares access to the memory allocated to the process. When a thread is not running, the memory is used to store the processor content between allocations. The term thread can also be used to refer to multiple non-running threads in a computer system process. Cloud computing

[0109] The term "cloud computing" is generally used to describe a computing model that supports on-demand access to a shared pool of computing resources such as computer networks, servers, software applications, and services, and that allows for the rapid configuration and release of resources with minimal administrative effort or service provider interaction.

[0110] Cloud computing environments (sometimes referred to as cloud environments or the cloud) can be implemented in a variety of different ways to best suit different needs. For example, in a public cloud environment, the underlying computing infrastructure is owned by an organization that provides its cloud services to other organizations or the public. In contrast, a private cloud environment is typically used only by or within a single organization. Community clouds are intended to be shared by multiple organizations within a community, while hybrid clouds contain two or more types of clouds (such as private clouds, community clouds, or public clouds) bound together through data and application portability.

[0111] In general, cloud computing models can deliver some of the responsibilities that were previously provided by an organization's own information technology department as a service layer in a cloud environment for consumers to use (either inside or outside the organization, depending on the public / private nature of the cloud). Depending on the specific implementation, the precise definition of each cloud service layer's provided or internal components or functions may vary, but common examples include: Software as a Service (SaaS), where consumers use software applications running on cloud infrastructure, and the SaaS provider manages or controls the underlying cloud infrastructure and applications. Platform as a Service (PaaS), where consumers can use software programming languages and development tools supported by the PaaS provider to develop, deploy, and otherwise control their own applications, and the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything below the runtime execution environment). Infrastructure as a Service (IaaS), where consumers can deploy and run any software applications and / or provide processing, storage, networking, and other basic computing resources, and the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS), where consumers use database servers or database management systems running on cloud infrastructure, and the DBaaS provider manages or controls the underlying cloud infrastructure, applications, and servers, including one or more database servers. In a cloud computing environment, there is no insight into the application or application data. For planned operations that require disconnection, using the techniques discussed in this document, sessions can be released and then rebalanced later without interrupting the application.

[0112] The above basic computer hardware and software, as well as the cloud computing environment, are presented to illustrate the basic underlying computer components that can be used to implement the exemplary implementations. However, the exemplary implementations are not necessarily limited to any specific computing environment or computing device configuration. Instead, the exemplary implementations can be implemented in any type of system architecture or processing environment that those skilled in the art understand, based on this disclosure, to be capable of supporting the features and functions of the exemplary implementations presented herein. Hardware Overview [[ID=�]]

[0113] According to one implementation, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices can be hard-wired to perform the techniques, or can include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or can include one or more general-purpose hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices can also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to implement the techniques. The special-purpose computing devices can be a desktop computer system, a portable computer system, a handheld device, a network device, or any other device that combines hard-wired and / or program logic to implement the techniques.

[0114] For example, Figure 7 is a block diagram showing a computer system 700 on which the present invention can be implemented. The computer system 700 includes a bus 702 or other communication mechanism for conveying information and a hardware processor 704 coupled to the bus 702 for processing information. The hardware processor 704 can be, for example, a general-purpose microprocessor.

[0115] The computer system 700 also includes a main memory 706, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 702 for storing information and instructions to be executed by the processor 704. The main memory 706 can also be used to store temporary variables or other intermediate information during execution of instructions by the processor 704. When stored in a non-transitory storage medium accessible to the processor 704, these instructions cause the computer system 700 to become a special-purpose machine customized to perform the operations specified in the instructions.

[0116] The computer system 700 also includes a read-only memory (ROM) 708 or other static storage device coupled to the bus 702 for storing static information and instructions for the processor 704. A storage device 710, such as a magnetic disk or optical disk, is provided and coupled to the bus 702 for storing information and instructions.

[0117] The computer system 700 can be coupled via the bus 702 to a display 712, such as a cathode ray tube (CRT), for displaying information to a computer user. An input device 714 including alphanumeric keys and other keys is coupled to the bus 702 for communicating information and command selections to the processor 704. Another type of user input device is a cursor control 716, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 704 and controlling cursor movement on the display 712. This input device typically has two degrees of freedom in two axes (i.e., a first axis (e.g., x) and a second axis (e.g., y)), which enables the device to specify positions within a plane.

[0118] The computer system 700 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, when combined with the computer system, cause or program the computer system 700 to be a special-purpose machine. According to one implementation, the techniques herein are performed by the computer system 700 in response to one or more sequences of one or more instructions contained in the main memory 706 being executed by the processor 704. Such instructions can be read into the main memory 706 from another storage medium, such as the storage device 710. Execution of the instruction sequence contained in the main memory 706 causes the processor 704 to perform the processing steps described herein. In an alternative implementation, hardwired circuitry may be used in place of or in combination with software instructions.

[0119] The term "storage medium" as used herein refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as the storage device 710. Volatile media includes dynamic memory, such as the main memory 706. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge.

[0120] Storage media is distinct from but can be used in conjunction with transmission media. Transmission media participates in the transfer of information between storage media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise the bus 702. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infrared data communications.

[0121] It can involve various forms of media to transfer one or more sequences of one or more instructions to the processor 704 for execution. For example, the instructions can initially be carried on the disk or solid-state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 700 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on the bus 702. The bus 702 carries the data to the main memory 706, and the processor 704 retrieves and executes the instructions from the main memory 706. The instructions received by the main memory 706 can optionally be stored on the storage device 710 before or after being executed by the processor 704.

[0122] The computer system 700 also includes a communication interface 718 coupled to the bus 702. The communication interface 718 provides two-way data communication coupled to a network link 720 connected to a local network 722. For example, the communication interface 718 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, the communication interface 718 can be a Local Area Network (LAN) card to provide a data communication connection to a compatible LAN. A wireless connection can also be implemented. In any such implementation, the communication interface 718 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0123] The network link 720 generally provides data communication to other data devices through one or more networks. For example, the network link 720 can provide a connection to a main computer 724 or a data device operated by an Internet Service Provider (ISP) 726 through the local network 722. In turn, the ISP 726 provides data communication services through the global packet data communication network now commonly referred to as the "Internet" 728. Both the local network 722 and the Internet 728 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks and signals on the network link 720 and through the communication interface 718 (which carry digital data to and from the computer system 700) are example forms of transmission media.

[0124] The computer system 700 can send messages and receive data, including program code, through a network, the network link 720, and the communication interface 718. In the Internet example, the server 730 can transmit the request code of an application through the Internet 728, the ISP 726, the local network 722, and the communication interface 718.

[0125] The received code can be executed by the processor 704 upon receipt and / or stored in the storage device 710 or other non-volatile memory for later execution. Compute Nodes and Clusters

[0126] A compute node is a combination of one or more hardware processors, each sharing access to byte-addressable memory. Each hardware processor is electronically coupled to registers on the same chip as the hardware processor and is capable of executing instructions that reference a memory address in the addressable memory and cause the hardware processor to load the data at that memory address into any register. Additionally, the hardware processor can access its own separate dedicated memory that is not accessible to other processors. One or more hardware processors can operate under the control of the same operating system.

[0127] The hardware processor can include multiple core processors on the same chip, each core processor (“core”) being capable of separately executing machine code instructions in the same clock cycle as another core in the multiple cores. Each core processor can be electronically coupled to connect to scratchpad memory that is not accessible to any other core processor in the multi-core processor.

[0128] A cluster contains compute nodes, with each node communicating with each other via a network. Each node in the cluster can be coupled to a network card or network integrated circuit on the same board as the compute node. Network communication between any two nodes is via the network card or network integrated circuit on one node and the network card or network integrated circuit on the other node. The network can be configured to support remote direct memory access.

[0129] In the foregoing specification, implementations of the present invention have been described with reference to numerous specific details that may vary depending on the implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. The sole and exclusive indicator of the scope of the present invention and what the applicant desires to be the scope of the present invention is within the literal and equivalent scope of a set of claims issued in this application, in the specific form in which such claims are issued, including any subsequent amendments.

Claims

1. A computer-implemented method, the method comprising: Receiving first input music sequence data for a first portion of an original music signal; Generating first output music sequence data by a machine learning (ML) model, at least partially based on the first input music sequence data; While receiving subsequent input music sequence data corresponding to a subsequent portion of the original music signal that is temporally after the first portion of the original music signal, generating second output music sequence data for the generated music signal by the ML model, at least partially based on a specific portion of the first output music sequence data; Wherein the first output music sequence data of the generated music signal corresponds to a portion of the generated music signal that is temporally aligned with the subsequent portion of the original music signal.

2. The method of claim 1, further comprising: Receiving second input music sequence data corresponding to a second portion of the original music signal; Generating the second output music sequence data for the generated music signal by the ML model, at least partially based on the specific portion of the first output music sequence data and the second input music sequence data; Wherein the second input music sequence data corresponds to a second portion of the original music signal, the second portion being temporally after the first portion of the original music signal and before the subsequent portion of the original music signal.

3. The method of claim 1, further comprising: Determining one or more first music metric values of the first input music sequence data; Configuring the ML model to generate the second output music sequence data at least partially based on the one or more first music metric values.

4. The method of claim 3, wherein the first music metric value is a metric value of polyphony or intensity.

5. The method of claim 1, further comprising: Determining one or more first music metric values of the first input music sequence data; Configuring the ML model to generate the first output music sequence data at least partially based on the one or more first music metric values; Receiving second input music sequence data corresponding to a second portion of the original music signal; Determining one or more second music metric values of the second input music sequence data; Reconfiguring the ML model to generate the second output music sequence data at least partially based on the one or more second music metric values.

6. The method of claim 1, further comprising: Receiving a variance configuration metric value; Receiving second input music sequence data corresponding to a second portion of the original music signal; Generating the second output music sequence data for the generated music signal by the ML model, at least partially based on the specific portion of the first music output sequence data and the second input music sequence data; Wherein the specific portion of the first music output sequence data is a portion of the first music output sequence data by an amount of the variance configuration metric value.

7. The method of claim 1, further comprising: Receive the beats per minute (BPM) metric value of the generated music signal; Determine, at least in part, the bar time duration of the generated music signal based on the BPM metric value; Wherein the first part of the original music signal corresponds to a specific number of bar time durations.

8. The method according to claim 1, wherein the ML model is a large language model (LLM), the first input music sequence data is a first input sequence of music tokens, and the first output music sequence data is a first output sequence of music tokens, and the method further comprises: Receiving a first MIDI data sequence for the first part of the original music signal; Determining, at least in part, the first input sequence of music tokens based on the first MIDI data sequence; Generating, by the LLM model, the first output sequence of music tokens for the generated music signal, at least in part, based on the specific part of the first output music sequence data.

9. The method according to claim 8, wherein the LLM is trained, at least by adjusting the parameters of the LLM, to generate a music token output sequence given a music token input sequence, thereby reducing the error in generating a known token output sequence corresponding to the next token input sequence.

10. The method according to claim 8, further comprising: Determining the time delay between a first MIDI event and a second MIDI event in the first MIDI data sequence; Generating a music token for the time delay between the music token of the first MIDI event and the music token of the second MIDI event.

11. The method according to claim 8, further comprising: Determining one or more first music metric values of the first input music sequence data; Generating one or more music tokens indicating the one or more first music metric values in the first input sequence of music tokens.

12. One or more non-transitory computer-readable media storing an instruction set, wherein the instruction set includes instructions that, when executed by one or more processors, cause the following operations: Receiving first input music sequence data for a first part of an original music signal; Generating first output music sequence data, at least in part, based on the first input music sequence data by a machine learning (ML) model; When receiving subsequent input music sequence data corresponding to a subsequent part of the original music signal that is temporally after the first part of the original music signal, generating second output music sequence data for the generated music signal, at least in part, based on a specific part of the first output music sequence data by the ML model; Wherein the first output music sequence data of the generated music signal corresponds to the part of the generated music signal that is temporally aligned with the subsequent part of the original music signal.

13. A system comprising one or more processors and one or more storage media storing one or more computer programs executed by the one or more processors, wherein The one or more computer programs, when executed by the one or more processors, cause: Receiving first input music sequence data for a first portion of an original music signal; Generating first output music sequence data by a machine learning (ML) model, at least partially based on the first input music sequence data; When receiving subsequent input music sequence data corresponding to a subsequent portion of the original music signal that is temporally after the first portion of the original music signal, generating second output music sequence data for the generated music signal by the ML model, at least partially based on a specific portion of the first output music sequence data; Wherein the first output music sequence data of the generated music signal corresponds to a portion of the generated music signal that is temporally aligned with the subsequent portion of the original music signal.