System and method for generating media
Patent Information
- Application Number
- PCT/IL2026/050266
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2026-03-24
- Publication Date
- 2026-10-01
Smart Images

Figure IL2026050266_01102026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: P-633341-PC SYSTEM AND METHOD FOR GENERATING MEDIA CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of United States Provisional Patent Application No. 63 / 776,412, filed March 24, 2025, which is hereby incorporated by reference.FIELD OF THE INVENTION
[0002] The present invention relates to generation of media. More particularly, the present invention relates to systems and methods for generation of music using machine learning.BACKGROUND OF THE INVENTION
[0003] In the music industry, different musicians and ensembles produce a wide range of creative works, each with unique stylistic choices and decision-making processes. Attempting to replicate these nuanced, highly personal behaviors using machine learning (ML) can be challenging. Traditional ML approaches can struggle to balance specificity (e.g., capturing an individual musician’s style) with broad generalization (e.g., covering many possible musical contexts).
[0004] Some generative ML models that suggest “control” can drift away from the actual dataset’ s stylistic fidelity when heavily guided. Moreover, some ML architectures do not handle complex interactions between multiple musicians, nor do they integrate partial or inconsistent user inputs (for example, incomplete audio STEMs or loosely defined instructions), which can be akin to a realistic music production environment.
[0005] Neural network (NN) or connectionist approaches - using various deep learning methods such as convolutional neural networks (CNNs), transformers, or diffusion architectures - form the technical backbone of most modern ML systems. These networks typically require large-scale parallel computations on CPU or GPU hardware. However, merely deploying existing ML architectures has been insufficient for realistic emulation of real musicians’ micro-level decisions and audio.
[0006] Some artificial intelligence (Al)-driven music solutions can rely on static or fully pre- structured data. They cannot, for instance, take in partial or zero STEMs, layeredAttorney Docket No.: P-633341-PC instructions, and unique musician- specific nuance simultaneously - particularly when real-time or near-real-time feedback loops are desired.
[0007] Hence, there is a need for a system and method that can enable a flexible, multilayer input of audio and instructions, and / or adopt specialized training pipelines that separate general music understanding from musician- specific fine-tuning, and output new audio reflective of individual performance traits while allowing partial instructions and / or partial STEMs in a typical music production environment.SUMMARY OF THE INVENTION
[0008] An Al-driven framework to emulate and / or replicate the nuanced decision-making processes and unique stylistic choices of individual musicians, or, more general models capturing nuanced genres and styles, enabling the creation of virtual musicians and models that perform, adapt, and interact with musical data whether partially complete or complete, and handle user instructions.
[0009] There is this provided, in accordance with some embodiments of the invention, a method of generating audio, the method including: converting, by a processor, at least one audio condition to an audio spectrogram, applying, by the processor, a first machine learning (ML) algorithm to encode the spectrogram to a latent representation, masking, by the processor, at least one predefined portion of the latent representation, to generate at least one missing portion within the latent representation, applying, by the processor, a second ML algorithm to modify the at least one missing portion within the latent representation to create a complete latent representation, wherein the second ML algorithm includes architectural musical embeddings, decoding, by the processor, the complete latent representation into a new spectrogram using the first ML algorithm, and converting, by the processor, the new spectrogram to a waveform.
[0010] In some embodiments, a third ML algorithm may be applied to embed non-audio conditions to the latent representation. In some embodiments, the processor further receives at least one of: musical instructions, playback metadata, region definitions, Boolean parameters, and range parameters, and wherein modification of the at least one missing portion of the latent representation by the second ML algorithm is based on the information received by the processor. In some embodiments, the musical instructions include instructions for a target audio from a set of audio files, and wherein the maskingAttorney Docket No.: P-633341-PC is carried out at temporal locations correlated with said instructions for context-based audio generation.
[0011] In some embodiments, the at least one audio condition includes one or more of: musical instructions, playback metadata, and musical context. In some embodiments, the second ML algorithm modifies the at least one missing portion within the latent representation, based on the received at least one audio condition. In some embodiments, the second ML algorithm is trained using a dataset including at least one audio STEM and corresponding metadata, so as to interpret instructions when inferring missing instrumentation layers. In some embodiments, the architectural musical embeddings may be randomly omitted during training.
[0012] In some embodiments, the second ML algorithm applies at least one latent transformer-based diffusion mechanism to synthesize audio in accordance with the latent representation, at least one musical embedding, and the at least one audio condition. In some embodiments, audio channel information from encoding through decoding may be preserved to generate a multi-channel output waveform. In some embodiments, the at least one audio condition is empty. In some embodiments, the second ML algorithm employs a unidirectional mechanism.
[0013] There is this provided, in accordance with some embodiments of the invention, a system for generating audio, the system including: a dataset, including architectural musical embeddings, and a processor, configured to: convert at least one audio condition to an audio spectrogram, apply a first machine learning (ML) algorithm to encode the spectrogram to a latent representation, mask at least one predefined portion of the latent representation, to generate at least one missing portion within the latent representation, apply a second ML algorithm to modify the at least one missing portion within the latent representation to create a complete latent representation, wherein the second ML algorithm includes the architectural musical embeddings, decode the complete latent representation into a new spectrogram using the first ML algorithm, and convert the new spectrogram to a waveform.
[0014] In some embodiments, the processor is configured to apply a third ML algorithm to embed non-audio conditions to the latent representation. In some embodiments, the processor further receives at least one of: musical instructions, playback metadata, region definitions, boolean parameters, and range parameters, and wherein modification of theAttorney Docket No.: P-633341-PC at least one missing portion of the latent representation by the second ML algorithm is based on the information received by the processor. In some embodiments, the at least one audio condition includes one or more of: musical instructions, playback metadata, and musical context.
[0015] In some embodiments, the processor is configured to randomly omit the architectural musical embeddings during training. In some embodiments, the second ML algorithm applies at least one latent transformer-based diffusion mechanism to synthesize audio in accordance with the latent representation, at least one musical embedding, and the at least one audio condition. In some embodiments, the at least one audio condition is empty. In some embodiments, the second ML algorithm employs a unidirectional mechanism.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features and advantages thereof, may best be understood by reference to the following detailed description when read with the accompanied drawings. Embodiments of the invention are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate corresponding, analogous or similar elements, and in which:
[0017] Fig. 1 shows a block diagram of a computing device, according to some embodiments of the invention;
[0018] Fig. 2 shows a block diagram for a system for generating media, according to some embodiments of the invention; and
[0019] Figs. 3A-3B, show a flowchart for a method of generating media, according to some embodiments of the invention
[0020] Figs. 4A-4C, show a flowchart for data gathering for a predefined musician, according to some embodiments of the invention;
[0021] Figs. 5A-5C, show a flowchart for data gathering for a general musician, according to some embodiments of the invention;Attorney Docket No.: P-633341-PC
[0022] Figs. 6A-6D, show a flowchart for exemplary ML audio generation architecture, according to some embodiments of the invention;
[0023] Fig. 7, shows a flowchart for an encoder training loop, according to some embodiments of the invention;
[0024] Figs. 8A-8H, show a flowchart for a musician’s training loop, according to some embodiments of the invention; and
[0025] Fig. 9 shows a flowchart for a musician’s inference loop, according to some embodiments of the invention.
[0026] It will be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION OF THE INVENTION
[0027] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details.
[0028] In other instances, well-known methods, procedures, and components, modules, units and / or circuits have not been described in detail so as not to obscure the invention. Some features or elements described with respect to one embodiment may be combined with features or elements described with respect to other embodiments. For the sake of clarity, discussion of same or similar features or elements may not be repeated.
[0029] Although embodiments of the invention are not limited in this regard, discussions utilizing terms such as, for example, “processing”, “computing”, “calculating”, “determining”, “establishing”, “analyzing”, “checking”, or the like, may refer to operation(s) and / or process(es) of a computer, a computing platform, a computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer’s registers and / or memories into other data similarly represented as physical quantities within the computer’ s registersAttorney Docket No.: P-633341-PC and / or memories or other information non-transitory storage medium that may store instructions to perform operations and / or processes.
[0030] Although embodiments of the invention are not limited in this regard, the terms “plurality” and “a plurality” as used herein may include, for example, “multiple” or “two or more”. The terms “plurality” or “a plurality” may be used throughout the specification to describe two or more components, devices, elements, units, parameters, or the like. The term set when used herein may include one or more items.
[0031] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Additionally, some of the described method embodiments or elements thereof may occur or be performed simultaneously, at the same point in time, or concurrently.
[0032] Reference is made to Fig. 1, which is a block diagram of an example computing device, according to some embodiments of the invention. Computing device 100 may include a controller or processor 105 (e.g., a central processing unit processor (CPU), a chip or any suitable computing or computational device), an operating system 115, memory 120, executable code 125, storage 130, input devices 135 (e.g. a keyboard or touchscreen), and output devices 140 (e.g., a display), a communication unit 145 (e.g., a cellular transmitter or modem, a Wi-Fi communication unit, or the like) for communicating with remote devices via a communication network, such as, for example, the Internet or using a near field communication (NFC) sensor.
[0033] Controller 105 may be configured to execute program code to perform operations described herein. The system described herein may include one or more computing device(s) 100. For example, system 200 for generating audio may be, or may include computing device 100 or components thereof.
[0034] Operating system 115 may be or may include any code segment (e.g., one similar to executable code 125 described herein) designed and / or configured to perform tasks involving coordinating, scheduling, arbitrating, supervising, controlling or otherwise managing operation of computing device 100, for example, scheduling execution of software programs or enabling software programs or other modules or units to communicate.
[0035] Memory 120 may be or may include, for example, a Random Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAMAttorney Docket No.: P-633341-PC (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Memory 120 may be or may include a plurality of similar and / or different memory units. Memory 120 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM.
[0036] Executable code 125 may be any executable code, e.g., an application, a program, a process, task or script. Executable code 125 may be executed by controller 105 possibly under control of operating system 115. For example, executable code 125 may be a software application that performs methods as further described herein.
[0037] Although, for the sake of clarity, a single item of executable code 125 is shown in Fig. 1, a system according to embodiments of the invention may include a plurality of executable code segments similar to executable code 125 that may be stored into memory 120 and cause controller 105 to carry out methods described herein.
[0038] Storage 130 may be or may include, for example, a hard disk drive, a universal serial bus (USB) device or other suitable removable and / or fixed storage unit. In some embodiments, some of the components shown in Fig. 1 are omitted. For example, memory 120 may be a non-volatile memory having the storage capacity of storage 130. Accordingly, although shown as a separate component, storage 130 may be embedded or included in memory 120.
[0039] Input devices 135 may be or may include a keyboard, a touch screen or pad, a camera one or more sensors or any other or additional suitable input device. Any suitable number of input devices 135 may be operatively connected to computing device 100. Output devices 140 may include one or more displays or monitors and / or any other suitable output devices. Any suitable number of output devices 140 may be operatively connected to computing device 100.
[0040] Any applicable input / output (I / O) devices may be connected to computing device 100 as shown by blocks 135 and 140. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device or external hard drive may be included in input devices 135 and / or output devices 140.
[0041] Embodiments of the invention may include an article such as a computer or processor non-transitory readable medium, or a computer or processor non-transitoryAttorney Docket No.: P-633341-PC storage medium, such as for example a memory, a disk drive, or a USB flash memory, encoding, including or storing instructions, e.g., computer-executable instructions, which, when executed by a processor or controller, carry out methods disclosed herein. For example, an article may include a storage medium such as memory 120, computerexecutable instructions such as executable code 125 and a controller such as controller 105.
[0042] Such a non-transitory computer readable medium may be for example a memory, a disk drive, or a USB flash memory, encoding, including or storing instructions, e.g., computer-executable instructions, which when executed by a processor or controller, carry out methods disclosed herein.
[0043] The storage medium may include, but is not limited to, any type of disk including, semiconductor devices such as read-only memories (ROMs) and / or random- access memories (RAMs), flash memories, electrically erasable programmable read-only memories (EEPROMs) or any type of media suitable for storing electronic instructions, including programmable storage devices. For example, in some embodiments, memory 120 is a non-transitory machine-readable medium.
[0044] In some embodiments, a system may include or may be, for example, a personal computer, a desktop computer, a laptop computer, a workstation, a server computer, a network device, or any other suitable computing device.
[0045] A system according to embodiments of the invention may include components such as, but not limited to, a plurality of central processing units (CPUs), a plurality of graphics processing units (GPUs), or any other suitable multi-purpose or specific processors or controllers (e.g., controllers similar to controller 105), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units. A system may additionally include other suitable hardware components and / or software components.
[0046] As described hereinafter, at least some of the following terms may be used referring to the musical industry. A STEM can be a single audio file representing a track within a group of tracks, collectively representing the layers of a song. Master files can be finalized versions of songs, including all the STEMs within its audio content. Master files can be different from simply combining all the STEMs, as they typically go through additional musical processes that include specialized summation (SUM) functions, and / orAttorney Docket No.: P-633341-PC other static and dynamic manipulations. STEM files and / or master files may include stereo, 24-bit, 48kHz, WAV audio files.
[0047] These audio files, that when played produce sounds and / or music, may be “dry” (unprocessed) versions, or “wet” (processed) versions. The “dry” version may include raw, unprocessed sounds like a directly inputted guitar signal or a drum set recorded from close microphones without further processes. The “wet” version may include further processes and manipulations, based on the “dry” version.
[0048] Some data provided for each musician and / or musical piece may also include metadata, such as metadata files that provide detailed information about the recording and their musical instructions for each song within the dataset.
[0049] According to some embodiments, systems and methods are provided for enhanced generative framework to generate audio using machine learning (ML) algorithms, with a latent audio encoder-decoder, paired with a transformer-based diffusion and / or at least one musical embedding. ML embeddings may include numerical, high-dimensional vector representations of data (e.g., for text, images, audio) that capture semantic meaning and / or relationships, allowing models to process, compare, and cluster information.
[0050] Attention layers may be used to provide relationship information between the audio playback, musical instructions, metadata, and the target audio (e.g., the audio the model is aimed at recreating). Dynamic and / or layered input may include augmentation of partial or complete multi-STEM mixes, and / or metadata (including Booleans, range parameters, region definitions, region classification (e.g., harmony, melody, solo, etc.), and / or each region’s corresponding start and end timestamps).
[0051] In some embodiments, progressive specialization may be applied to create a hierarchical training process for ML algorithms that may encompass broad multiinstrument learning, and may be narrowed down to specific instrument groups, and / or focused on particular musicians. This approach may leverage randomized subsets of at least one audio STEM and / or layered metadata.
[0052] Emulation of musical micro decisions may enable a computing system to simulate the fine-grained thought process of a musician based on a corresponding musical context. Rather than merely generating static audio segments, embodiments of the system may simulate choices or actions by a musician at the level of phrasing, articulation, timing, and / or dynamics.Attorney Docket No.: P-633341-PC
[0053] Each note and / or expressive element in an output audio signal (e.g., a generated performance audio, and / or a generated STEM) may be shaped by prior musical content, forthcoming sections, and / or instructions in the metadata (e.g., prompts, user instructions, and / or metadata associated with training data and / or inference input). Thus, results that are not only consistent in style but also carry the “feel” of a human performer may be provided, complete with anticipations, subtle variations, and / or context-based decisions. The ML system may use transformer-based diffusion models, dedicated attention mechanisms, and / or latent audio representations to achieve the “feel” of the human performer, by integrating both temporal and harmonic cues to decide how each new moment unfolds.
[0054] Initially, an ML algorithm may be trained to learn general musical patterns (e.g., such as rhythmic, harmonic, melodic, structural, dynamic, and / or timbral patterns) across multiple genres, instruments, and / or contexts. Subsequent fine-tuning may be focused on instrument-specific nuances, such as drum techniques, or guitar articulations. The final training stage may be focused on an individual musician’s unique style, capturing microdecisions, habitual phrasing, and / or personalized effects, and / or, focused on a nuanced genre or style of music. For example, capturing based on training data associated with that musician, genre, and / or style (e.g., audio recordings, one or more STEMs, metadata and / or musical instructions). Using progressive specialization training, the system retains broad musical knowledge while adding increasingly specific layers of detail to emulate the decision-making process of a particular performer, genre, and / or style.
[0055] According to some embodiments, systems and methods are provided for at least one or more of the following: micro decision emulation, variable control mechanisms for user interaction, output flexibility to accommodate multiple modes of production (e.g., real-time streams or isolated takes), symbolic decisions and conditions that map musical parameters to discrete or continuous representations, real-time responses enabled by a unidirectional architecture, and / or inter-agent communication that may allow multiple virtual and / or real musicians, to interact in a shared musical environment. These elements may form a cohesive framework for generating adaptive, high-fidelity audio that is configured to mirror the intricate decision pathways a human musician may follow.
[0056] Symbolic decisions and / or conditions may transform a musician’s tendencies, instructions, and / or context into quantized representations. These may include barAttorney Docket No.: P-633341-PC position encodings, chord or key embeddings, region definitions (e.g., melody, ambience, solo), and various Boolean or range parameters. By using symbolic layers, the ML system may reduce complexity, focusing on how certain symbolic states (e.g., a chord progression or a strumming Boolean) correspond to audio patterns in latent space. Such layering may assists the ML system to better “understand” and replicate a musician’s style since it may handle higher-level structures, such as anticipating chord changes or reacting to sections labeled as “free,” where the human musician is permitted increased variability.
[0057] Reference is now made to Fig. 2, which shows a block diagram for a system 200 for generating audio, according to some embodiments of the invention. In Fig. 2, some hardware elements may be indicated by a solid line while some software elements may be indicated by a dashed line.
[0058] The system 200 may include a processor 201, in communication with a dataset 202 and with a server 203. The dataset 202 may include or store a set of audio files 204, with the target audio (e.g., as the ground truth) that may include the “wet” (e.g., processed version) and the “dry” (e.g., unprocessed version) and / or the multitrack version. The set of audio files 204 may include playback files with the at least one STEM of the playback audio and may include the master file as well.
[0059] The dataset 202 may include or store at least one condition 205 (e.g., a condition selected by the user). In some embodiments, the at least one condition 205 may include audio data, playback data, musical parameters and / or metadata. For example, the selected at least one condition 205 may be inputted to a ML algorithm for conditioning the algorithm to generate new audio.
[0060] Conditioning the ML algorithm may include guiding and / or controlling the output of the ML algorithm by feeding it specific, and / or additional input audio information, thereby transforming a general-purpose ML model into a specialized model for audio. For example, conditioning the algorithm may include conditioning on a specific label or following a specific style.
[0061] In some embodiments, the at least one condition 205 may be selected to be empty so that no selected input is provided to the system by the user. The ML algorithm may also be trained on audio files separately. For example, during training, the ML model may receive training examples that include audio target data and also “empty” conditions, so that the ML model does not need to rely on the condition in order to generate audio, butAttorney Docket No.: P-633341-PC rather be able to generate music even when receiving empty conditions. The model may receive a playback audio and generate compatible music, but may also generate audio from scratch, without a playback audio, since in training the ML algorithm may be trained with the playback condition as “empty”.
[0062] In some embodiments, the processor 201 may convert the at least one audio condition 205 to an audio spectrogram 206. Even if the at least one audio condition 205 is empty, the processor 201 may convert empty condition to an audio spectrogram 206 (e.g., without a playback audio).
[0063] The processor 201 may apply a first ML algorithm 207 to encode the spectrogram 206 to a latent representation 208. The latent representation may include a compressed, lower-dimensional encoding of high-dimensional data, capturing the essential, underlying structure of the high-dimensional data. In some embodiments, the first ML algorithm 207 may employ encoders, decoders, and / or enhancers that are trained using a specific musical dataset, with the output of the first ML algorithm 207 creating a unique codebook of symbols and latent spaces representing those musical datasets. Such training may allow for the emulation of the dataset’ s distinctive sound and behavior, thereby providing high personalization to the dataset’ s output.
[0064] In some embodiments, the processor 201 may mask at least one predefined portion of the latent representation 208 to determine at least one missing portion within the latent representation 208. The processor 201 may apply a second ML algorithm 209 to modify (or fill) the at least one missing portion within the latent representation 208 and to create a complete latent representation.
[0065] In some embodiments, the dataset 202 may include or store architectural musical embeddings 210. For example, the architectural musical embeddings may be predefined.
[0066] The second ML algorithm 209 may include musical concepts and / or theory (such as measures, keys, harmony, melody) that are represented as efficient architectural musical embeddings 210, aimed to reduce the resources needed while training the ML model to create an understanding of music as a concept within the weights of the ML model, and therefore focus more resources on the particular data’s semantic and phonetic / acoustic uniqueness. Thus, a more efficient training process may be achieved that requires less time to train, needs less data, and also may require “weaker” computing machines or uses less computing resources.Attorney Docket No.: P-633341-PC
[0067] In some embodiments, a third ML algorithm may be applied to embed non-audio conditions to the latent representation (such as tempo, time signature, region definitions, instruction templates, Boolean and range parameter sets).
[0068] The processor 201 may decode the complete latent representation into a new spectrogram 211 using at least one decoder of the first ML algorithm 207.
[0069] In some embodiments, the processor 201 may convert the new spectrogram 211 to a waveform 212, and / or send data corresponding to the waveform 212 to the server 203 to generate new audio 220 based on the waveform 212. For example, the processor 201 may send the generated waveform 212 to an audio playback device (e.g., a radio or a server based speaker) to play the new waveform.
[0070] A dedicated user interface may be employed to offer a spectrum of ways for users, such as producers, composers, or music hobbyists, to interact with a virtual musician that may be emulated by the system as the new waveform. A user may provide high-level stylistic suggestions (e.g., choosing musical region type and / or setting the Boolean groups to higher variance ranges) and / or detailed micro-guidance (e.g., referencing a specific musical part to emulate). The ML model may interpret these directives, drawing upon the personal style data of the emulated musician. The layered control framework may allow some users to specify nuances while other users may rely on more generic instructions. The ML system “fills in the gaps” by leveraging the musician’s learned tendencies, so if the inputs are vague, the ML model may still generate coherent musical content aligned with the musician’s style and provided context.
[0071] The control mechanism may allow dynamic interplay between user instructions and the musician’s style parameters. For instance, a producer may specify that the musician should use a distorted sound with moderate complexity. The ML system may accordingly interpret “distortion” as a Boolean (or set of Booleans) in the metadata, and “moderate complexity” as a range parameter that constrains how densely the musician plays. Even if the user does not address specific articulations, the system may autonomously determine when to add fills or vary dynamics, guided by both the user’s request and the musician’s historical data. This can ensure a truly interactive experience that mirrors working with an actual session player.
[0072] Output flexibility may provide support for continuous streams, discrete takes, and / or a hybrid approach to meet diverse production needs. In a real-time musical jam orAttorney Docket No.: P-633341-PC live performance scenario, the ML model may generate an uninterrupted audio stream, reacting to new input. In a musical studio context, the ML system may produce individual takes or segments, that may then be reassembled. These takes may be “stitched” together using an intelligent bridging process that ensures natural transitions, so producers may splice, layer, or choose the best segments just as they would with real recorded audio. Continuous mode may be especially useful for improvisational settings or performancebased installations, whereas isolated responses are preferred in typical music production workflows.
[0073] Real-time responses using a unidirectional approach may enable the ML system to continuously generate new audio without requiring extensive backward passes and / or re-computation. By maintaining only the relevant historical context and not needing future content, the model may respond on-the-fly. This may be useful for live applications, jam sessions, improvisational performances, and / or real-time collaboration with other Al or human musicians. The unidirectional design ensures that latency remains low and that the ML system may keep pace with typical musical performance standards, accommodating tempo changes, dynamic variations, and spontaneous instructions.
[0074] Such approach may be enabled by the ML system’s streaming capabilities, which only require knowledge of what has already occurred, allowing it to continuously emit new content. The encoder-decoder pipeline of the first ML algorithm, combined with a diffusion-based transformer structure, may produce partial outputs in alignment with realtime data. Because the system learned from a variety of real-world scenarios (including noise, effect chains, and incomplete metadata), it may remain robust in impromptu settings. Live band scenarios, rehearsal simulations, and / or remote collaborative sessions can all be facilitated by such consistent forward-only generation approach.
[0075] In some embodiments, inter-agent communication may facilitate multiple virtual musicians exchanging data in real time, either among themselves or with human users. Each virtual musician may interpret another musician’s outputs as a condition and / or “playback layer.” For example, if one musician is providing a lead guitar line, the other may adapt with a complementary bass groove, just as human ensemble members listen and respond to each other. This accordingly may extend to interactions with real performers, so that a human drummer and a virtual bassist coordinate on the fly.Attorney Docket No.: P-633341-PC Communication channels may treat each musician’s output as an evolving set of audio and symbolic cues, integrated into the cross-attention layers of the diffusion model.
[0076] Reference is now made to Figs. 3A-3B, which show a flowchart for a method of generating audio, according to some embodiments of the invention.
[0077] At least one audio condition may be received 301 (by a processor), for example selected by a user. The received at least one audio condition may be converted 302 (by the processor) to an audio spectrogram, and a first ML algorithm may be applied 303 (by the processor) to encode the spectrogram to a latent representation.
[0078] At least one predefined portion of the latent representation may be masked 304 (by the processor) to designate at least one missing portion within the latent representation, and a second ML algorithm may be applied 305 (by the processor) to modify the at least one missing portion within the latent representation to create a complete latent representation. The second ML algorithm may include musical concepts and / or theory that are represented as efficient architectural musical embeddings, aimed to create an understanding of music as a concept within the weights of the second ML algorithm, and therefore focus more computing resources on the particular data’s semantic and phonetic / acoustic uniqueness.
[0079] The complete latent representation may be decoded 306 (by the processor) into a new spectrogram using the first ML algorithm, for example using a decoder of the first ML algorithm. The new spectrogram may be converted 307 (by the processor) to a waveform.
[0080] In some embodiments, the waveform may be sent 308 to a dedicated server, in communication with the processor, to generate new audio based on the waveform. The generated new audio may be played on an external device, such as an audio playback device (e.g., a radio).
[0081] Reference is now made to Figs. 4A-4C, which show a flowchart for data gathering for a predefined musician, according to some embodiments of the invention.
[0082] Fig. 4A shows determining a playback distribution for use in recording sessions and dataset creation. For example, the system (or an operator) may select, for each song or segment, a distribution over possible subsets of available STEMs (e.g., drums and bass only, full band minus lead instrument, or sparse mixtures) and may further determine mixing weights, normalization targets, and / or augmentation parameters for each selectedAttorney Docket No.: P-633341-PC subset. In addition, the system (or an operator) may determine control parameters and design decisions for the session and / or dataset, such as region definitions, instruction templates, Boolean and range parameter sets, and target coverage levels for different musical roles or articulations, and may associate such controls with the selected STEM combinations. The musician output to be recorded may be specified (e.g., instrument, wet and / or dry capture, multi-take requirements, and / or expected role within the arrangement) so that each recording is aligned to the selected playback subset and associated controls. Such playback distribution and control selection may be configured to promote dataset diversity and robustness by ensuring that a predefined musician is recorded against varied musical contexts and incomplete combinations of accompaniment, thereby enabling subsequent training to support conditioning on partial, full, and / or empty playback inputs.
[0083] Fig. 4B shows an exemplary recording loop for collecting training data from a predefined musician. A playback corresponding to the selected STEM subset may be prepared (for example by removing the musician’s instrument from the mix and aligning the playback to a musical grid and / or timing reference), and a set of region-based instructions and control parameters may be generated and / or selected for the session. The predefined musician may then receive the playback and associated instructions and record one or more takes, and the resulting musician output (e.g., dry and / or wet STEMs, and / or other derived representations) may be captured and stored together with the playback conditions and instruction metadata, thereby enabling the recorded performance to be correlated with the conditioning context used during collection.
[0084] Fig. 4C shows converting the recorded outputs and associated controls into dataset conventions suitable for training. For example, the collected files and metadata may be cleaned to remove or correct errors and inconsistencies, musical instructions and playback metadata may be converted into a standardized machine-readable format (e.g., formatted JSON files), and the audio and metadata assets may be organized according to predefined directory structures, naming conventions, and indexing schemes.
[0085] Figs. 5A-5C show a flowchart for data gathering for a general musician, according to some embodiments of the invention.
[0086] Fig. 5A shows initiating a collection workflow for a general musician dataset. For example, the system may identify candidate source material (e.g., songs, projects, and / or multitrack sessions), determine which instruments and roles are to be represented, andAttorney Docket No.: P-633341-PC define baseline dataset targets such as tempo and time signature ranges, genre coverage, and minimum numbers of examples per instruction type and / or control-parameter combination. The system may further establish how playback subsets and instructions will be generated when a predefined musician identity is not used as the organizing key, thereby enabling collection across multiple performers, sessions, or sources while maintaining consistent conditioning conventions.
[0087] Fig. 5B shows an exemplary collection loop for a general musician dataset. For example, the system may iteratively select and / or generate playback subsets (including randomized combinations of STEMs), generate corresponding instruction and control parameter sets, and ingest musician outputs from recordings and / or existing multitrack sources. During each iteration, the system may validate alignment between playback, instructions, and the musician output (e.g., synchronizing to a timing grid and verifying region timestamps), and may record provenance information indicating the source session, performer, and / or context, thereby allowing aggregation of training examples across multiple musicians while preserving consistent conditioning labels.
[0088] Fig. 5C shows converting the collected general musician data into dataset conventions suitable for training. For example, the system may perform cleaning to remove or correct errors, convert musical instructions and playback metadata into a standardized machine-readable format (e.g., formatted JSON files), and organize audio and metadata assets into a predefined structure so that training pipelines can retrieve, for any example, the selected playback STEM subset, the aligned instruction / control parameters, and the corresponding musician output. Such conversion may further include normalization and / or validation checks to ensure consistent sampling rates, channel formats, and metadata schemas across heterogeneous sources.
[0089] Figs. 6A-6D, show a flowchart for an exemplary ML audio generation architecture, according to some embodiments of the invention.
[0090] These figures 6A-6D show how raw audio may be converted to spectrogram form, encoded to a latent representation, that may be masked or partially corrupted, and then reconstructed via a diffusion-based transformer. Metadata, instructions, and partial or full playback may be fed in parallel as conditions, guiding the generative process in both training and inference.Attorney Docket No.: P-633341-PC
[0091] From receiving a set of audio files and conditions, to converting the audio files into spectrograms, encoding the spectrograms into latent space, performing inpainting and / or prediction with diffusion-based transformers, and reconstructing the waveform, each step may be coordinated by metadata that includes arrangement data, instructions, and user parameters. The layered approach ensures that final results may be grounded in real musical logic.
[0092] A set of audio files and at least one predefined condition may be received by the system, typically a processor and / or server that orchestrates subsequent steps. The at least one condition may specify a style, a musician profile, and / or constraints such as “solo guitar with moderate fills.” The system begins by preprocessing the audio, combining or normalizing at least one STEM if necessary, and converting the audio to spectrogram form so they can be embedded in latent space.
[0093] At least one portion of the latent representation may be masked by the system, prompting the second ML algorithm to predict the missing segments. This can happen because the system may be, for example, asked to fill in an eight-bar gap or spontaneously create a bridging passage. The masked latent representation may be fed into a diffusion or transformer-based inference loop, which may use prior context to produce a plausible completion.
[0094] The predicted missing portions may be merged with the partial latent representation, creating a complete latent version of the audio that aligns with all instructions and constraints. Such comprehensive latent representation may then be decoded back into a spectrogram via the decoder of the first ML algorithm, which was trained to preserve high fidelity and nuance. By bridging partial instructions or incomplete context, the ML system may demonstrate a powerful generative ability that parallels real musicians who fill in the blanks when improvising or jamming.
[0095] The new spectrogram may be subsequently converted into a waveform. If the user requested a “wet” output, any relevant effect instructions are applied at this stage, possibly using convolution-based reverb or specialized distortion modules trained on the musician’s preferences. The final audio may then be streamed, stored, and / or delivered to a server for distribution, mixing, or other musical workflows. Because the system maintains a multi-stage pipeline, each transformation, from encoding to decoding, adheres to the constraints gleaned from the metadata and user instructions.Attorney Docket No.: P-633341-PC
[0096] In some embodiments, the processor converts the set of audio files to audio spectrograms, capturing frequency and time-domain properties. These spectrograms may serve as input to a first machine-learning encoder that projects the spectrograms into a latent space where temporal patterns, harmonic content, and other musical features become more tractable. The system may manipulate, modify and / or fill in missing segments within the latent space before ultimately decoding the representation back into an audio spectrogram.
[0097] In some embodiments, the system’s encoders, decoders, and diffusion components may be separately trained on general music data and musician- specific data. For instance, an encoder may learn general spectral features from at least one random STEM, while an associated decoder becomes adept at reconstructing or up-sampling to professional-quality audio. The second ML algorithm may be a specialized diffusion transformer to predict missing or masked portions in the latent space, bridging partial context to produce a coherent audio segment.
[0098] In some embodiments, the processor masks certain portions of the latent representation to replicate tasks like partial re-generation or inpainting. The second ML algorithm may then infer plausible content for these masked sections, effectively learning how a musician would complete a phrase if it were missing or how to respond to a new chord or drum fill. Such approach builds versatility, letting users remove existing sections in a track and ask the system to fill them in with the same style.
[0099] In some embodiments, after the ML model predicts the missing segments, those segments may be merged with the original latent representation to form a complete audio representation. The ML model decodes the representation back into a high-resolution spectrogram, and finally may convert that spectrogram into a temporal waveform. Such waveform may be output as raw PCM data, a WAV file, and / or a compressed format, depending on user needs.
[0100] In some embodiments, the processor uses a unidirectional and / or streaming approach to iteratively refine portions of the audio in real time, sending partial waveforms to a server and / or a local audio engine. Such design may be useful for live contexts where latency may be minimal, or for multi-musician setups where each agent’s output may be fed into the others’ next time-step. The real-time generation ensures that the ML system stays in sync with ongoing events or new instructions.Attorney Docket No.: P-633341-PC
[0101] According to some embodiments, once the ML system has generated a “dry” track, the system may apply post-processing effect chains, guided by either deterministic instructions or learning-based predictions. Such multi-stage pipeline may separate the musician’s fundamental performance decisions from the engineering or mixing steps. Users who prefer full creative control over mixing may bypass the post-processing, while those who want a polished product may let the system apply “wet” transformations that replicate their musician’s usual studio environment.
[0102] The ML audio generation architecture may be designed to process incomplete and / or deterministic instructions, bridging raw audio content with symbolic metadata. This entails feeding the model with a rich set of parallel data: spectrograms of combined playback, real-time instructions and / or constraints, symbolic region definitions (including region classification (e.g., harmony, melody, solo, etc.) and each region’s corresponding start and end timestamps), Boolean parameters for effects, and user- specified ranges for dynamics or complexity. The architecture may prioritize data alignment so that every second of music or instruction is synchronized in the model’s latent space. If certain instructions are omitted or incomplete, the system may infer plausible musical decisions based on the known context.
[0103] A transformer-based diffusion model may be employed to interpret the symbols and generate intermediate representations. Ultimately, the output may be either raw (“dry”) and / or enhanced (“wet”) audio, depending on whether effect chains and postprocessing instructions may be activated. Such design effectively captures microdecisions by layering the learned behaviors of the musician with instructions. The system may thus replicate not only the audio content but also the reasoning behind each choice, resulting in performances that may mirror human improvisation or carefully arranged playing.
[0104] One of the data layers may be a playback layer that encodes the audio context-within training often a random and / or partial subset of STEMs that are merged and normalized. Such playback may be translated into a spectrogram and embedded into latent space, enabling the system to reason about tempo, rhythm, harmonic structure, dynamics, and emotion as it would during real performance. While training, the playback layer may incorporate automated volume scaling, noise injection, and / or other augmentations toAttorney Docket No.: P-633341-PC improve robustness, ensuring that the model may handle real-world scenarios like changed mix levels and / or noisy backgrounds.
[0105] An output layer may capture the musician’s performance in a latent or symbolic representation. This layer may encode multiple variations of the same performance, including “dry,” “wet,” multi-track expansions, or even MIDI data, depending on the scenario. By synchronizing the musician’s recording with the playback, the model gains context about how the musician’s choices evolve throughout a piece of music. In training, the system may compares predicted outputs to recorded data, learning how a particular musician typically responds to a given input.
[0106] Metadata layers may provide high-level instructions and fine-grained parameters. They may include region definitions (melody, ambience, solo, etc.), chord progressions, Boolean toggles for specific articulations or effects, and range limits for dynamics or complexity. During inference, such metadata may be user-driven or partially generated from other modules. The layering may be useful as it gives the model a robust framework for “understanding” what is expected in each part of the music, from broad arrangement considerations to tight, localized instructions.
[0107] The instruction layer may embed rich temporal metadata that goes beyond simple chord or tempo data. For example, the instruction layer may specify roles in the arrangement (e.g., melody, counterpoint, rhythm), define partial freedom in “free” sections where the musician can explore more, or set Boolean flags for “clean” vs. “distorted” sound. Additionally, the instruction layer may organize time into discrete blocks, like bars or sections, each mapped to a unique combination of instructions. This information may be integrated with the playback to ensure the generated performance logically evolves with the structure of the song and the instructions.
[0108] Free regions in an arrangement may encourage the model to draw on the musician’s personal style rather than explicit instructions regarding musical role. In these areas, the system may maintain certain constraints, like an upper bound on complexity, but otherwise replicate how the musician would spontaneously play. These free sections may be used for capturing the unique imprint of an individual performer. They also demonstrate the generative system’s capacity for creativity within the boundaries of learned behavior, allowing for variability each time the model generates new content.Attorney Docket No.: P-633341-PC
[0109] Boolean groups may organize parameters that are represented as on / off or that may coexist simultaneously, reflecting the reality that a musician may use multiple articulation techniques in overlapping ways. For instance, one Boolean may represent “palm muting,” another “pick harmonics,” and another “legato,” all triggered at varying intensities. By modeling these states as Booleans, the system may remain flexible in how it combines them, ensuring that the musician’s typical layering of effects or techniques is represented accurately.
[0110] Range parameters may further refine the performance by establishing minimum and maximum allowable levels for dynamics, complexity, and / or fills. A wide dynamic range may mean the musician is free to play from pianissimo to fortissimo within a segment, whereas a narrower range constrains them to remain moderately loud. Complexity handles the density and intricacy of notes, while fills indicates how frequently the musician may insert transitions or embellishments. Because these parameters may not be fixed points but continuous intervals, the system can interpret them adaptively in response to the rest of the music.
[0111] The ranges may not be static, as they are interpreted in the broader musical context at each moment. If the system detects a high-energy lead-in from the playback, it may decide to push the upper edge of the dynamic range. Conversely, if the song transitions to a subdued section, the performance may shrink toward the lower bound. These continuous range parameters, combined with symbolic instructions, replicate the nuanced interplay of a real musician reacting to both past and upcoming musical events.
[0112] In some embodiments, the ML system incorporates deep temporal correlations to emulate the forward-looking strategy a musician uses. The knowledge of “what came before” shapes immediate responses, but the system also accounts for upcoming changes indicated by arrangement metadata. For instance, if a big chorus is about to start, the model may escalate volume or complexity just prior to that. Such forward anticipation is integral to realistic emulation and a key advantage of using advanced cross-attention or transformer-based approaches.
[0113] The dynamic ranges, Boolean groupings, and regions collectively may define a field of possibilities that the musician may explore. Within each bar and / or time slice, these constraints shape the generation process, ensuring consistency with the musician’s established style. Rather than producing random changes, the system may justify eachAttorney Docket No.: P-633341-PC note’s amplitude, timing, and / or articulation by referencing the instructions and the musician’s typical tendencies. The layered structure thus captures not only the final audio but the underlying reasoning, bridging explicit instructions and emergent musical intelligence.
[0114] The musical representations used by the ML system may include bar encodings, chord and key embeddings, region definitions, and more. These representations may be used for aligning the system’s generative outputs to the structural grid of the piece. While the playback layer encodes raw audio, the symbolic layer may parse each bar and / or chord change into a discrete token that the system can interpret. In combination, these encodings may provide a holistic (or architectural) view of the musical environment in which the musician is operating.
[0115] The musical representations may also include bar-level positional signals, for example implemented as a saw wave or sinusoidal function that resets at bar boundaries. Such musical representations may assist the model to keep track of exactly where it is in a measure, aligning potential fills or transitions with typical musical timing. When combined with chord and / or key embeddings, the system may gain a robust sense of harmony at each point, potentially controlling which pitches or intervals are used for melodic lines.
[0116] The musical representations may also include keys projected onto a circular representation that captures the closeness of related keys. Minor keys may be rotated on the complex plane to align with relative majors, allowing the model to interpret them in a consistent manner. Chords may be updated at their change points, like every measure or partial measure, to keep harmonic context fresh. These chord embeddings may also track chord quality (major, minor, sevenths, and / or extended chords), giving the system fine-grained input for melodic or harmonic improvisation.
[0117] The musical representations may also include “energy graphs” or “energy representations”, where the system marks strong downbeats, accents, or band “hits.” These points may assist shape the generated performance, ensuring that the musician accentuates or syncs with important rhythmic events. For instance, a drummer may naturally place a fill leading up to a marked accent in the next bar, or a guitar part may add emphasis where the energy graph spikes.Attorney Docket No.: P-633341-PC
[0118] Additional contexts such as tempo, time signature, arrangement sections, or explicit user-labeled genre may all be integrated. Tempo may be globally uniform or vary measure-by-measure (rubato or tempo ramp). The time signature could be something common like 4 / 4 or unusual like 7 / 8, letting the system adapt. Genre tags (e.g., “rock,” “jazz,” or “pop”) may inform broad stylistic boundaries, ensuring the system adopts relevant tonalities, instrument choices, or rhythmic grooves. Arrangement sections (e.g., verse, chorus, and bridge) may give the system cues on how to transition from one thematic idea to another.
[0119] In some embodiments, the system may be trained to emulate a particular musician by combining broad music knowledge with that musician’s personal recordings. The musician’s data may include at least one “dry” STEM recorded in a controlled environment alongside detailed instructions about how they intended to play. Because these STEMs may be unprocessed, the system can learn the musician’s baseline sound without artificial reverb, compression, or distortion although it may later reintroduce these via “wet” effect chains if desired.
[0120] A user interface (UI) may allow users to upload base audio, specify regions, define instructions, and / or select or fine-tune a musician profile. A user may also route audio from other virtual musicians into the system so it may treat those signals as a playback layer. In production scenarios, such audio routing may be used to coordinate an entire band of Al-driven performers. In a live context, this may be used to jam with real musicians who feed the system real-time audio.
[0121] Referring now to Fig. 7, that shows a flowchart for an encoder training loop, according to some embodiments.
[0122] Fig. 7 shows an encoder training loop for training at least one encoder and at least one decoder to map audio representations between spectrogram space and a latent representation. Training examples may include spectrograms derived from one or more STEMs and / or playback mixtures, and the encoder may be trained to produce a compact latent code from which the decoder reconstructs a target spectrogram and / or waveform with high fidelity. A reconstruction loss (and, in some embodiments, additional perceptual and / or adversarial losses) may be calculated between the reconstructed output and a corresponding target, and the encoder-decoder parameters may be updatedAttorney Docket No.: P-633341-PC iteratively until the latent representation and codebook capture sufficiently expressive musical structure for subsequent diffusion-based generation and inpainting.
[0123] Figs. 8A-8H show a flowchart for a musician’s training loop, according to some embodiments.
[0124] In some embodiments, the training process may be organized in three stages: general music understanding, specialized musician training, and musician- specific fine-tuning. In the first stage, the system may use massive collections of recorded at least one STEM in random subsets to develop a broad correlation-based music model. This establishes general knowledge of timing, harmony, and instrumentation. In the second stage, the system may refine the system’s understanding for a narrower set, such as a particular instrument family (e.g., guitars), allowing it to capture more detailed stylings. Finally, the model may be trained on a specific musician’s recordings, capturing their unique dynamic range preferences, favorite articulation Booleans, and typical approach to fills and / or solos.
[0125] In the first stage, the model may be trained using large sets of STEMs from many genres, instruments, and styles. The system may randomly pick partial segments of a song’s at least one STEM, sum them, and feed them into the model as a single playback track. This may foster robust correlation modeling, teaching the system to handle incomplete data and / or partial references. By intentionally limiting certain STEMs or sections, the model learns to fill in missing pieces, effectively performing an inpainting task on the audio domain.
[0126] A random generator may select random start times, subsets of instruments, and / or time lengths (e.g., up to 30 seconds) to further diversify training. The resulting subset may be typically normalized so that peak levels stay within range, avoiding clipping. The system may then be asked to reconstruct and / or predict how a typical musician or instrument group would respond. These tasks may reinforce an internal representation of how different tracks interlock and how musical roles shift throughout a piece.
[0127] During general training, some segments of the output layer may be masked or corrupted, requiring the model to predict and / or generate plausible musical content. This may foster a powerful form of inpainting that is later relevant when users want to remove or replace certain notes in a final composition. The system thus gains the capability to fill in short or long gaps with stylistically appropriate audio.Attorney Docket No.: P-633341-PC
[0128] In a second stage, specialized training may be performed for specific instrument groups, such as drums, keyboards, and / or guitars. The essential parameters learned in the first stage are preserved, serving as the foundation. New or adapter layers may then refine the model for the intricacies of that instrument family. For instance, a drum-specific model may learn the typical velocity distribution on snare hits or the interplay of kick and hi-hat. A guitar- specific model may learn typical chord voicings, picking styles, or how a guitarist transitions from clean to distortion.
[0129] In a third stage of further refinement, the system may apply musician- specific fine-tuning on top of the instrument specialization. Here, the system may use real recordings from an individual musician, capturing their signature flourishes, effect preferences, and dynamic shaping. Additional Booleans or range parameters may be introduced for particular personal quirks, like a preference for “fretboard tapping” or “aggressive palm, muting” in certain sections. Such layering may ensure the model’s knowledge of general guitar playing is not overwritten but is specialized to reflect the personal style of the chosen musician.
[0130] ‘Dry’ audio may be valuable for training because it captures the unprocessed essence of the musician’s performance. The system may then generate the “wet” chain if the user wants to replicate the final polished studio sound. This may involve reverb, echo, or specialized IR convolution. By separating dry from wet in the dataset, the model learns both the raw performance style and how it may typically be engineered, giving end-users the freedom to generate fully processed tracks and / or the raw recordings for further mixing.
[0131] According to some embodiments, noise augmentation may be applied to the playback at least one STEM, adding room noise, vinyl static, or hiss. Doing so prevents the system from overfitting to perfectly clean audio and allows it to remain robust when dealing with real-world or user-supplied recordings. Compression and / or equalization changes may also be randomly applied, teaching the system that a musician’s response does not drastically change when the volume or tonal shape of the playback shifts slightly.
[0132] Metadata augmentation can broaden the system’s adaptive capacity by omitting certain fields and / or introducing new Booleans mid-training. For example, half of the training examples may lack chord information, forcing the system to rely on the audio itself. In other examples, new effect Booleans could be randomly toggled on, so theAttorney Docket No.: P-633341-PC system learns to interpret unanticipated instructions gracefully. This may ensure the model remains flexible and doesn’t become dependent on a specific, rigid metadata schema.
[0133] Textual (or context) augmentation may teach the system to interpret synonyms or paraphrased instructions. For instance, “gentle groove” may appear as “soft pattern” or “light approach” in different training samples. By decoupling the underlying meaning from exact wording, the system becomes better at understanding user instructions that vary in phrasing or language style. This also aligns with the broader aim of letting novices or experts communicate with the system in natural language.
[0134] In some embodiments, a multi-stage training workflow may be introduced. For general music training, the model may get randomized subsets of STEMs to learn broad correlation modeling across instruments, time, and genres. It encodes partial or incomplete audio into latent representations, forcing robust music understanding.
[0135] For instrument-specific training, the model may be fine-tuned for particular instrument families (e.g., guitars, drums, keyboards), preserving prior general knowledge while adding specialized layers for timbral and stylistic nuance typical of that instrument group.
[0136] For musician- specific training, a further fine-tuning layer may capture the personal style and micro-decision-making unique to a specific musician. Boolean toggles (e.g., “clean,” “distorted,” “slap,” “picking”), range parameters (e.g., dynamic, complexity, fills), and instructions can then be integrated to guide the model in replicating that musician’s typical response under a given context.
[0137] Fig. 9 shows a flowchart for a musician’s inference loop, according to some embodiments.
[0138] During inference (media generation), a flexible input (an unfinished piece that may have any number of STEMs from zero to many, partial and / or detailed instructions) may be converted to a latent representation. A multi-layer instruction set (including metadata such as chord changes, arrangement markers, region definitions, or Boolean groups) may be combined with a musician- specific layer. The system receives or “listens to” these inputs, simulating how real musicians adapt and feedback on each other’s parts. The final output may be decoded into new audio, delivering either a dry waveform and / or a processed “wet” waveform consistent with the specified musician’s sonic style.Attorney Docket No.: P-633341-PC
[0139] Such architecture addresses prior limitations by allowing partial and / or zero STEMs for “starting from scratch” compositions, handling complex instruction sets that may be partial, micro-detailed, or omitted, and providing a robust feedback loop among multiple “virtual musicians”. The architecture may also maintain resource efficiency through latent encoding, diffusion-based transformers, and cross-attention layers dedicated to multiple data streams. Thus, producers, engineers, and other users may gain an advanced tool for generating high-quality, stylistically accurate musical content, harnessing the distinctive approaches of real or hypothetical musicians while preserving full or partial control over the final audio.
[0140] The deterministic instructions decoded alongside the dry audio may specify effect types, reverb levels, and / or other processing details. For instance, if the user toggled “overdrive” in the Boolean group, the system’s final stage may generate the signal with a distortion profile that matches the musician’s typical pedal.
[0141] The ‘dry’ audio may also serves an essential purpose in the training loop, where it acts as a baseline. By learning from multiple examples of dry performance in tandem with the corresponding wet outputs, the system may master how a musician transitions from raw input to finished track, such learning, in turn, may create a more flexible environment in which users can request any combination of partial dryness or specialized wetness.
[0142] In some embodiments, users may create an account within a platform that hosts multiple virtual musicians. The users may select a musician profile and / or train their own from scratch, upload base tracks, and define instructions using a user interface. Each instruction may be anchored to a region in time or to a textual anchor like “melody” or “solo.” The system may then process all relevant data and outputs the corresponding performance. Over time, the system may refine or update the musician’s model as more training data is added, continuously improving the fidelity of the emulation.
[0143] In some embodiments, users may create multi-track projects by subdividing songs into “regions” that carry distinct instructions. Virtual musicians may interpret these region-based instructions to yield cohesive multi-STEM outputs. The user may further manipulate and / or combine these outputs, highlight specific takes, and finalize the arrangement in a manner analogous to a digital audio workstation (DAW) environment,Attorney Docket No.: P-633341-PC only here, each track may be performed by an Al musician that understands context and personal style.
[0144] The process of dataset creation for training musicians’ models may include capturing “dry” audio in a controlled environment. Musicians may be asked to record themselves playing along to reference playbacks while viewing or listening to specific instructions. For each take, these instructions may specify certain dynamics, effects toggles, and / or arrangement roles. The raw STEMs-along with metadata describing how the musician responded, may be stored in a structured dataset.
[0145] In some embodiments, skilled musicians, may be selected for their unique playing styles, tonal preferences, and interpretive range. By collecting documented recordings from diverse performers, the system gains a broad palette of real-world data. Each musician’s dataset may include multiple genres, tempos, and / or time signatures, ensuring comprehensive coverage.
[0146] To capture the ‘dry’ audio, the environment may be kept free from extraneous hardware or software effects. A guitarist, for instance, may record a direct signal from their instrument’s pickup, or a drummer may record using close mics and minimal outboard processing. This may ensure that the recorded audio strictly represents the musician’s technique. During performance, the musician hears the playback tracks to stay synchronized, and they see the instructions that define the roles, Boolean toggles, or dynamic targets. Such instructions may guide the musician’s real decisions, effectively mirroring how a user may instruct the virtual version.
[0147] During recording sessions, musicians may wear headphones so they can hear the backing track and any click or metronome. This may ensure accurate timing and consistent volume reference. All the details-like which region is being played, or which Booleans are active-get logged in the metadata.
[0148] Alongside playback, a dynamic metadata script may be displayed, reminding the musician to switch from clean to distorted or to apply a subtle fill at the end of a measure. Because these instructions are not always rigid, the musician may interpret them in personally stylistic manners. Such instructions may form the kind of variation the system needs to learn how a real musician thinks in response to open-ended requests.
[0149] Equipped with the playback and instructions, the musician may perform multiple takes, capturing all possible permutations of instructions in short or long segments. TheAttorney Docket No.: P-633341-PC resulting at least one STEM becomes part of the training data. The system may then correlate each segment’s raw audio with the instructions, embedding the unique context that led to each performance choice. This may assist the final model replicate not just the outcome but the process of decision-making.
[0150] Concurrent with the audio recording, the metadata may be documented. This includes timing references, instructions, Boolean states, dynamic and / or complexity ranges, and any relevant textual descriptors. If the musician spontaneously deviates from the instructions in an artistic manner, that too may be logged for the model to learn from. The goal is to compile a robust, multi-dimensional snapshot of each performance.
[0151] Post-recording, minimal edits may be applied-such as trimming silence and / or normalizing levels-before storing the files. Because the system needs the data to remain as close to the raw playing as possible, the use of equalization, compression, or reverb is avoided, preserving the authenticity of the musician’s real output. In contrast, “wet” versions may be recorded in parallel for training on post-processing styles.
[0152] Such process may be repeated across many songs, styles, and / or sessions to ensure a sufficiently large dataset. Over time, multiple audio segments may be collected, each labeled with thorough metadata. This may enable the system to generalize not only across different songs but across various instructions and conditions. The curated dataset may form the backbone for a highly specialized musician model that captures personal style down to subtle dynamic swells and signature articulations.
[0153] According to some embodiments, advantages of this invention include the ability to operate primarily in latent space, drastically accelerating generation speeds without sacrificing audio fidelity. Layered temporal and contextual cues-like bar-level chord embeddings and region-based instructions-help the system maintain structural coherence.
[0154] Randomized subsets of STEMs may expand the data’s coverage, fostering robust correlation modeling. Deep temporal correlations ensure each note is influenced by prior buildup, future transitions, and / or real-time constraints, yielding outputs that feel genuinely alive.
[0155] While certain features of the invention have been described in detail, numerous modifications, substitutions, and changes are possible. For instance, different neural network architectures could be substituted without deviating from the core layered approach. The concept of Booleans, range parameters, and partial instructions may beAttorney Docket No.: P-633341-PC extended to any domain beyond music where micro-decision processes matter. By capturing the reasoning behind each creative step, the invention opens new possibilities for real-time collaboration, advanced music production, and the emulation of highly personalized artistic styles. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes.
[0156] Various embodiments have been presented. Each of these embodiments may of course include features from other embodiments presented, and embodiments not specifically described may include various features described herein.
Claims
Attorney Docket No.: P-633341-PC CLAIMS1. A method of generating audio, the method comprising:converting, by a processor, at least one audio condition to an audio spectrogram; applying, by the processor, a first machine learning (ML) algorithm to encode the spectrogram to a latent representation;masking, by the processor, at least one predefined portion of the latent representation, to generate at least one missing portion within the latent representation;applying, by the processor, a second ML algorithm to modify the at least one missing portion within the latent representation to create a complete latent representation, wherein the second ML algorithm comprises architectural musical embeddings;decoding, by the processor, the complete latent representation into a new spectrogram using the first ML algorithm; andconverting, by the processor, the new spectrogram to a waveform.
2. The method of claim 1, further comprising applying, by the processor, a third ML algorithm to embed non-audio conditions to the latent representation.
3. The method of claim 1, wherein the processor further receives at least one of: musical instructions, playback metadata, region definitions, Boolean parameters, and range parameters, and wherein modification of the at least one missing portion of the latent representation by the second ML algorithm is based on the information received by the processor.
4. The method of claim 3, wherein the musical instructions comprise instructions for a target audio from a set of audio files, and wherein the masking is carried out at temporal locations correlated with said instructions for context-based audio generation.
5. The method of claim 1, wherein the at least one condition comprises one or more of:musical instructions, playback metadata, and musical context.Attorney Docket No.: P-633341-PC6. The method of claim 1, wherein the second ML algorithm modifies the at least one missing portion within the latent representation, based on the received at least one condition.
7. The method of claim 1, wherein the second ML algorithm is trained using a dataset comprising at least one audio STEM and corresponding metadata, so as to interpret instructions when inferring missing instrumentation layers.
8. The method of claim 7, further comprising randomly omitting the architectural musical embeddings during training.
9. The method of claim 1, wherein the second ML algorithm applies at least one latent transformer-based diffusion mechanism to synthesize audio in accordance with the latent representation, at least one musical embedding, and the at least one condition.
10. The method of claim 1, further comprising preserving audio channel information from encoding through decoding to generate a multi-channel output waveform.
11. The method of claim 1, wherein the at least one condition is empty.
12. The method of claim 1, wherein the second ML algorithm employs a unidirectional mechanism.
13. A system for generating audio, the system comprising:a dataset, comprising architectural musical embeddings; anda processor, configured to:convert at least one audio condition to an audio spectrogram;apply a first machine learning (ML) algorithm to encode the spectrogram to a latent representation;mask at least one predefined portion of the latent representation, to generate at least one missing portion within the latent representation;Attorney Docket No.: P-633341-PC apply a second ML algorithm to modify the at least one missing portion within the latent representation to create a complete latent representation, wherein the second ML algorithm comprises the architectural musical embeddings;decode the complete latent representation into a new spectrogram using the first ML algorithm; andconvert the new spectrogram to a waveform.
14. The system of claim 13, wherein the processor is configured to apply a third ML algorithm to embed non-audio conditions to the latent representation.
15. The system of claim 13, wherein the processor further receives at least one of: musical instructions, playback metadata, region definitions, boolean parameters, and range parameters, and wherein modification of the at least one missing portion of the latent representation by the second ML algorithm is based on the information received by the processor.
16. The system of claim 13, wherein the at least one condition comprises one or more of:musical instructions, playback metadata, and musical context.
17. The system of claim 13, wherein the processor is configured to randomly omit the architectural musical embeddings during training.
18. The system of claim 13, wherein the second ML algorithm applies at least one latent transformer-based diffusion mechanism to synthesize audio in accordance with the latent representation, at least one musical embedding, and the at least one condition.
19. The system of claim 13, wherein the at least one condition is empty.
20. The system of claim 13, wherein the second ML algorithm employs a unidirectional mechanism.