Systems and methods of real-time musical accompaniment using artificial intelligence
The AI model system addresses the challenge of real-time synchronization with live music by training to track and generate virtual accompaniment, improving live performances and education with synchronized virtual instruments and vocalists.
Patent Information
- Application Number
- PCT/CA2025/050730
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-05-23
- Publication Date
- 2025-12-04
AI Technical Summary
Existing artificial intelligence models fail to generate virtual vocalists and virtual musical instruments in real-time synchronization with live music audio streams, limiting their application in live performances and music education.
A system and method for training an AI model to track live vocalists and instruments in real-time, generating synchronized virtual accompaniment based on the genre of the live music stream, using a computing device and server infrastructure for processing and storage, with features for customization and post-processing.
Enables real-time generation of synchronized virtual accompaniment with minimal latency, enhancing live performances and music education through AI-generated virtual backup vocalists and instruments.
Smart Images

Figure CA2025050730_04122025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS OF REAL-TIME MUSICAL ACCOMPANIMENTUSINGARTIFICIAL INTELLIGENCETECHNICAL FIELD
[0001] The present disclosure is in the field of computer technology, and in particular the field of artificial intelligence in music.BACKGROUND
[0002] In recent years artificial intelligence (Al) models have been trained to generate accompanying virtual vocalists and accompanying virtual musical instruments in many musical genres through text prompts and other means of input such as keys on a keyboard, buttons on a touchscreen or buttons and dials on specialized input controllers. Purely generative Al models do not track live music audio streams comprised of live vocalists and live musical instruments in real time to generate virtual vocalists and virtual musical instruments in synchronization with the live music audio stream and in a style based on the genre of the live music audio stream. A real-time Al model for live music would allow musicians to give live performances with realtime virtual backup vocalists and real-time virtual musical instruments in select musical genres, provide music students with a way to improve their music skills, and provide a simplified way to create social media content.
[0003] There exists a continuing desire to advance and improve technology related to real-time musical accompaniment using artificial intelligence.SUMMARY
[0004] Systems, methods and techniques are disclosed for training an Al model on datasets of music audio files comprised of select genres of music to track vocalists and musical instruments, to provide Al model inference on a live music audio stream to track live vocalists and live musical instruments in real time to generate accompanying virtual vocalists and virtual musical instruments in a simulated music audio stream in synchronization with the live music audio stream and in a style based on the genre of the live music audio stream, a style based on a genre selected by the Al model, or a style based on a genre or sheet music selected by a user of the Al model. A system for Al model training and inference may include a computing device comprised of a processor and communicatively coupled memory. The computing device may be further comprised of a display and one or more user input devices, wireless and wired communication interfaces, audio input / output ports and devices, and data input / output ports. A server, comprised of a processor and communicatively couple memory, and communicatively coupled to the computing device through a communications network, may provide computing and data storage services to the computing device. The system may further comprise the Al model and other applications stored in the memory of the computing device or the memory of server, or combinations of the memory of computing device and the memory of the sever.
[0005] In some embodiments, the simulated audio stream may be played through wired or wireless headphones, earphones, or loudspeakers or combinations of wired and wireless headphones, earphones, and loudspeakers to accompany the live music audio stream in real time.
[0006] In some embodiments, the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream may be transmitted over a communications network.
[0007] In some embodiments, the Al model may be further trainable by an Al trainer or a user of a computing device running the Al model, by vocalists, and by musicians after training on the initial datasets is completed, to provide Al inferences for particular musical compositions to suit the styles and tastes of the user, the vocalists, and the musicians.
[0008] In some embodiments, the Al model may be trained with sheet music in addition to music audio fdes and may be trained to generate sheet music for tracked live music audio streams and simulated music audio streams and store the generated sheet music in the memoryof the computing device or the memory of the server or combinations of the memory of the computing device and the memory of the server.
[0009] In some embodiments, the Al model may store a live music audio stream and the simulated music audio stream in one combined music audio fde or as separate music audio fdes in the memory of the computing device or the memory of the server or combinations of the memory of the computing device and the memory of the server.
[0010] In some embodiments, the Al model may provide post-processing features to add a musical introduction and conclusion to the stored music audio fdes, allow the live vocalists and live musical instruments to be tuned or replaced, and provide options for additional or replacement simulated music audio streams.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the accompanying drawings, which illustrate one or more example embodiments,
[0012] FIGURE 1 shows a method for synchronizing with a live music audio stream and generating a simulated music audio stream in real time in synchronization with the live music audio stream and in the style of the genre of the live music audio stream; and
[0013] FIGURE 2 shows system diagram with a vocalist, wireless headphones, and a computing device (smartphone); and
[0014] FIGURE 3 shows system diagram with a vocalist, wireless headphones, a musical instrument (guitar), a microphone, a computing device (tablet computer) and a server.DETAILED DESCRIPTION
[0015] The terms computer program, application, applet, app, or script, as used in the present disclosure, refer to a set of instructions executable by one or more computer processors. A computer program, application, applet, app, or script may be standalone, may run under a computer operating system, or may be integrated within other computer programs, applications, applets, apps, scripts, or systems, such as a computer operating system. A computer program, application, applet, app, or script may be comprised of one or more routines.
[0016] A computing device or computer, in the context of the present disclosure, refers to a device having a processor and memory communicatively coupled to the processor and can run computer programs, applications, applets, apps and scripts. A computing device may optionally include a display or the ability to drive one or more displays, some method of user input, such as a keyboard, discrete buttons, mouse or touchscreen, data I / O ports, audio I / O ports, one or more microphones, one or more speakers, wireless and wired communication interfaces such as but not limited to Internet, Bluetooth, Wi-Fi, and Cellular Network, and may be portable or non-portable. Examples of computing devices include, without limitation, desktop computers, embedded computers, servers, laptops, tablets, smartphones, smartwatches, virtual reality (VR) headsets and augmented reality (AR) headsets.
[0017] A server, cloud computer, cloud server, Internet server, Internet cloud server, local server, or remote server, in the context of the present disclosure, all refer to a computer that provides computing, data and data storage services, and may provide software license registration services, to other computers, through a network / Intemet connection. A server, referred to in the singular form in the present disclosure, may be comprised of a one or more computers in one location or may be comprised of more than one computer in a distributed computer system. A server may be implemented as a virtual machine running on a cloud / utility computing service.
[0018] An Al model is a computer program that may be trained on representative datasets with supervised, reinforced, and unsupervised learning to analyze new datasets to find patterns, make predictions, draw conclusions, and conduct actions with or without human intervention. The analysis of new datasets is referred to as inference. Al model parameters are configuration variables that are internal to the Al model and are learned from historical training data.
[0019] A sound is a vibration that propagates as an acoustic wave through air and may be heard by a person or animal. Sounds may be converted to analog electrical signals by a microphone and converted to digital form by an analog to digital converter (ADC) with a sampling rate sufficient to avoid aliasing and with a bit depth sufficient to avoid audible quantization noise and may processed by a program such as an Al model running on a computing device. Sampled and processed sounds represented in digital form and sounds generated by a computer program such as an Al model represented in digital form may be converted to analog form by a digital to analog converter (DAC), amplified and played through a speaker to produce sounds that may be heard by a person or animal.
[0020] Music is the arrangement of sounds to create one of or a combination of two or more of form, harmony, melody, and rhythm.
[0021] Live music is music that is performed in real time and may be comprised of audio and video components. Audio-only live music that is converted to digital form in real time is referred to as a live music audio stream.
[0022] A singer or vocalist is a person making musical sounds by voice to sing a song: a set of words, also called lyrics, set to music. A vocalist in a live music audio stream is referred to as a live vocalist. In some forms of music, singing is done alongside music produced by musical instruments. Vocal music is music comprised entirely of one or more vocalists making musical sounds. A virtual vocalist is a simulated vocalist generated by an Al model.
[0023] A physical musical instrument may be played by a vocalist or a musician. A virtual musical instrument is a simulated musical instrument generated by an Al model or another computer program. Physical musical instrument types may include percussive, woodwind, string, brass, and keyboard. Examples include but are not limited to drums, flutes, guitars, trumpets, and pianos. Synthesizers are electronic musical instruments played physically by musicians with the sound produced by the musical instrument picked up with a microphone or electronic sensor and electronically processed or altered, or the sound may be entirely synthetic, controlled by the musician through a keyboard or a device that mimics a physical musical instrument. A synthesizer may also play music entirely on its own from a preprogrammed set of instructions or from sheet music. The output of a musical synthesizer may be sent to an Al model through an electronic port or by wireless transmission. Instrumentalmusic is music comprised entirely of any combination of physical, synthesized, or virtual musical instruments.
[0024] A live music audio stream may be comprised of one or more audio channels. A live music audio stream may be stored as it is received as a music audio file on the memory of a computing device, or the memory of a server communicatively coupled to the computing device.
[0025] A musical composition is an original piece of music, either vocal music, instrumental music or a combination of both vocal music and instrumental music.
[0026] In some forms of music, a song structure comprised of an introduction (intro), a repeating verse and chorus, a bridge, and a conclusion (outro) is used. The chorus usually has the same lyrics throughout the song whereas the verse usually has a new set of lyrics every time its music appears.
[0027] A beat is a basic or whole repeating unit of time in a musical piece. In some forms of music, beats may be divided into half beats, quarter beats, eighth beats, etc.
[0028] A note is a distinct and isolatable sound produced by a vocalist or musical instrument, which may be played for the duration of multiple beats, a single beat, half beat, quarter beat, eighth beat, etc. of the musical composition.
[0029] A note may have a pitch, with a fundamental frequency.
[0030] Timbre is the harmonic content, spectral content, vibrato, tremolo, and time variation of a note that distinguishes one voice from another and one musical instrument from another.
[0031] An octave is the interval between a first note and a second note with double or half the first note’s pitch. In many forms of music, there are a defined number of notes in an octave and notes have defined pitches. For example, Western Music has 12 notes per octave and the note A above middle C has a fundamental frequency or pitch of 440 Hertz (cycles per second). Some forms of music may further divide notes into semitones.
[0032] A bend in a note is a smooth transition to a different frequency than the note’s standard frequency. A note may be bent to another note.
[0033] A chord is two or more notes played simultaneously.
[0034] A rhythmic cycle is a repeating cycle of beats in music, involving the same or similar repeating vocal or musical instrument sound patterns. In Western Music meter refers to regularly occurring patterns and accents such as bars. Equivalences to meter are found in other forms of music. An example is Tala in Indian Music.
[0035] Timing or division is the arrangement or grouping of beats within a rhythmic cycle.
[0036] A time signature in Western Music is a convention that denotes how many note values of a particular type are contained in each bar.
[0037] A downbeat is the first beat of a rhythmic cycle in music.
[0038] Tempo is the pace at which music is played and may be measured in beats per minute in some forms of music. Rubato is a fluctuation of the tempo, often a speeding up and then slowing down the tempo.
[0039] A scale is a graduated set of notes. A scale ordered by increasing pitch is an ascending scale, and scale ordered by decreasing pitch is a descending scale. Most scale patterns have the same pattern of notes in every octave. An example is the C Major scale in Western Music. A key is a group of notes that creates a tonal center. In some forms of music, additional rules are added to frameworks akin to scales and keys. An example is the Raga system in Indian Music; each Raga has a melodic structure in addition to defined notes.
[0040] Musical traditions may be based on geographic regions. Examples include but are not limited to Western Music, Indian Music, and Persian Music.
[0041] A musical genre is a set of music that shares styles, traditions, or conventions. Examples of Western Music genres include but are not limited to Rock, Pop, Hip-hop, Blues, Rhythm and Blues, Country, and Classical.
[0042] A musical subgenre is a style of music within a genre. An example of a subgenre is Rap Music within the genre of Hip-hop Music. In the context of this disclosure, a defined subgenre handled in the same way as a genre.
[0043] Some forms of music have a tradition of musical notation written as music symbols to indicate notes, chords, timings, and rhythms of compositions. An example is sheet music in Western Music, which may be represented, without limitation, in hand-written, printed or electronic format. Other forms of music, such as Indian Music have a limited tradition of written musical notation but do have an oral musical notation.
[0044] In some embodiments, a system for tracking a live music audio stream to generate an accompanying simulated music audio stream is disclosed.
[0045] The system may be comprised of a Computing Device with a processor, an input unit communicatively coupled to the processor of the Computing Device for receiving audio signals to generate a live music audio stream, a memory communicatively coupled to the processor of the Computing Device for holding an artificial intelligence model (Al Model) and training datasets and further holding instructions executable by the processor.
[0046] The Al Model may be trained using the training datasets to track at least one vocalist or at least one musical instrument or a combination of the at least one vocalist or the at least one musical instrument in a live music audio stream and to generate at least one virtual vocalist or at least one virtual musical instrument or a combination of the at least one virtual vocalist or the at least one virtual musical instrument in a simulated music audio stream to accompany the live music audio stream in real time, and in synchronization with the live music audio stream.
[0047] A simulated music audio stream generated for real time applications has no perceivable timing delay or latency compared to the live music audio stream it is intended to be used with. It is generally accepted by the music industry that the maximum acceptable latency is 12 milliseconds before a discrepancy in timing is perceived by musicians.
[0048] The training datasets may comprise recorded training music audio fdes of representative compositions in select genres of music.
[0049] The Al Model may be trained to determine combinations of music synchronization parameters comprised of beats, downbeats, rhythmic cycles, tempos, timings, meters, and time signatures in the training datasets and in real time in live music audio streams.
[0050] The Al Model may be trained to determine combinations of music description parameters comprised of notes, timbres, bends, scales, keys, chords and genres in the training datasets and in real time in live music audio streams.
[0051] The Al Model may be trained to correlate sheet music with vocalists or musical instruments or combinations of vocalists and musical instruments in the training datasets and in real time in live music audio streams.
[0052] The Al Model may be trained to a select a virtual vocalist or a virtual musical instrument or combinations of virtual vocalists and virtual musical instruments to generate the simulated music audio stream in a style based on at least one selected from the group of the genre of the live music audio stream, a genre selected by the Al Model, a genre selected by a user of the Al Model, sheet music selected by the user, a particular composition selected by the user, a genre created by the user or a trainer of the Al Model, or any combination thereof. A style based on a genre may closely follow the conventions of a genre or may include improvisations that deviate from the conventions of the genre.
[0053] The system may include a server, with a processor and memory communicatively coupled to the processor, and with the server communicatively coupled to the Computing Device, to provide computing and data storage services to the Computing Device. The server may store part or all of the Al Model and may run part or all of the Al Model. The sever may store part or all of the training datasets.
[0054] The Al Model training and inference computations may be done entirely on the Computing Device, partly on the Computing Device and partly on the server, or entirely on the server with Computing Device acting as a user interface. The system may include operating systems, programs, applications, applets, apps, and scripts running on the Computing Device and on the server that provide a framework for the Al Model. The user of the Computing Device may be a vocalist, a musician, or a person assigned to operate the Computing Device. A trainer or an Al trainer is a person who trains Al models.
[0055] The system may include at least one audio input comprised of one or more analog microphones connected to or included in the Computing Device and sampled through one or more analog-to-digital converters (ADCs) in the Computing Device, or one or more XLR or line inputs connected to one or more external microphones and sampled through one or more ADCs in the Computing Device, or one or more digital audio inputs comprised of but notlimited to USB ports, MIDI ports, S / PDIF ports, Wi-Fi, Internet, or Bluetooth, all of which may be used to generate or add to a live music audio stream.
[0056] The system may include audio outputs on the Computing Device comprised of but not limited to headphone outputs, line outputs, XLR outputs, the audio outputs converted to analog using digital-to-analog converters (DACs) and amplified, or digital outputs such as but not limited to USB ports, MIDI ports, S / PDIF ports, Wi-Fi, Internet, and Bluetooth.
[0057] The system may include training datasets comprised of recorded training music audio files of representative compositions in additional traditions and genres of music for the Al Model. The recorded training music audio files may be stored in the memory of the Computing Device or the memory of a server accessible by the Computing Device through a communications network. The Al training and inference computations may be done on the Computing Device or the server.
[0058] In some embodiments, the system may allow the simulated music audio stream to be output through an audio output device communicatively coupled to the processor to accompany the live music audio stream in real time. Audio output devices may include, without limitation, wired or wireless headphones, earphones, or loudspeakers.
[0059] Feedback cancellation may be used on the live music audio stream if the simulated music audio stream is played through loudspeakers to isolate the live music audio stream. The simulated music audio stream may be cancelled or removed from the live music audio stream by measuring the delay from playing the simulated music audio stream to when it is picked up in the live music audio stream, or by calibrating the delay before a live performance.
[0060] In some embodiments, the system may allow the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream to be transmitted over a communications network.
[0061] In some embodiments, the system may allow the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream to be stored in the memory of the Computing Device or the memory of theserver or a combination of the memory of the Computing Device and the memory of the server, for future playback and post-processing.
[0062] In some embodiments, a method for training, using a processor of a Computing Device, an Artificial Intelligence model (Al Model), the Al Model stored in a memory communicatively coupled to the processor, on training datasets comprised of recorded training music audio files of representative compositions in select genres of music, to track a live music audio stream generated by an audio input unit communicatively coupled to the processor of the Computing Device, and generate an accompanying simulated music audio stream is disclosed.
[0063] The method may comprise training the Al Model to track at least one vocalist or at least one musical instrument or a combination of the at least one vocalist or the at least one musical instrument in a live music audio stream and to generate at least one virtual vocalist or at least one virtual musical instrument or a combination of the at least one virtual vocalist or the at least one virtual musical instrument in a simulated music audio stream to accompany the live music audio stream in real time, and in synchronization with the live music audio stream
[0064] The method may comprise training the Al Model to determine combinations of music synchronization parameters comprised of beats, downbeats, rhythmic cycles, tempos, timings, meters, and time signatures in the training datasets and in live music audio streams.
[0065] The method may comprise training the Al Model to determine combinations of music description parameters comprised of notes, timbres, bends, scales, keys, chords, and genres in the training datasets and in live music audio streams.
[0066] The method may comprise training the Al Model to correlate sheet music with vocalists or musical instruments or combinations of vocalists and musical instruments in the training datasets and in real time in live music audio streams.
[0067] The method may comprise training the Al Model to select a virtual vocalist or a virtual musical instrument or combinations of virtual vocalists and virtual musical instruments to generate the simulated music audio stream in a style based on at least one selected from the group of the genre of the live music audio stream, a genre selected by the Al Model, a genre selected by a user of the Al Model, sheet music selected by the user, a particular composition selected by the user, a genre created by the user or a trainer of the Al Model, or any combinationthereof. A style based on a genre may closely follow the conventions of a genre or may include improvisations that deviate from the conventions of the genre.
[0068] The method may comprise storing the Al Model or the training datasets or a combination of the Al Model and the training datasets in the memory of the Computing Device or the memory of a server comprised of a processor and memory communicatively coupled to the processor, or a combination of the memory of the Computing Device and the memory of the server.
[0069] The method may comprise running the Al Model on the Computing Device or on the server, or a combination of the Computing Device and the server.
[0070] The method may comprise the outputting the simulated music audio stream through an audio output device communicatively coupled to the processor of the Computing Device to accompany the live music audio stream in real time. Audio output devices may include, without limitation, wired or wireless headphones, earphones, or loudspeakers.
[0071] The method may comprise transmitting the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream over a communications network.
[0072] The method may comprise storing the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream, in the memory of the computing device or the memory of the server or a combination of the memory of the computing device and the memory of the server, for future playback and post-processing.
[0073] Training of the Al Model may begin with supervised training: Al trainers may indicate beats and downbeats in a first set of training music audio files using any suitable means, such as but not limited to marking beats and downbeats on a graphical user interface displaying each training music audio file on a display connected to the Computing Device, marking the beats and downbeats on sheet music associated with each training audio file on a graphical user interface on a display connected to the Computing Device, or hitting a key on a keyboard, a button on a touchscreen or a button and on a specialized input controller while listening to and playing each training audio file through the Al Model at normal playback speed. The Al Modeltraining may be reinforced by training the Al Model on a second set of training audio music fdes played at normal speed and correcting the Al Model inferences on when the beats and downbeats occur in each training audio music fde. After the Al Model reliably learns how to determine beats and downbeats, unsupervised training for beat and downbeat detection may be done on a third set of training audio music fdes, which may be played at higher-than-normal speed. Testing of the Al Model’s ability to determine beats and downbeats in real time may done on live music audio streams. From the determined beats and downbeats, the Al Model may determine the rhythmic cycle and tempo of the live music audio stream.
[0074] During Al training, the Al Model may be directed to by the trainer, or the Al Model may on its own, correlate intentional cues, such as but not limited to emphasis or loudness of certain notes or beats or varying the tempo of a composition, and unintentional cues, such as but not limited to arrangements of notes, from a vocalist or a musical instruments or a combination of vocalists and musical instruments in the training datasets, with the indicated and determined beats and downbeats in the supervised training, reinforced training, and unsupervised training to determine beats and downbeats in live music audio streams.
[0075] In addition to determining the beats, downbeats, rhythmic cycles and tempo in live music audio streams, the Al Model may be trained to recognize the timings, meters, time signatures, notes, timbres, bends, scales, keys, chords, genres, and styles from intentional and unintentional cues in the training datasets and in live music audio streams using supervised, reinforced, and unsupervised training.
[0076] The Al Model may be trained to track and replicate the timbres and playing patterns of musical instruments in training datasets comprised of music audio fdes of representative compositions in select genres of music for select virtual musical instruments.
[0077] The Al Model may be trained to track and replicate the lyrics, timbres, and singing patterns of vocalists in training datasets comprised of music audio fdes of representative compositions in select languages and select genres of music for select virtual vocalists.
[0078] In some embodiments, the Al Model may be trained to track and replicate the lyrics, timbres, and singing patterns of a particular vocalist on training datasets comprised of music audio fdes of representative compositions of the vocalist to generate one or more virtual vocalists in synchronization and in harmony with the vocalist for select compositions.
[0079] In some embodiments, the Al Model may be trained to synchronize with particular compositions in live music audio streams or in music audio files, store the music synchronization parameters for each composition, and use the stored music synchronization parameters for a particular composition to synchronize with a composition in a live music audio stream.
[0080] In some embodiments, the Al Model may be trained to add an introduction (intro) or a conclusion (outro) or combinations of an intro and outro in a style based on at least one selected from the group of the genre of the live music audio stream, a genre selected by the Al model, a genre selected by a user of the Al model, sheet music selected by the user, a particular composition selected by the user, a genre created by the user or a trainer of the Al model, or any combination thereof to a simulated music audio stream.
[0081] In some embodiments, the Al Model may be trained to store music synchronization parameters from one or more repetitions of a rhythmic cycle or one or more repetitions of a chorus and verse for a particular composition and use the music synchronization parameters to synchronize with a composition in a live music audio stream.
[0082] In some embodiments, the Al Model may be trained to monitor for changes in combinations of rhythmic cycle, tempo, timing, meter, and time signature in a live music audio stream and change combinations of rhythmic cycle, tempo, timing, meter, and time signature in the simulated music audio stream to resynchronize a simulated music audio stream with the live music audio stream.
[0083] In some embodiments, the Al Model may be trained to monitor for style changes in the live music audio stream and respond with style changes to the simulated music audio stream.
[0084] In some embodiments, the Al Model may be trained to accept as selections from the Al Model user, by text prompts, through a graphical user interface, or by audio prompts, combinations of the rhythmic cycle, tempo, timing, meter and time signature, scale, key, genre, virtual vocalists and virtual instruments, and when to start and stop the simulated music audio stream.
[0085] In some embodiments, the Al Model may be trained to store the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream in the memory of the Computing Device or the memory of a server communicatively coupled to the Computing Device, or a combination of the memory of the Computing Device and the memory of the server, for future playback and postprocessing.
[0086] In some embodiments, the Al Model may be trained to start or stop or combinations of start and stop a virtual vocalist or musical instrument or combinations of virtual vocalists and musical instruments in the simulated music stream on cues from a vocalist or a musical instruments or combinations of vocalists and musical instruments in the live music audio stream. Cues include, without limitation, varying the loudness of a note, word or syllable in a composition, varying the tempo of the composition, or singing or playing a series of notes or words.
[0087] In some embodiments, the Al Model may be trained to start or stop or combinations of start and stop a virtual vocalist or musical instrument or combinations of virtual vocalists and musical instruments in the simulated music stream, and when beats and downbeats occur, without limitation, from the user actions of pressing a button on a screen with a mouse, a key on a keyboard, a key on a controller, or a button on a touchscreen on the Computing Device running the Al Model.
[0088] In some embodiments, the Al Model may be trained to generate a virtual vocalist or a virtual musical instruments or combinations of virtual vocalists and musical instruments in a simulated music audio stream when a vocalist or musical instrument or combinations of vocalists and musical instruments in the live music audio stream are silent for a period of time, as directed by the vocalists, musicians or by a user of the Computing Device.
[0089] In some embodiments, the Al Model may be designed or trained to generate a simulated music audio stream with compensation for latency from the generation of the simulated music audio stream to playback of the simulated music audio stream through an audio output device, by a time difference specified by a user of the Computing Device or a time difference measured by the Al Model. A latency may be introduced by the Al Model’s computer processing or by wireless links to the audio output device.
[0090] In some embodiments, the Al Model may trained to run entirely in generative mode without a live music audio stream using, without limitation, a text, graphical or audio prompt interface in a select language to allow a user for the Al Model to select combinations of the rhythmic cycle, tempo, timing, meter, scale, key, and genre, specify when to start and stop a virtual vocalist or a virtual musical instrument or combinations of virtual vocalists and virtual musical instruments, to generate a simulated music audio stream in a style based on at least one selected from the group of the genre of the live music audio stream, a genre selected by the Al Model, a genre selected by a user of the Al Model, sheet music selected by the user, a particular composition selected by the user, a genre created by the user or a trainer of the Al Model, or any combination thereof.
[0091] In some embodiments, the Al Model inference may be run on recorded music audio fdes in real time in addition to live music audio streams.
[0092] In some embodiments, the Al Model may be trained, in addition to Western Music, on other musical traditions in select languages. An example is but not limited to Indian Music: the Al Model may be trained on datasets of recorded training music audio fdes to determine the Tala, which has elements of a beat, downbeat, rhythmic cycle, tempo, timing and meter and determine the Raga, which has a melodic structure in addition to defined notes.
[0093] In some embodiments, combinations of a live music audio stream, generated sheet music for the live music audio stream, a simulated music audio stream, and generated sheet music may for the simulated music audio stream be stored as files in combinations of the memory of the Computing Device or the memory of a server for future playback and postprocessing. The Al Model may be trained to allow the user to, using without limitation, a text, graphical or audio prompt interface in a select language, without limitation in a genre of the user’s choosing, fix synchronization issues, add an introduction, add a conclusion, tune, modify, or remove a live vocalist or a live musical instrument or a combination of live vocalist and live musical instruments, and add to, tune, modify, remove or replace a virtual vocalist or virtual musical instrument or combinations of virtual vocalists and virtual musical instruments in the stored files.
[0094] In some embodiments, the live music audio stream may be comprised of more than one audio channel.
[0095] In some embodiments, the simulated music audio stream may be comprised of more than one audio channel.
[0096] In some embodiments, the Al Model may be run on a smartphone, tablet, laptop computer, desktop computer, embedded computer, VR headset, AR headset, or similar Computing Device.
[0097] In some embodiments, the Al Model may partly run on a local server with a local area network (LAN) connection to the Computing Device or on a remote server with a wide area network (WAN) connection to the Computing Device. In some embodiments the Al Model may fully run on a server with the Computing Device acting as a client. It is generally accepted by the music industry that the maximum acceptable latency is 12 milliseconds (ms) before a discrepancy in timing is perceived by musicians. If the LAN or WAN latency is greater than 12ms, the Al Model may be designed or trained to measure the latency and generate a simulated music audio stream that has compensation for the latency and is played in real time by the Computing Device or generate a simulated music audio stream that is buffered by the Computing Device and played in sync with the live music audio stream by the Computing Device.
[0098] In some embodiments, the if the Al Model for select genres exceeds the memory or processing capabilities of a Computing Device, the Al Model may be divided into a number of smaller Al models based on, without limitation, musical traditions, genres, groupings of compositions, or individual compositions, one or more such smaller Al models selected and run on the Computing Device by the user of the Computing Device from the memory of the Computing Device or the memory of a server.
[0099] In some embodiments, the Al Model may partly run on a server to synchronize with a live music audio stream and detect the genre of a live input audio stream, and thereafter run on the Computing Device.
[0100] In some embodiments, the Al Model may be implemented as one or more computer programs, applications, applets, apps, or scripts running on one or more Computing Devices or servers. Part or all of the functionality of the Al Model may be integrated into a third-party application or into an operating system running on a Computing Device or server.
[0101] A live music audio stream may be comprised of combinations of vocalists and musical instruments. A reference to a single vocalist or a single musical instrument in this disclosure may be replaced by the plural form and a reference to plural vocalists and plural musical instruments may be replaced by the singular form.
[0102] Referring to Figure 1, there is provided a method to generate a simulated music audio stream in real time in sync and in the style of the genre of a live music audio stream using a trained Al Model. At Box 10, the beats, downbeats, rhythmic cycle, tempo, timing, meter, and time signature in the in the live music audio stream are determined by the Al Model. At Box 20, the notes, timbres, bends, scales, keys, and chords of vocalists and musical instruments in the live music audio stream are determined by the Al Model. At Box 30, the genre of the live music audio stream is determined by the Al Model. At Box 40, simulated vocalists and simulated musical instruments are selected for the genre of the live music audio stream by the Al Model, and a simulated music audio stream is generated in real time in sync with and in the style of the genre of the live music audio stream by the Al Model.
[0103] Referring to Figure 2, an embodiment of the present disclosure is shown. A vocalist (110) sings into a smartphone (120) running a trained Al Model. The Al Model converts the sampled analog audio into a live music audio stream and determines the beats, downbeats, rhythmic cycle, tempo, timing, meter, time signature, notes, timbre, bends, scale, key, chords, and genre in the live music audio stream, selects virtual vocalists and musical instruments or uses virtual vocalists and musical instruments selected by the vocalist, and generates a simulated music audio stream in real time in sync with and in the style of the genre of the live music audio stream, and plays the simulated music audio stream through wireless headphones (130) worn by the vocalist. Smartphone to wireless headphone latency may be measured and compensated by the Al Model. Alternately, wired headphones may be used to reduce or eliminate the wireless headphone latency.
[0104] Referring to Figure 3, an embodiment of the present disclosure is shown. A vocalist (210) plays a guitar (220) and sings into a microphone (230) connected to a tablet (240). The tablet is wirelessly connected to a server (250). The tablet computer and server together run a trained Al Model. The Al Model running on the tablet converts the sampled analog audio into a live music audio stream and sends it in real time to the server. The server determines the beats, downbeats, rhythmic cycle, tempo, timing, meter, time signature, notes, timbre, bends, scale, key, chords, and genre in the live music audio stream, selects virtual vocalists and musicalinstruments or uses virtual vocalists and musical instruments selected by the vocalist on the tablet, and generates a simulated music audio stream in real time in sync with and in the style of the genre of the live music audio stream. The server sends the simulated music audio stream to the tablet in real time which plays the simulated music audio stream through wireless headphones (260) worn by the vocalist. Server to tablet and tablet to wireless headphone latency may be measured and compensated by the Al Model on the server. Alternately, a wired network connection from the table to the server and wired headphones may be used to reduce or eliminate the server to tablet and tablet to headphone latencies.
[0105] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Accordingly, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and "comprising," when used in this specification, specify the presence of one or more stated features, integers, steps, operations, elements, and components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and groups.
[0106] It is contemplated that any part of any aspect or embodiment discussed in this specification may be implemented or combined with any part of any other aspect or embodiment discussed in this specification.
[0107] While particular embodiments have been described in the foregoing, it is to be understood that other embodiments are possible and are intended to be included herein. It will be clear to any person skilled in the art that modifications of and adjustments to the foregoing embodiments, not shown, are possible.
Claims
CLAIMS1. A system for tracking a live music audio stream to generate an accompanying simulated music audio stream, the system comprising:(a) a computing device with a processor;(b) an input unit communicatively coupled to the processor of the computing device for receiving audio signals to generate a live music audio stream;(c) a memory, communicatively coupled to the processor of the computing device for holding an artificial intelligence model (Al model) and training datasets and further holding instructions executable by the processor to:(i) train the Al model using the training datasets to track at least one vocalist or at least one musical instrument or a combination of the at least one vocalist or the at least one musical instrument in a live music audio stream and to generate at least one virtual vocalist or at least one virtual musical instrument or a combination of the at least one virtual vocalist or the at least one virtual musical instrument in a simulated music audio stream to accompany the live music audio stream in real time, and in synchronization with the live music audio stream.
2. The system of claim 1, wherein the training datasets comprise recorded training music audio fdes of representative compositions in select genres of music.
3. The system of claim 1, wherein the Al model is trained to determine combinations of music synchronization parameters in the training datasets and in live music audio streams.
4. The system of claim 1, wherein the Al model is trained to determine combinations of music description parameters in the training datasets and in live music audio streams.
5. The system of claim 1, wherein the Al model is trained to correlate sheet music with vocalists or musical instruments or combinations of vocalists and musical instruments in the training datasets and in live music audio streams.
6. The system of claim 1, wherein the Al model is trained to select a virtual vocalist or a virtual musical instrument or combinations of virtual vocalists and virtual musicalinstruments to generate the simulated music audio stream in a style based on at least one selected from the group of the genre of the live music audio stream, a genre selected by the Al model, a genre selected by a user of the Al model, sheet music selected by the user, a particular composition selected by the user, a genre created by the user or a trainer of the Al model, or any combination thereof.
7. The system of claim 1 , wherein a server, with a processor and memory communicatively coupled to the processor, and with the server communicatively coupled to the computing device, provides computing and data storage services to the computing device.
8. The system of claim 1, wherein the simulated music audio stream is output through an audio output device communicatively coupled to the processor to accompany the live music audio stream in real time.
9. The system of claim 1, wherein the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream is transmitted over a communications network.
10. The system of claims 1, wherein the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream, is stored in the memory of the computing device or the memory of the server or a combination of the memory of the computing device and the memory of the server, for future playback and post-processing.
11. A method for training, using a processor of a computing device, an Artificial Intelligence model (Al model), the Al model stored in a memory communicatively coupled to the processor, on training datasets comprised of recorded training music audio files of representative compositions in select genres of music, to track a live music audio stream generated by an audio input unit communicatively coupled to the processor of the computing device, and generate an accompanying simulated music audio stream, the method comprising:(a) training the Al model to track at least one vocalist or at least one musical instrument or a combination of the at least one vocalist or the at least one musical instrument in a live music audio stream and to generate at least one virtual vocalist or at least one virtual musical instrument or a combination of the at least one virtual vocalist or the at least one virtual musical instrument in a simulatedmusic audio stream to accompany the live music audio stream in real time, and in synchronization with the live music audio stream.
12. The method of claim 11, wherein the Al model is trained to determine combinations of music synchronization parameters in the training datasets and in live music audio streams.
13. The method of claim 11, wherein the Al model is trained to determine combinations of music description parameters in the training datasets and in live music audio streams.
14. The method of claim 11, wherein the Al model is trained to correlate sheet music with vocalists or musical instruments or combinations of vocalists and musical instruments in the training datasets and in real time in live music audio streams.
15. The method of claim 11, wherein the Al model is trained to select a virtual vocalist or a virtual musical instrument or combinations of virtual vocalists and virtual musical instruments to generate the simulated music audio stream in a style based on at least one selected from the group of the genre of the live music audio stream, a genre selected by the Al model, a genre selected by a user of the Al model, sheet music selected by the user, a particular composition selected by the user, a genre created by the user or a trainer of the Al model, or any combination thereof.
16. The method of claim 11 , wherein the Al model or the training datasets or a combination of the Al Model and the training datasets is stored in the memory of the Computing Device, or the memory of a server comprised of a processor and memory communicatively coupled to the processor, or a combination of the memory of the Computing Device and the memory of the server.
17. The method of claim 11, wherein the Al model is run on the Computing Device or on the server, or a combination of the Computing Device and the server.
18. The method of claim 11, wherein the simulated music audio stream is output through an audio output device communicatively coupled to the processor of the Computing Device to accompany the live music audio stream in real time.
19. The method of claim 11, wherein the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream is transmitted over a communications network.
0. The method of claim 11, wherein the live music audio stream or the simulated music audio stream or a combination of the live music audio stream and the simulated music audio stream, is stored in the memory of the computing device or the memory of the server or a combination of the memory of the computing device and the memory of the server, for future playback and post-processing.
Citation Information
Patent Citations
System and method for a networked virtual musical instrument
US12154533B2
Real-time integration and review of musical performances streamed from remote locations
US20210225344A1
Automated music production
WO2020121225A1