Method and system for near-live performance and recording of live Internet music without time delay

Through the coordinated work of network cache, storage, timing and mixing modules, the problem of musicians performing and recording Internet music without delay in different geographical locations is solved, and time accuracy and recording fidelity between musicians are achieved, adapting to real-time performance and recording under different network conditions.

CN114120942BActive Publication Date: 2025-09-02SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110702792.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-25
Filing Date
2021-06-24
Publication Date
2025-09-02
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

The prior art cannot realize that musicians can perform and record live Internet music in real time without delay in different geographical locations. The network delay causes performance inaccuracy to fail to meet the music creation requirements.

Method used

By using network cache, storage, timing and mixing modules, combined with Internet bandwidth testing and quality/delay setting modules, the player's performance is transmitted at full resolution or lower resolution, and the master clock synchronization is used to ensure that the performance and timing information of each player are seamlessly spliced ​​and mixed.

Benefits of technology

It realizes synchronous performance and recording between musicians without delay, ensuring the time accuracy and fidelity of the final recording results, and adapts to real-time performance and recording requirements under different network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120942B_ABST
    Figure CN114120942B_ABST
Patent Text Reader

Abstract

An exemplary method includes a processor executing instructions stored in a memory, the instructions for generating an electronic count, binding the electronic count to a first performance to generate a master clock, and transmitting a first performance by a first musician and first timing information to a network cache, storage, timing, and mixing module. The first performance by the first musician can be recorded locally at full resolution and can be transmitted to a full resolution media server, and the first timing information can be transmitted to the master clock. The first performance by the first musician is transmitted to a sound device of a second musician, and the second musician composes a second performance, which and the second timing information are transmitted to the network cache, storage, timing, and mixing module. The first performance and the second performance are mixed together with the first timing information and the second timing information to generate a first mixed audio, which can be transmitted to the sound device of a third musician.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is related to U.S. non-provisional patent application serial number 16 / 912,569, filed concurrently on June 25, 2020, entitled “Methods and Systems for Performing and Recording Live Internet Music Near Live with no Latency,” all contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to the fields of music performance and recording as well as network latency and synchronization.

[0004] Description of Related Technology

[0005] Music is typically recorded with some combination of simultaneous and asynchronous performances. That is, some or all musicians play the music simultaneously, and it is recorded as a single performance. Initially, all music was recorded as a single performance, with all musicians playing simultaneously. In the 1950s, Les Paul pioneered the multi-track recorder, which allowed him to play a second part of music over a pre-recorded part. After that, musicians began recording one or more instruments in the initial recording and then adding other instruments later—this is called overdubbing.

[0006] For the past 20 years, musicians have wished they could perform live with other musicians in different locations. While this has been possible to some extent, network latency has been too long for most musical styles to create useful recordings. Good musicians can detect notes or drum beats that are "out of tune" with inaccuracies as low as a few milliseconds. Even at the speed of light, it takes about 13 milliseconds to travel from Los Angeles to New York (26 milliseconds round trip), making this latency too long for musicians to play together in real time. Summary of the Invention

[0007] The exemplary embodiments provide systems and methods for nearly live performance and recording of live Internet music without delay.

[0008] The exemplary method includes a processor executing instructions stored in a memory, the instructions being used to generate an electronic count-in, binding the electronic count-in to a first performance to generate a master clock, and transmitting a first performance of a first musician and first timing information to a network cache, storage, timing, and mixing module. The first performance of the first musician can be recorded locally at full resolution and can be transmitted to a full-resolution media server, and the first timing information can be transmitted to the master clock. Alternatively, a lower-resolution version of the first performance of the first musician can be transmitted to a compressed audio media server, and the first timing information can be transmitted to the master clock.

[0009] Subsequently, according to an exemplary embodiment, the first performance of the first musician is transmitted to the sound device of the second musician, and the second musician creates a second performance, which and the second timing information are transmitted to the network cache, storage, timing, and mixing module. The first performance and the second performance are mixed together with the first timing information and the second timing information to produce a first mixed audio, which can be transmitted to the sound device of the third musician. The third musician creates a third performance and third timing information, which is mixed with the first mixed audio to produce a second mixed audio. This process is repeated until the last musician has completed the performance and recording.

[0010] An exemplary system for network caching, storage, timing, and mixing of media includes: an internet bandwidth test module configured to ping a network and determine a bandwidth to a first user device; a quality / latency setting module communicatively coupled to the internet bandwidth test module, the quality / latency setting module configured to determine a resolution of the media based on the bandwidth; and a network audio mixer communicatively coupled to the quality / latency setting module, the network audio mixer configured to transmit the media to the first user device at the determined resolution. The system includes: a full-resolution media server configured to receive a time synchronization code for the media and a master clock from the first user device; and / or a compressed media server configured to receive a time synchronization code for the media and a master clock from the first user device.

[0011] Subsequently, according to various exemplary embodiments, the Internet bandwidth test module pings the network and determines the bandwidth to the second user device in order to determine the resolution of the media to be transmitted to the second user device. In other exemplary embodiments, the media is a single mixed track that combines performances of multiple musicians, and the performances have a range of resolutions. In this case, both the full-resolution media server and the compressed media server transmit the media to the network audio mixer, which transmits the media to the second user device. The system receives the performance from the second user device and mixes it with the single mixed track.

[0012] An exemplary system for managing internet bandwidth, latency, quality, and mixing of media includes a processor that executes instructions stored in a memory, the instructions controlling: a component for measuring bandwidth over time, a component for varying levels of compression, and a component for seamlessly stitching together various resolutions using a common time code while quality varies over time. All components are communicatively coupled to one another and connected to a single fader via a bus. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and further objects, features and advantages of the present invention will become apparent upon consideration of the following detailed description of certain specific embodiments of the present invention, particularly when read in conjunction with the accompanying drawings, in which like reference numerals are used to designate like parts throughout the various drawings, and in which:

[0014] Figure 1 is a high-level diagram of the architecture, showing musicians, web services, and listeners.

[0015] Figure 2 Provides more details about the first player, network stack, and transport stack.

[0016] Figure 2A Shows how time relates to music samples.

[0017] Figure 2B It is shown that this can be used with video as well as audio.

[0018] Figure 3 The network and transport stack associated with the second (and other) musicians are shown.

[0019] Figure 4 Shows how to connect musicians in a chain via the network and transport stack and how to optimize playback synchronization and bandwidth.

[0020] Figure 5 It shows how network caching, storage, timing, and mixing modules work together as music passes from one musician to the next.

[0021] Figure 6 Shows how Internet bandwidth, latency, quality, and mix work together.

[0022] Figure 7 Shows how a single performance can be produced from different resolutions.

[0023] Figure 8 An exemplary jam band scene is shown.

[0024] Figure 9 Exemplary timing for a jam band scene is shown.

[0025] Figure 10 An exemplary dramatic podcast scene is shown. DETAILED DESCRIPTION

[0026] The elements identified throughout the text are exemplary and may include various alternatives, equivalents, or derivatives thereof. Various combinations of hardware, software, and computer-executable instructions may be utilized. Program modules and engines may include routines, programs, objects, components, and data structures that implement the performance of specific tasks when executed by a processor, which may be general or special-purpose. The computer-executable instructions and associated data structures stored in a computer-readable storage medium represent examples of programming means for performing the steps of the method and / or implementing the specific system configurations disclosed herein.

[0027] This disclosure describes a mechanism for allowing musicians to play together continuously in real time, based on the sounds of the previous musicians. If there are many musicians playing a song together, the first person starts, and although the music may be delayed by a few milliseconds to reach the second person, the second person will play according to what they heard, and for them, both performances are perfectly timed. Now, the third person hears the performance of the first two people (in time with each other), just as the second person heard it, and although the time they heard it may be later than the actual time of the performance, they will play according to what they heard in time, and for them, all three instruments are perfectly timed. This can continue indefinitely.

[0028] To achieve this, a serial recording is required. However, since the audio is transmitted over the network, the quality can easily degrade. That is, once the music starts being played by one musician, it cannot be paused or slowed down, but the bitrate (quality) can be reduced to achieve accurate timing. It is recommended that each performance be recorded in the cloud (for example, on a network server) at full resolution and also compressed if necessary. It may also need to be buffered locally to maintain fidelity so that when the final performance reaches the cloud it is at full resolution. In this way, even if the quality heard by the musicians during the performance degrades slightly, the quality of its recording and transmission to the cloud does not need to be sacrificed, so the final result will be full fidelity and perfectly timed when all are played back in the end.

[0029] like Figure 1 As can be seen in the figure, the entire system consists of the individual musicians (and their equipment and software as well as the recording) as well as network caching, storage, timing and mixing components. The scenario is as follows:

[0030] The first musician (101) begins by saying or generating an electronic count (typically 1, 2, 3, 4). In various exemplary embodiments, there is a signal (digital data or audio data) that indicates the start of the segment and a prompt to let the other musicians know when to start. In some cases, there may be a click track (metronome) played by the first (and possibly subsequent) musicians. In other cases, this may be a countdown on the sound or a pickup of an instrument. Alternatively, there may be a visual prompt such as that given by a conductor. In any case, this first mark (again, not necessarily the strong beat) is absolutely bound to the first performance, and together they become the "master clock" that will be used to synchronize all local clocks and performances. Using NTP or the Network Time Protocol would be the simplest, but NTP is typically only accurate to within 100 milliseconds. All participants' performances must be bound to a common clock that is accurate to less than 1 millisecond. The performance and timing information (102) of the first musician (101) is sent to the network cache, storage, timing and mixing module (103).

[0031] Each musician's performance is recorded locally at full resolution. This performance is ultimately transmitted to a full resolution media server (104). This performance can be sent in real time, but may not be sent in real time. In the absence of optimal bandwidth, the performance can be sent later.

[0032] If there is not enough bandwidth to send the full-resolution audio without delay, a lower-resolution version of the first musician's performance can be sent to the compressed audio media server (105). This lower-resolution version should be sufficient to enable the following musicians to hear the parts in front of them and play accordingly. This lower-resolution version should be of the highest possible quality and, under ideal network conditions, should be almost indistinguishable from the full-quality version. However, depending on bandwidth conditions, the full-resolution audio may have to be sent later.

[0033] At the same time, and as part of the same media file (both full resolution and compressed), timing information is sent to the master clock (106). Audio is typically recorded at 44.1, 48, or 96 kHz, and so by definition, there is a clock that is much more accurate than the 1 millisecond required herein. The timestamp associated with the audio recording is used to set the clock and synchronize it.

[0034] When the second musician (107) hears music from the full-resolution media server (104) or the compressed audio media server (105) according to the network bandwidth, the second musician (107) joins its performance. The performance of the second musician is now sent to the network cache, storage and timing module (103) where the audio and timing information are stored. Meanwhile, the audio of the first two musicians is combined (or mixed) by the network audio mixer (108) and sent to the third musician (109) together with the timing information. The performance of the third musician is sent back to the network cache, storage and timing module (103), where the new audio and timing information are stored together with other performances and then sent to the other musicians until the last musician (110) has performed and recorded.

[0035] The network audio mixer (108) not only combines the performances of the individual musicians so that they can be heard by each other, but also combines the cumulative performances of all the musicians so that the audience (111) can hear them. As will be described in more detail below, the network audio mixer (108) not only combines the different tracks (or performances), but also combines them in a way that provides maximum fidelity. So, for example, if a musician's performance is at a lower resolution due to bandwidth limitations, but their bandwidth is improved, then their quality will also be improved. In addition, the full resolution version will eventually make it to the full resolution media server (104), and no matter when that resolution arrives at the server, people who hear it after that will hear it at full resolution. In the long run, this means that if the music is played later (for example, two hours after a live performance), it will be played at full resolution. In some cases, the resolution of some musicians whose bandwidth is increased can increase the resolution of their parts as their performance unfolds.

[0036] Figure 2 Details are provided for the recording and initial transmission of audio and timing information. Later in the process, those systems and musicians should be able to accurately identify reliable starting points for synchronization. For example, imagine a musician counting (e.g., 1, 2, 3, 4). When the word "one" is recorded, it has a specific, recognizable waveform based on the audio waveform sample occurring at a specific time. By definition, a digital waveform is sampled at a certain frequency (e.g., 44.1kHz, 48kHz, 96kHz, etc.), and position is always associated with time. Figure 2A A sample of the cello playing the tone A4 is shown. The fundamental frequency is 440 Hz, which is approximately 2 1 / 4 milliseconds (the disturbances in the waveform are harmonics and other noises such as from bowing). Once you find a common point in the recording, you can easily calculate the number of milliseconds from that point to anywhere in the clip.

[0037] The same timing information can be applied to video. For example, if the first musician is the conductor, then the other musicians can still follow in time (even if not at the same time). In practice, they may need a common rhythm, such as a click track or drum loop, but in theory, there is nothing to prevent them all from following the same conductor or other visual cue (such as for a film score). See Figure 2B , which is similar to Figure 2 , except that the microphone (201) has been replaced by a camera (213) and the recording of a video (214) has been added to the recording element, which has been synchronized with the local recording (203, 204) by the sampling clock (202).

[0038] Back to Figure 2 , the first musician (200) sounds on a microphone (201), which starts a clock (202) with some audio (or video as described above). The sound is recorded at full fidelity (203) and prepared for transmission. Starting with the recording device turned on and connected to the network, the network is polled to test the bandwidth. If the bandwidth is sufficient, the full fidelity (lossless) version (203) is transmitted along with the timing information (205). However, if the bandwidth is insufficient, a software module in the first musician's recording environment can compress the audio to a smaller file size. For example, the audio codec AAC is considered to have reasonable fidelity at 128 kilobits per second (kbps) created from a 48kHz recording. The uncompressed file will be streamed at a rate of 1536kbps - even using lossless compression, which is still about 800kbps. [Note: Multiple files of any given resolution will result in a higher resolution file when played together than if the instrument were recorded as a single recording.] For example, 16 channels of 16-bit 48k audio, when mixed together, will have higher resolution than 2 channels of 16-bit 48k audio.] Learn more about balancing latency, bandwidth, and quality later in this disclosure.

[0039] Regarding the transport format, the clock will always be tied to every version of every recording (both lossless and compressed). When looking at the transport stack (205), it should be treated as two separate streams, each with the same corresponding time / sync code. This way, when the music arrives at the network caching, storage, timing, and mixing components (servers / services) (206), if the service has to switch between resolutions (208, 209), it can use the common (master) clock (207) to maintain perfect synchronization. When other musicians' performances are combined, this will be done by the network audio mixer (210).

[0040] Figure 3The joining of the second musician (300) is shown. Audio and possibly video come from network caching, storage, timing and mixing services (301), where the media from the first musician is stored and transmitted over the Internet using a transport stack (302) protocol including lossless audio (304), which is bound to timing information (303), and, subject to bandwidth limitations, compressed audio (305) is also bound to timing information (303). Video can be included in this entire process, and those practicing in the audio-visual field can easily build using video based on the data in this disclosure. If there is enough bandwidth, compressed audio may not be needed. When audio arrives, it will first enter a mixing module (306), which will feed information to the second musician monitor (307) (which may be headphones). As the second musician plays or sings, the sound will enter the mixing module either by direct injection (for electronic instruments or acoustic-electric pickups such as piezoelectric or magnetic pickups) or through a microphone (308) where it is combined (mixed) with the audio from the first and second musicians so that both parts can be heard when they play together.

[0041] The second player is recorded losslessly (310) and time stamped using the same clock synchronization (309) as the original recording. The audio from the second player is sent back to the Network Caching, Storage, Timing and Mixing Service (NCSTMS) (301) with the same time code as it received from the original recording using the same transport stack protocol (312). Since the NCSTMS already has the audio of the first player and the same synchronized time code, it is not necessary to send the audio of the first player back to the NCSTMS. It should be noted that there is a network audio mixer at the NCSTMS that mixes the performances of the different players together. This is separate from the mixers at the individual player locations.

[0042] Figure 4 Shown is playback synchronization and bandwidth optimization (408).As mentioned above, synchronization is based on the common time code shared between all resolutions of audio (and video).Sometimes it may be necessary to weigh between quality and time delay.Suppose that a musician (musician N) transmits (lossless compression) with full resolution at a rate of 800kbps, and the next musician (musician N+1) has less bandwidth.For example, if based on carrying out throughput test to the network, for musician N, streaming at a rate of 800kbps, then she / he will have to cache enough music to make the time delay be 15 seconds.But, if musician N receives and sends audio at a rate of 128kbps, then the time delay will only be 75 milliseconds.Play synchronization and bandwidth optimization module (408) can select resolution, and therefore selects required bandwidth to send audio to musician N+1.

[0043] For more details on this, see Figure 5and Figure 6 .

[0044] Figure 5 Shown is a musician N (500). In order to know the possible available bandwidth between musician N (500) and the NNCS™ module (501), the Internet Bandwidth Test module (502) is used. "Ping" the network and finding the bandwidth between two points is a fairly standard practice, and anyone who practices in this field can use this capability. Based on the available bandwidth, the quality / delay setting module (503) will determine the media resolution (e.g., the resolution at which the network audio mixer should send the media to musician N). Figure 6 1 . In more detail (see Figure 5). Depending on the bandwidth, musician N will send his media to either a full resolution media server (506) or a compressed media server (507) and the synchronized time code that goes into the master clock (505). It should be noted that "server" means any server configuration from a hard drive on a home computer to a server array widely distributed on the Internet. Similarly, a "compressed media server" can include multiple resolutions of video and / or audio and can also be distributed. In order to send the media to musician N+1 (508) (the next musician in the chain), the bandwidth must be tested again by the Internet bandwidth test module (502). This determines at what resolution the media is sent to musician N+1. It should be noted that the media sent to musician N+1 is not all of the individual recordings that the musician has played before, but rather a single mixed track that combines all of their performances. Assume, for example, that musician N+1 is the fifth musician in the chain, and that the previous musicians have the following bandwidth limitations on the quality of their performances: musician 1, 800 kbps (completely lossless); musician 2, 450 kbps; musician 3, 800 kbps; musician 4, 325 kbps; and musician 5, 800 kbps. The media will come from a combination of the full-resolution media server (506) and the compressed media server (507), where it will be fed into the network audio mixer (504). The combined "mix" will be sent to musician N+1. Note that in the combined mix, the parts from musicians 1 and 3 will have a higher resolution than the parts from musicians 2 and 4. Note also that the only media that will be sent back to the NCS™ module will be the new performance of musician 5, since the other performances have already been cached. Therefore, any bandwidth limitations connected to Player 5 will only affect the quality of Player 5's part, and even then, only affect the players in the chain - not the final listener who will be able to receive the full fidelity of all players (depending on when they listen).

[0045] Figure 6The bandwidth, quality, latency, and mixing components of the system are shown. Bandwidth affects music quality in two directions. Upload bandwidth affects the quality of the initial transmission of each performance (and subsequent transmissions of the same performance, still at full resolution). Download bandwidth affects the quality that musicians hear when they play together.

[0046] The operating environment of the uploading musician will have its own bandwidth measurement capabilities, so that, for example, the full bandwidth (605) may sometimes appear, or the compression level (606, 607, 608, 609) may vary depending on the bandwidth. The system will seamlessly stitch together the various resolutions using a common time code (where only the quality (not the timing) will vary over time). In practice, all of these will be connected to a single fader via a bus to achieve this level of musician in the mix (there may be a human operating the fader, or there may be an algorithm doing the mixing). This is the case for the second musician in the chain (610, 611, 612, 613) and so on to the Nth musician (614, 615, 616). These levels are combined in the mix, and it is this mix that is transmitted to the next musician in the chain (508) with its bandwidth. It should be noted that the transmission bandwidth from the NCS™ to any individual musician is typically (and usually is today) sent at the appropriate bandwidth to ensure there is no delay. This is independent of the upload bandwidth from each musician. For example, if a musician has particularly low bandwidth, they may receive a lower-quality stream. However, they will still be recorded in full fidelity in their local environment, and the quality of their performance reaching low-latency listeners will reflect their upload bandwidth. Of course, as mentioned before, once their full-resolution performance has been uploaded, subsequent listeners will hear it in full resolution (depending on the listener's bandwidth, of course).

[0047] To clarify the discussion of different resolutions, see Figure 7 It may be helpful. This shows how audio at different resolutions can be recorded and stored. Note that the different resolutions from the first musician (701) are displayed as different waveforms over time (702). Subsequent musicians will hear the performance from the first musician at variable resolutions but as a single performance. The second musician can also be recorded at multiple resolutions (703), as will the subsequent musician (704). As described above, these different performances will be mixed together by the mixing engineer using faders (602, 603, 604) so ​​that they can be heard by subsequent musicians or audience members. Note again that once the higher resolution portions of the audio have been uploaded to the network caching, storage, timing and mixing components, they can be used in subsequent mixes (e.g., after the performance is over) to improve quality.

[0048] As a use case, let's see Figure 8The jam band scenario shown. Let's assume there are 6 musicians, playing drums (801), percussion (802), bass (803), piano (804), and two guitars (805 and 806). They are all connected to NCSTM (807), as are the listeners (808). Suppose you start the drummer, and then after two measures, the percussionist and bassist join. Other musicians can join immediately, or after a few measures. Each musician can only hear the musicians before them in order, but you can change the order through the layout.

[0049] See also Figure 9 , the actual time on the clock (901) is moving forward without stopping, but the actual measure number (902) is moving in time with the musicians. The drummer's measure 1 (903) is the beginning, but each subsequent musician's measure 1 (904) is a little behind - each measure is a little longer than the previous one. After the drummer (905) starts, the percussionist (906), the bassist (907), and the keyboardist (908) are followed. Suppose a guitarist (909) starts just after the keyboardist but before the second guitarist, but she / he wants to be able to hear the other guitars during the solo. In this case, when we say "start in front", we mean "network order", not to be confused with musical order. She / he (or the mixing engineer at a predetermined prompt) may press reset or "change position", and they will start hearing the audio at the new position.

[0050] exist Figure 9 In the example, the gray areas (911, 912, and 913) represent someone laying out a chord. So, assuming a total delay of 2 seconds, when the guitarist hits the switch, they hear the music 2 seconds after where they are, but all the players are playing. Therefore, if I'm laying out a bar or two, I can rejoin while listening to the other players. Choreographing this would probably be easier if there were an interactive chord diagram tracking where I was in the song, but the players will probably be pretty good at identifying where they are pretty quickly.

[0051] Now, in this imaginary jam band scenario, the musicians can take turns laying out and then returning to hear the other musicians play—even the drummer or percussionist can lay out and return a few beats later, only to hear the other musicians. You don't have to go to the end of the line. Maybe the singer always goes last, and "Return" only takes you to the second-to-last, or maybe you only return one or two spots. For example, the drummer and percussionist could switch places. There could be a lot of question-and-answer-style performances, but you won't hear the answer until the final play.

[0052] Another use case would be a drama podcast scenario. In this case, Figure 10As shown, we have many actors creating a nearly live performance online. This can be scripted, or spontaneous, like an interview or a reality show. We can do what we did above, but we also have some other options. Spoken language is not as time-sensitive as music, so we may have more time to perform. Moreover, performances are more serial than parallel and more flexible in their fidelity requirements. In a jam band scenario, while a musician lays out a few bars, they can be placed at the back of the line. Furthermore, the time between performances can be compressed. Let's imagine a performance with six actors (1001, 1002, 1003, 1004, 1005, and 1005). For convenience, let's assume that actors 5 and 6 (1005 and 1006) are in the same location. Tracking time (1007), we start with actor 1 (1001), who speaks for less than a minute. Actor 2 (1002) is listening, and for them, it is real time. Now, actor 1 is planning to rejoin in less than a minute. Let's assume, for the sake of argument, that the delay between actor 1 and actor 2 is 100 milliseconds. Once actor 1 is finished, they can skip the queue. However, there are two constraints: 1) actor 1 doesn't want to miss anything actor 2 has to say; 2) actor 1 wants to hear at least the last part of actor 2's part as unchanged as possible, so their timing and pitch changes are as natural as possible. Therefore, the solution is as follows: when actor 1 skips the queue, they are 100 milliseconds behind actor 2—that is, actor 2 has already spoken for 100 milliseconds. Therefore, when actor 1 jumps back into the queue, they must make up those 100 milliseconds. This is a common technique for speeding up a recording without changing its pitch. So, when actor 1 jumps back into the queue, they hear actor 2 playing back from the recording, but sped up. If this is sped up by 10% (a barely perceptible change in pitch) and the total delay is 100 milliseconds, actor 1 will hear actor 2's voice in real time, at actor 1's actual speed. This can continue indefinitely, with multiple actors entering and catching up as necessary. As with the music recording scenario, the final product (for spoken word with added sound effects) may be only a few minutes behind the live broadcast.

[0053] Modifications can be made without departing from the basic teachings of the invention.A variety of alternative systems can be used to implement the various methods described herein, and various methods can be used to achieve certain results from the above-described systems.

Claims

1. A method for nearly live performance and recording of live Internet music without delay, the method being performed by a processor executing instructions stored in a memory, the instructions comprising: Producing electron counts; binding the electronic count to the first performance to produce a master clock; transmitting the first performance and first timing information of the first musician to the network caching, storage, timing and mixing module; receiving a first performance by the first musician through a sound device of a second musician, and composing a second performance by the second musician; transmitting the second performance and second timing information to the network caching, storage, timing and mixing module; receiving a first mixed audio through a sound device of a third musician, and the third musician composing a third performance, the first mixed audio including the first performance and the second performance and the first timing information and the second timing information; transmitting the third performance and third timing information to the network caching, storage, timing and mixing module; as well as A second mixed audio is received through a sound device of a fourth musician, and the fourth musician composes a fourth performance, the second mixed audio including the third performance and the third timing information and the first mixed audio.

2. The method of claim 1 , further comprising locally recording the first performance of the first musician at full resolution and transmitting it to a full resolution media server; and transmitting the first timing information to the master clock.

3. The method of claim 1 , further comprising transmitting one or more lower resolution versions of the first performance by the first musician to a compressed audio media server; and transmitting the first timing information to the master clock.

4. The method of claim 1, further comprising: The combined performances of the individual musicians are received from the network audio mixer on the plurality of sound devices so as to be heard by each other, and the combined cumulative performances of all the individual musicians are received on the plurality of sound devices so as to be heard by an audience.

5. The method of claim 1 , further comprising: Audio having increased resolution is received at a sound device from a network audio mixer.

6. The method of claim 1, further comprising: The electronic counts have a specific and identifiable waveform that occurs at specific times based on audio waveform samples.

7. The method of claim 1, wherein the electronic count is a video.

8. The method of claim 1, wherein the electronic counts are audio and video.

9. The method of claim 1, further comprising: Activate the recording device; Poll the network to test bandwidth; transmitting full-fidelity digital data with said timing information if said bandwidth is sufficient; and If the bandwidth is insufficient, the audio is compressed to a smaller file size.

10. The method of claim 1, further comprising: The first timing information includes timing information for a lossless version and a compressed version of each recording.

11. The method of claim 10, further comprising: Synchronization is maintained when switching between the lossless version and the compressed version while streaming the recording.

12. A system for nearly live performance and recording of live Internet music without delay, the system comprising: Memory; and one or more processors coupled to the memory, wherein the processors execute instructions stored in the memory, the instructions comprising: Producing electron counts; binding the electronic count to the first performance to produce a master clock; transmitting the first performance and first timing information of the first musician to the network caching, storage, timing and mixing module; receiving a first performance by the first musician through a sound device of a second musician, and composing a second performance by the second musician; transmitting the second performance and second timing information to the network caching, storage, timing and mixing module; receiving a first mixed audio through a sound device of a third musician, and the third musician composing a third performance, the first mixed audio including the first performance and the second performance and the first timing information and the second timing information; transmitting the third performance and third timing information to the network caching, storage, timing and mixing module; and A second mixed audio is received through a sound device of a fourth musician, and the fourth musician composes a fourth performance, the second mixed audio including the third performance and the third timing information and the first mixed audio.

13. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to cause the processor to perform the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Global time server for high accuracy musical tempo and event synchronization

    US20180196393A1