Audio synthesis for synchronous communication

By generating and mixing future portions of an audio stream, and using machine learning models to predict and compensate for latency, the synchronization problem caused by computer network latency is solved, enabling synchronized synthesis and latency hiding of audio streams.

CN120077430BActive Publication Date: 2026-06-02ROBLOX CORP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ROBLOX CORP
Filing Date
2023-10-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Due to computer network latency, people in different locations cannot perform or speak synchronously. In particular, network latency, input latency, and processing latency cause the observed performance or speech to always lag behind the actual time.

Method used

By receiving an audio stream, a synthetic audio stream is generated that predicts future parts of the performance. The first and second audio streams are then mixed to form a synchronously synthesized combined audio stream. A machine learning model is used to predict and compensate for latency, ensuring that the audio streams are synchronized across different devices.

Benefits of technology

It enables the synchronous synthesis of audio streams on different devices, hides transmission delays, makes the listener unaware of the delay, and provides a synchronized performance or speaking experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077430B_ABST
    Figure CN120077430B_ABST
Patent Text Reader

Abstract

A computer-implemented method includes receiving a first audio stream of a performance associated with a first client device. The method also includes, during a time window of the performance, wherein the time window is less than a total time of the performance: generating a synthesized first audio stream that predicts a future portion of the performance based on audio features of the first audio stream, and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream, wherein the time window is advanced, and the generating and mixing are repeated until the performance is completed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is an international application and has priority to U.S. Patent Application No. 17 / 959,736, filed on October 4, 2022, entitled “SYNTHESIZING AUDIO FOR SYNCHRONOUS COMMUNICATION”, filed pursuant to 35 U.S. SC § 119(e), the entire contents of which are incorporated herein by reference. Background Technology

[0003] For multiple people in different locations but connected via a computer network, it is impossible to perform music, recite, or talk synchronously. This is because, due to the transmission delay of the computer network, the performance of person P0 observed by person P1 is always in the past. If each of the n-1 performers P1, P2, P3...P(n-1) each delays precisely by an appropriate amount of time relative to performer P0, then P0 will observe that everyone else is synchronized with each other, but they themselves cannot synchronize with the others.

[0004] This latency can be caused by one or more of at least three types of latency: network latency, input latency, and processing latency. Network latency occurs when there is insufficient network transmission time due to delays in the use of physical devices in the network or processing time at nodes. Input latency occurs when users experience delayed responses, or when users are using client devices with limited functionality. Processing latency occurs when latency is intentionally introduced for tuning and analysis purposes.

[0005] The background description provided herein is intended to introduce the background of this disclosure. The work done by the present inventor, to the extent described in this background section, and in all aspects of the specification that may not have constituted prior art at the time of application, does not expressly or implicitly acknowledge as prior art to this disclosure. Summary of the Invention

[0006] This application generally relates to a system and method for audio synthesis in synchronous communication. According to one aspect of this application, a computer-implemented method includes: receiving a first audio stream of a performance associated with a first client device. The method further includes, during a time window of the performance, wherein the time window is less than the total performance time: generating a synthesized first audio stream based on audio features of the first audio stream to predict future portions of the performance; and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream of the synchronously synthesized first and second audio streams; wherein the time window is shifted forward, and the generation and mixing are repeated until the performance is complete.

[0007] In some embodiments, the method further includes: in response to receiving a first audio stream, determining a performance identifier associated with a performance of the first audio stream, and receiving reference audio based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a different starting point than the reference audio, and the generation of the synthesized first audio stream is also based on the time offset. In some embodiments, generating a synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and the generation of the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof. In some embodiments, the audio features of the first audio stream include one or more speaker identifiers detected in the first audio stream. In some embodiments, the method further includes: determining that the time difference between the first audio stream and a second audio stream exceeds a threshold time difference; and generating graphical data for displaying a user interface, the user interface including user guidance and movement indicators for the performance, the movement indicators prompting the performer associated with a second client device to perform in a manner that reduces the time difference between the first audio stream and the second audio stream. In some embodiments, the method further includes: modifying the combined audio stream to match the acoustic effects of the environment in which the second client device is located. In some embodiments, generating the synthesized first audio stream includes: identifying a portion of the performance that is skipped in the first audio stream; and synthesizing the first audio stream to correct the skipped portion of the performance. In some embodiments, the method further includes: synchronizing the combined audio stream to match the performer's graphically displayed movements.

[0008] In some embodiments, a device includes a processor and a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations including: receiving a first audio stream of a performance associated with a first client device; during a time window of the performance, wherein the time window is less than the total performance time: generating a synthesized first audio stream based on audio features of the first audio stream to predict future portions of the performance; and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream of the synchronously synthesized first and second audio streams; wherein the time window is shifted forward, and the generation and mixing are repeated until the performance is complete.

[0009] In some embodiments, in response to receiving a first audio stream, a performance identifier for a performance associated with the first audio stream is determined; and a reference audio is received based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a different starting point than the reference audio, and the generation of the synthesized first audio stream is also based on the time offset. In some embodiments, generating a synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and the generation of the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

[0010] In some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by one or more computers, cause the one or more computers to perform an operation comprising: receiving a first audio stream of a performance associated with a first client device; during a time window of the performance, wherein the time window is less than the total duration of the performance; generating a synthesized first audio stream based on audio features of the first audio stream to predict future portions of the performance; and mixing the synthesized first audio stream with a second audio stream associated with a second client device to form a combined audio stream of the synchronously synthesized first and second audio streams; wherein the time window is shifted forward, and the generation and mixing are repeated until the performance is complete.

[0011] In some embodiments, in response to receiving a first audio stream, a performance identifier for a performance associated with the first audio stream is determined; and a reference audio is received based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a start point that the reference audio has, and the generation of the synthesized first audio stream is also based on the time offset. In some embodiments, generating a synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and the generation of the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

[0012] This application advantageously describes a metaverse engine and / or metaverse application that, based on the audio characteristics of a first audio stream, generates a synthesized first audio stream that predicts future portions of a performance, and blends the synthesized first audio stream with a second audio stream associated with a second client device to form a combined audio stream of synchronously synthesized first and second audio streams, thereby making the user generating the audio stream appear to be performing or speaking synchronously. The generation of the synthesized first audio stream and the blending of the synthesized first and second audio streams can be performed on different devices, including combinations of a first audio device, a second audio device, and a server. Therefore, the method is distributed across multiple devices, and any delays in the streaming, synthesis, or blending process are hidden through the blending step, so that users listening to the performance do not perceive any delay. Attached Figure Description

[0013] Figure 1 This is a block diagram of an example network environment for audio synthesis for synchronous communication according to some embodiments described herein.

[0014] Figure 2 This is a block diagram of an example computing device for audio synthesis for synchronous communication according to some embodiments described herein.

[0015] Figure 3A This is a block diagram of an example architecture of a machine learning model based on some embodiments described herein.

[0016] Figure 3B This is a block diagram of another example architecture of a machine learning model based on some embodiments described herein.

[0017] Figure 4 This is an example user interface that guides a user to change the timing aspects of a performance, according to some embodiments described herein.

[0018] Figure 5 This is an example flowchart illustrating data transmission between a client device and a server according to some embodiments described herein.

[0019] Figure 6 This is an example flowchart of a synthesized audio stream based on some embodiments described herein.

[0020] Figure 7 This is an example flowchart of synthesizing an audio stream for synchronous communication according to some embodiments described herein.

[0021] Figure 8 This is an example flowchart illustrating the use of a server to synthesize audio for synchronized communication, based on some embodiments described herein. Detailed Implementation

[0022] Network environment 100

[0023] Figure 1 A block diagram of an example environment 100 for audio synthesis in synchronous communication is shown. In some embodiments, environment 100 includes a server 101 and client devices 115a…n coupled via a network 105. Users 125a…n may be associated with their respective client devices 115a…n. Figure 1 In the accompanying drawings, letters following reference numerals (e.g., "115a") indicate references to elements having that particular reference numeral. Reference numerals without a letter following them (e.g., "115") indicate general references to embodiments of the elements having that reference numeral. In some embodiments, environment 100 may include Figure 1 Other servers or devices not shown. For example, server 101 may be multiple servers 101.

[0024] Server 101 includes one or more servers, each server including a processor, memory, and network communication hardware. In some embodiments, server 101 is a hardware server. Server 101 is communicatively coupled to network 105. In some embodiments, server 101 sends data to and receives data from client device 115. Server 101 may include a metaverse engine 103 and a database 199.

[0025] In some embodiments, the metaverse engine 103 includes code and routines for facilitating communication between two or more user-associated client devices 115 within a virtual metaverse, such as between friends in the same location within the metaverse, within the same metaverse experience, or within a metaverse application. In the metaverse, users interact with people who have different characteristics (e.g., different ages, regions, languages, etc.).

[0026] In some embodiments, the metaverse engine 103 performs some or all of the following steps: generating a synthesized first audio stream based on audio features of a first audio stream to predict future portions of the performance, and mixing the synthesized first audio stream with a second audio stream associated with a second client device to form a combined audio stream of the synchronously synthesized first and second audio streams. Different embodiments regarding the division of steps will be discussed in detail below.

[0027] In some embodiments, the Metaverse Engine 103 is implemented in hardware, including a Central Processing Unit (CPU), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), any other type of processor, or any combination thereof. In some embodiments, the Metaverse Engine 103 is implemented using a combination of hardware and software.

[0028] Database 199 may be a non-transitory computer-readable storage device (e.g., random access memory), cache, drive (e.g., hard disk drive), flash drive, database system, or another type of component or device capable of storing data. Database 199 may also include multiple storage components (e.g., multiple drives or multiple databases) distributed across multiple computing devices (e.g., multiple server computers). Database 199 may store data associated with Metaverse Engine 103, such as datasets for training machine learning models, reference audio, etc.

[0029] Client device 115 may be a computing device including memory and a hardware processor. For example, client device 115 may include a mobile device, tablet computer, mobile phone, wearable device, head-mounted display, mobile email device, portable game console, portable music player, e-reader device, or other electronic device capable of accessing network 105.

[0030] Client device 115a includes metaverse application 104a, and client device 115n includes metaverse application 104b. In some embodiments, user 125a uses metaverse application 104a on client device 115a to generate communication (e.g., a first audio stream), which is sent to metaverse engine 103 on server 101. Server 101 then sends the communication to metaverse application 104b on client device 115b of user 125n.

[0031] In some embodiments, the metaverse application 104 performs some or all of the following steps: generating a synthesized first audio stream based on the audio features of a first audio stream to predict future portions of the performance, and mixing the synthesized first audio stream with a second audio stream associated with a second client device to form a combined audio stream of the synchronously synthesized first and second audio streams. For example, the metaverse application 104a on client device 115a can generate the synthesized first audio stream, and the metaverse application 104b on client device 115n can mix the synthesized first audio stream with the second audio stream. In other embodiments, the synthesis and mixing steps can both be performed on server 101, and the metaverse application 104b on client device 115n outputs the mixed synthesized first and second audio streams through a speaker.

[0032] In the illustrated embodiment, entities in environment 100 are communicatively coupled via network 105. Network 105 may include public networks (e.g., the Internet), private networks (e.g., local area networks (LANs) or wide area networks (WANs)), wired networks (e.g., Ethernet), and wireless networks (e.g., 802.11 networks). Network or wireless LAN (WLAN), cellular network (e.g., Long Term Evolution (LTE) network), router, hub, switch, server computer, or a combination thereof. Although Figure 1 The diagram shows a network 105 coupled to server 101 and client device 115, but in practice, one or more networks 105 may be coupled to these entities.

[0033] Example 200 of computing devices

[0034] Figure 2 This is a block diagram of an example computing device 200, which can be used to implement one or more features described herein. The computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In some embodiments, the computing device 200 is a server 101. In some embodiments, the computing device 200 is a client device 115.

[0035] In some embodiments, computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a microphone 241, a speaker 243, a display 245, and a storage device 247, each coupled via a bus 218. Depending on whether computing device 200 is server 101 or client device 115, some components of computing device 200 may be absent. For example, in an instance where computing device 200 is server 101, the computing device may not include microphone 241 and speaker 243. In some embodiments, computing device 200 includes... Figure 2 Additional components not shown.

[0036] The processor 235 can be coupled to the bus 218 via signal line 222, the memory 237 can be coupled to the bus 218 via signal line 224, the I / O interface 239 can be coupled to the bus 218 via signal line 226, the microphone 241 can be coupled to the bus 218 via signal line 228, the speaker 243 can be coupled to the bus 218 via signal line 230, the display 245 can be coupled to the bus 218 via signal line 232, and the storage device 247 can be coupled to the bus 218 via signal line 234.

[0037] Processor 235 includes an arithmetic logic unit, a microprocessor, a general-purpose controller, or other processor array for performing calculations and providing instructions to a display device. Processor 235 processes data and may include various computing architectures, including Complex Instruction Set Computer (CISC) architecture, Reduced Instruction Set Computer (RISC) architecture, or architectures that implement combinations of instruction sets. Although Figure 2 A single processor 235 is shown, but multiple processors 235 may be included. In different embodiments, processor 235 may be a single-core processor or a multi-core processor. Other processors (such as a graphics processing unit), operating system, sensors, display, and / or physical configuration may be part of computing device 200.

[0038] Memory 237 stores instructions and / or data that can be executed by processor 235. These instructions may include code and / or routines for performing the techniques described herein. Memory 237 may be a Dynamic Random Access Memory (DRAM) device, Static RAM, or other memory device. In some embodiments, memory 237 may also include non-volatile memory, such as a Static Random Access Memory (SRAM) device or flash memory, or similar permanent storage devices and media, including hard disk drives, Compact Disc Read Only Memory (CD-ROM) devices, DVD-ROM devices, DVD-RAM devices, DVD-RW devices, flash memory devices, or other high-capacity storage devices for long-term storage of information. Memory 237 includes code and routines that can be used to execute the Metaverse Engine 103, as described in more detail below.

[0039] I / O interface 239 can provide functionality that enables computing device 200 to connect to other systems and devices. The interface device can be part of computing device 200 or separate from and communicate with computing device 200. For example, network communication devices, storage devices (such as memory 237 and / or storage device 247), and input / output devices can communicate via I / O interface 239. In another example, I / O interface 239 can receive data from server 101 and transfer data to metaverse engine 103 and its components (such as synthetic machine learning module 204). In some embodiments, I / O interface 239 can be connected to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone 241, sensors, etc.) and / or output devices (display device, speaker 243, monitor, etc.).

[0040] Examples of interface devices that can be connected to I / O interface 239 include display 245, which can be used to display content (such as images, videos, and / or user interfaces of the output applications described herein) and receive touch (or gesture) input from a user. Display 245 may include any suitable display device, such as a liquid crystal display (LCD), a light-emitting diode (LED) or plasma display, a cathode ray tube (CRT), a television, a monitor, a touch screen, a 3D display, or other visual display device.

[0041] Microphone 241 includes hardware for detecting audio generated by user 125. For example, microphone 241 can detect user 125 singing, user 125 playing the violin, etc. Microphone 241 can transmit audio to metaverse engine 103 via I / O interface 239.

[0042] Speaker 243 includes hardware for generating playback audio. For example, speaker 243 receives instructions from metaverse engine 103 to generate perceptible playback audio based on a digital composite audio stream generated by metaverse engine 103. Speaker 233 translates the instructions into audio and generates a composite audio stream for the user.

[0043] Storage device 247 stores data related to the metaverse engine 103. For example, storage device 247 may store training datasets for training machine learning models, reference audio, etc. In embodiments where computing device 200 is server 101, storage device 247 is equivalent to... Figure 1 Database 199.

[0044] Example Metaverse Engine 103 or Example Metaverse Application 104

[0045] Figure 2 A computing device 200 is shown executing an example metaverse engine 103 or an example metaverse application 104, which includes a performance recognition module 202, a synthesis machine learning module 204, a mixing module 206, a post-processing module 208, and a user interface module 210. Although these modules are exemplified as part of the same metaverse engine 103 or the same metaverse application 104, those skilled in the art will understand that these modules can be implemented by any computing device 200. For example, the performance recognition module 202 and the synthesis machine learning module 204 may be part of a client device 115, while the mixing module 206 may be part of a server 101 to reduce the computational requirements of the client device 115.

[0046] The performance recognition module 202 determines a performance identifier associated with the audio stream. In some embodiments, the performance recognition module 202 includes a set of instructions executable by the processor 235 to determine the performance identifier associated with the audio stream. In some embodiments, the performance recognition module 202 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0047] An audio stream can be a performance or interpretation of known content (e.g., a re-recording of a known song, singing a known melody, reading text material, etc.). In this context, a performance identifier refers to the known content being performed in the audio stream. Performance identifiers enable the identification and / or prediction of various attributes of the audio stream. For example, if the performance is based on written musical notation (including notations of rhythm and sounds played by various instruments), the rhythm of the performance and the individual sounds played by each performer can be identified. In another example, if the performance includes singing or reading based on text, upcoming words / phrases can be identified.

[0048] In some embodiments, the performance recognition module 202 receives an audio stream after obtaining user permission. For example, when the performance recognition module 202 is part of client device 115, it receives the audio stream from microphone 241 via I / O interface 239 through network 105. In another example, when the performance recognition module 202 is part of server 101, it receives the audio stream from client device 115 via I / O interface 239. This audio stream is part of a performance. The performance may be a song sung by multiple users, a speech, recitation, music from an instrumental performance, etc. In some embodiments, the user interface module 210 obtains permission to use the audio stream before the performance recognition module 202 receives it.

[0049] The performance recognition module 202 generates a fingerprint (i.e., an audio fingerprint) from at least a portion of the audio stream by generating a spectrogram of the audio stream whose frequency varies over time and which has the highest amplitude frequency in the audio stream. Then, it generates a hash value for this spectrogram. Since the audio stream is performed live (rather than pre-recorded), the performance recognition module 202 needs to consider pitch differences, including timing inconsistencies and timbre characteristics. The performance recognition module 202 compares the fingerprint of the audio stream with a set of fingerprints for known songs to identify a match, which is associated with the performance identifier of the audio stream. For example, the performance identifier could be "Happy Birthday" or "Moonlight Sonata, Third Movement."

[0050] The performance recognition module 202 can receive reference audio associated with a performance identifier. For example, the performance recognition module 202 can retrieve the reference audio from the storage device 247.

[0051] After obtaining user permission, the performance recognition module 202 can determine the performance type of the audio stream. For example, the audio stream may include a solo singing, a solo speech, a solo instrumental performance, etc. In some embodiments, the performance recognition module 202 determines the performance type based on unique aspects of different instruments and vocals, such as recognizing different frequencies, timings, pitches, tempos, etc. For example, the performance recognition module 202 can determine that in the metaverse, "Happy Birthday" is performed with a kazoo, while the third movement of "Moonlight Sonata" is performed with a piano. In some embodiments, the performance recognition module 202 can also determine the style of the performance, such as whether the singer is performing a blues version of the song or a rock version of the song, etc.

[0052] In some embodiments, the performance recognition module 202 obtains permission from user 125 to identify one or more speakers in the audio stream and associates the one or more speakers with speaker identifiers. This permission may include permission to use the audio stream for recognition purposes, permission to store user information, etc. User 125 is provided with guidance that information can be stored (e.g., temporarily stored until the end of the performance), and is offered options to deny permission and select storage attributes (e.g., local storage only, storage for X hours, etc.). After obtaining permission, the performance recognition module 202 may, for example, determine that the audio stream associated with the first client device 115 includes a solo performance of "Happy Birthday" or a multi-person performance of the song. The performance recognition module 202 may identify different performers of the song based on a mechanism similar to that described above, where each performer is associated with a specific speaking or singing style based on rhythm, pitch, frequency, etc.

[0053] The synthesis machine learning module 204 trains one or more machine learning models to synthesize audio streams. In some embodiments, the synthesis machine learning module 204 includes a set of instructions executable by the processor 235 to train the machine learning models to synthesize audio streams. In some embodiments, the synthesis machine learning module 204 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0054] The synthetic machine learning module 204 may use one or more (e.g., two) different training datasets. In some embodiments, the datasets and their corresponding structures may depend on whether the synthetic machine learning module 204 is stored on the client device 115 or the server 101.

[0055] Example machine learning module 204 on client device 115

[0056] In an embodiment where the synthetic machine learning module 204 is stored in the client device 115, the synthetic machine learning model 204 is trained using a training dataset of manually labeled audio streams to achieve supervised learning.

[0057] The synthetic machine learning module 204 may be a neural network such as a deep neural network (DNN), which includes layers that identify progressively refined features and patterns in the audio stream, where the output of one layer serves as the input to subsequent layers. The output layer generates the synthesized audio stream. In some embodiments, a first layer (or a first set of layers) outputs a mapping of the audio stream to a reference audio of the performance to determine the position of the audio stream relative to the reference audio, a second layer (or a second set of layers) outputs a prediction of the future time offset of the audio stream, and the output layer synthesizes the audio stream based on the outputs of the previous layers. In some embodiments, the machine learning model may include Long Short-Term Memory (LSTM) nodes in a recurrent neural network (RNN) trained for sequential processing of the audio stream. This machine learning model can receive any audio stream as input and perform the same analysis to synthesize the output.

[0058] The synthesized audio stream takes into account network latency, input latency, and / or processing latency that cause audio delay. Network latency occurs when there is a delay in content reception from the transmitting node to the receiving node in the network due to the use of physical devices in the network or the delay in processing time at the node. Input latency occurs when there is a delay in user response, or when the user is using a client device with limited functionality. Processing latency occurs when latency is intentionally introduced for conditioning analysis or when processing resources are insufficient. Because network latency is variable, the actual end-to-end latency of each audio stream can only be determined by the client device 115, where the audio stream is played back through the speaker 243.

[0059] In some embodiments, the synthesis machine learning module 204 advantageously considers different types of input delays. For example, a performance might be Vivaldi's *The Four Seasons*, where the first audio stream is a violin performance with a delay in the third note (determined using performance identifiers that support the identification of notes and beats in *The Four Seasons*). The machine learning model can adjust the first audio stream based on the input delay by skipping the sixteenth note in the output to compensate for the missed third note in the synthesized first audio stream. In another example, the synthesis machine learning module 204 can compensate for processing delays introduced by modulation by adding delays at the end of sentences in the speech. Adding delays at the end of sentences can be advantageous because delays at the end of sentences have a lower impact on the listener's perception compared to delays at any point within a sentence.

[0060] Go to Figure 3A , Figure 3A A block diagram of an example architecture for a machine learning model 300 is shown. The machine learning model 300 includes a bottleneck backbone 305, a sub-model decoder 310, and an audio waveform generator 315. The bottleneck backbone 305 is a general bottleneck backbone used to extract features from the input waveform of the audio stream. The sub-model decoder 310 is trained for a specific task: determining the temporal relevance of the audio stream compared to a reference audio (e.g., determined based on performance identifiers).

[0061] The bottleneck backbone 305 and sub-model decoder 310 are designed similarly to machine learning models that implement speech-to-text synthesis. The synthesis machine learning module 204 can train the bottleneck backbone 305 using a large number of, for example, millions of {audio, text} samples to output bottleneck features. For example, the bottleneck backbone 305 can be trained in a manner similar to the DeepSpeech system, but instead of being trained to predict text, it uses previous layers that store embeddings of audio features mapped to reference audio.

[0062] The synthetic machine learning module 204 can train the sub-model decoder 310 by freezing the weights of the bottleneck backbone 305 after training is complete, and provide the sub-model decoder 310 with bottleneck features and audio features (such as pitch, phase, and rate) extracted from the audio stream input waveform. In some embodiments, the sub-model decoder 310 also receives speaker identifiers for distinguishing different speech in the audio stream.

[0063] Once the bottleneck backbone 305 and the sub-model decoder 310 have been trained, the bottleneck backbone 305 receives the audio stream generated by the client device 115 as input. In some embodiments, the bottleneck backbone 305 includes a fully connected (FC) layer followed by an LSTM layer, which is used to segment the audio stream into progressively abstract audio data representations.

[0064] The bottleneck backbone 305 receives the audio stream as the input waveform and Mel-frequency cepstral coefficients (MFCC) features sampled from overlapping samples of the audio stream. MFCC features are a cepstral representation of the audio stream used to represent audio stream information (such as timbre representation). Audio stream samples are parameterized using time windows. For example, time windows shorter than 100 milliseconds can be used; while less accurate than longer time windows, they are computationally efficient and suitable for real-time execution on client device 115. Although the computational processing power of client device 115 is lower than that of server 101, the network latency is less than that incurred when processing the audio stream on server 101, allowing client device 115 to still generate accurate results using short time windows.

[0065] The bottleneck backbone 305 encodes the audio and outputs bottleneck features. These bottleneck features are then transmitted to the sub-model decoder 310. In some embodiments, the bottleneck backbone 305 is a subset of audio frames; for example, bottleneck features are output for one frame out of every five frames in the audio stream. In some embodiments, the sub-model decoder 310 also receives audio-specific features associated with the input waveform, such as the pitch, phase, and rate of the audio stream, as well as a speaker identifier.

[0066] In some embodiments, the sub-model decoder 310 includes a fully connected FC layer followed by an LSTM layer. The LSTM layer is used to obtain the temporal correlation of the audio stream. The sub-model decoder 310 outputs MFCC features, which are used as input to the audio waveform generator 315.

[0067] The audio waveform generator 315 receives MFCC features and generates an output waveform that presents the MFCC features as synthesized audio.

[0068] Throughout the process, machine learning model 300 learns the temporal mapping between positions in the audio stream and synthesizes future frames of the audio stream. For example, if the performance is "Happy Birthday," machine learning model 300 can map the audio stream to the reference audio by determining that the audio stream includes the first two lines of the song and by identifying the difference between the audio stream and the reference audio when the audio stream's starting point differs from the reference audio. This is called the temporal offset between the audio stream and the reference audio. Machine learning model 300 can also output the rate of the audio stream compared to the reference audio. For example, a user might sing "Happy Birthday" faster than the reference audio. Machine learning model 300 can predict the timing of the next frame of the audio stream based on the temporal offset, the rate of the audio stream, and the rate of the song's future temporal offset based on the rate of the audio stream compared to the reference audio, thereby outputting a synthesized audio that incorporates these features.

[0069] In some embodiments, the machine learning model 300 can implement a multilingual bottleneck extractor instead of performing the bottleneck extraction described above. The multilingual bottleneck extractor is trained to distinguish senones in multiple languages. The output features are language-independent and robust to variations caused by language, speaking style, speech rate, etc. The machine learning model 300 is trained using supervised learning, and its output includes senone posteriors.

[0070] Example machine learning model on client device 115

[0071] In an embodiment where the synthetic machine learning module 204 is stored on server 101, the synthetic machine learning model 204 can be trained using unsupervised learning with a training dataset containing unlabeled audio streams.

[0072] Go to Figure 3B , Figure 3B A block diagram of another example architecture of machine learning model 350 is shown. Machine learning model 350 includes a vector quantization variational autoencoder (VQ-VAE) 355, a VQ-VAE codebook 360, a prior model 365, and a VQ-VAE decoder 370.

[0073] The synthetic machine learning module 204 trains VQ-VAE 355 and VQ-VAE codebook 360 using a training dataset including audio stream waveforms. VQ-VAE 355 and VQ-VAE codebook 360 train corresponding modules of the machine learning model 350 by comparing output waveforms with input waveforms. The synthetic machine learning module 204 trains a prior model 365 on the code vector output from the VQ-VAE codebook 360, the code vector output depending on additional features such as one or more audio features including pitch, phase, and rate, and optionally, a speaker identifier.

[0074] Machine learning model 350 is trained to receive audio streams parameterized by time windows. In some embodiments, the audio stream length is 300-500 milliseconds. Longer audio streams require greater computation and larger model sizes, but produce more accurate results.

[0075] In some embodiments, the VQ-VAE 355 receives an input waveform of an audio stream and converts it into a latent representation. The VQ-VAE 355 employs an autoregressive network structure comprising multiple one-dimensional (1D) convolutional blocks. The VQ-VAE codebook 360 receives the latent representation as input, and the bottleneck quantizes the latent representation into discrete code vectors using a predefined codebook.

[0076] The prior model 365 receives discrete code vectors from the VQ-VAE codebook 360, along with audio features such as pitch, phase, and rate of the input waveform, and (optionally) one or more speaker identifiers from the audio stream. Audio features are computed based on the input waveform, and the one or more speaker identifiers allow the network to model user- or genre-specific codes. The prior model 365 uses a transform layer to extrapolate the code vectors in time and outputs edited, resampled code vectors. The VQ-VAE decoder 370 is an autoregressive audio waveform generator. The VQ-VAE decoder 370 receives the edited, resampled code vectors and synthesizes an output waveform based on them.

[0077] The mixing module 206 mixes audio streams from multiple client devices 115 to form a combined audio stream that synchronizes one or more synthesized audio streams with real audio streams. In some embodiments, the mixing module 206 includes a set of instructions executable by the processor 235 to form the combined audio stream. In some embodiments, the mixing module 206 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0078] In some embodiments, the mixing module 206 receives the output of the synthesis machine learning module 204 and synchronizes the audio streams based on that output. For example, the mixing module 206 may receive and synthesize a first audio stream and a second audio stream to generate a combined audio stream. Depending on whether the mixing is performed on server 101 or client device 115, the mixing module 206 may introduce different amounts of latency into the different audio streams to compensate for various delays and ensure audio stream synchronization. For example, if the mixing is performed on server 101, the mixing module 206 needs to consider the latency of the receiver. The different factors of audio stream mixing will be discussed in more detail below.

[0079] The mixing module 206 synchronizes the synthesized first and second audio streams to generate a combined audio stream. The time window length of the audio stream may vary depending on the device storing the synthesis machine learning module 204. For example, when the synthesis machine learning module 204 is stored on the client device 115, the time window is less than 100 milliseconds. As another example, when the synthesis machine learning module 204 is stored on the server 101, the time window is 300-500 milliseconds.

[0080] The mixing module 206 can be synchronized on the same or a different device than the device where the synthesis machine learning module 204 is stored. The device may include server 101, client device 115a that sends the first audio stream, and client device 115b that receives the first audio stream.

[0081] In the first example, client device 115b performs audio stream synthesis and synchronization. Synthesis machine learning module 204 receives global timestamp data packets from all other client devices 115 via server 101 and synthesizes the audio streams from the other client devices 115. Mixing module 206 generates a local mix based on the real audio stream from client device 115b and the synthesized audio from the other client devices 115.

[0082] In some embodiments, the first example is a preferred example. The first example has several advantages. The time required to predict future portions of the audio stream is reduced, thereby improving the quality of the audio stream and reducing interactive latency. Furthermore, the client device 115b is able to recover from high latency. Although a client device with powerful computing capabilities is required for local synthesis and synchronization of audio, better hardware can provide a better experience. Finally, a benefit of using this architecture is that the synthesized audio remains on the client device 115b.

[0083] Some potential drawbacks of the first example are that the computational complexity is O(n) for an input stream of n performers. Furthermore, since all processing is performed on client device 115b and involves a large amount of processing, this could potentially deplete the battery of client device 115b. These processes may also be complex, requiring different implementations for different types of client devices 115, and can only run on a subset of client devices 115 with sufficient computing power.

[0084] In the second example, server 101 performs the synthesis and synchronization of the audio streams. Since all client streams are received and processed at the server, the advantages of the second example include lower computational complexity O(1). In this example, there is no need for computation on the client side to generate the combined audio stream because server 101 receives the independent audio stream from each client and provides each client with a synthesized and synchronized combined audio stream. The infrastructure is controlled because it is all part of server 101. Furthermore, even in the worst-case scenario, the prediction time is shorter than that of client-side processing because even in the worst-case scenario, the audio stream latency is only the latency between the client and the server. In client-side processing, the latency could be the latency between client device 115a and server 101, and then between server device 115b.

[0085] In the third example, client device 115a performs audio stream synthesis, while server 101 performs audio stream synchronization. Client device 115a synthesizes audio streams with worst-case latency, while server 101 synchronously mixes (n-1) streams at the same time point by delaying each audio stream as needed. The advantages of this third example include: O(1) input and output streams for each client device 115, and O(1) processing time at the server. Furthermore, better hardware on client device 115a results in better sound quality than other client devices 115. For example, better hardware leads to superior processing, which is especially critical in scenarios requiring fast audio stream processing to maintain near real-time audio transmission.

[0086] In the fourth example, client device 115a performs the synthesis of the audio stream, and client device 115b performs the synchronization of the audio stream. This is suitable for serverless peer-to-peer streaming.

[0087] In instances where post-processing is not performed and mixing module 206 is not performing synthesis on client device 115b, mixing module 206 instructs I / O interface 239 to send the combined audio stream to client device 115b, which can play the combined audio stream. In instances where post-processing is not performed and mixing module 206 is located on client device 115b, mixing module 206 instructs I / O interface 239 to provide the combined audio stream to speaker 243 for playback.

[0088] Post-processing module 208 processes the combined audio stream. In some embodiments, post-processing module 208 includes a set of instructions executable by processor 235 to process the combined audio stream. In some embodiments, post-processing module 208 is stored in memory 237 of computing device 200 and is accessible and executable by processor 235.

[0089] In some embodiments, the post-processing module 208 can modify the combined audio stream to match the acoustics of the environment in which the client device 115b is located, taking into account echoes, background noise, etc. In some embodiments, the post-processing module 208 performs audio cleanup or noise suppression on the combined audio stream to avoid situations where, for example, the first audio stream includes background car horns while the space where the second audio stream is located is completely silent, resulting in a dissonance caused by the combined audio stream suddenly switching from having background noise to having no background noise.

[0090] User interface module 210 generates a user interface. In some embodiments, user interface module 210 includes a set of instructions executable by processor 235 to generate the user interface. In some embodiments, user interface module 210 is stored in memory 237 of computing device 200 and is accessible and executable by processor 235.

[0091] User interface module 210 generates a user interface for user 125 associated with client device 115. This user interface can be used to initiate audio communication with other users, participate in games or other experiences in the metaverse, send text messages to other users, initiate video calls with other users, etc.

[0092] In some embodiments, before a user joins the metaverse, the user interface module 210 generates a user interface that includes information about how user information is collected, stored, and analyzed. For example, the user interface requires the user to grant permission to use any information related to the user. The user is informed that they can delete their user information and can also choose which types of information to provide for different purposes. The use of information complies with relevant regulations, and the data is stored securely. Data collection is not performed in specific areas or for specific user categories (e.g., based on age or other demographic characteristics). Data collection is temporary (i.e., data is periodically deleted), and data is not shared with third parties. Some data may be anonymized, aggregated across users, or modified to make it impossible to identify a specific user.

[0093] In some implementations, the user interface obtains user permission before sending the audio stream to server 101 or other client device 115. The user interface may include different levels of granularity for user permissions. For example, a user may specify that synthesized audio can only be generated on client device 115 and not on the server.

[0094] In some embodiments, the mixing module 206 determines that the time difference between the first audio stream and the second audio stream exceeds a threshold time difference and sends an instruction to the user interface module 210 to generate a user interface. This user interface can guide the user 125 on how to change the singing rate so that the audio stream can be synchronized with other audio streams.

[0095] Go to Figure 4 , Figure 4 An example user interface 400 is shown that guides the user to change the timing aspects of a performance. In this example, the user interface 400 includes user guidance for the performance and a movement indicator 405 that prompts the performer associated with the client device 115b to perform in a way that reduces the time difference between the first audio stream and the second audio stream. For example, the movement indicator 405 could move slower than the user's performance to suggest that the user slow down the performance rate.

[0096] In some embodiments, the user interface module 210 illustrates the synchronization of the combined audio stream with the graphically displayed actions of a performer. For example, if the combined audio stream is a speech performed by a user as a graphically displayed virtual character, the user interface module 210 can synchronize the combined audio stream with the virtual character's lip movements, actions, etc.

[0097] Example Method

[0098] Figure 5As an example flowchart 500, it illustrates data transmission between a client device 115 and a server 101 according to some embodiments described herein. Flowchart 500 includes a first client device 510, a server 515, and a second client device 520. Thick lines represent network data transmission between the three devices, and thin lines represent data transmission within the first client device 510.

[0099] First client device 510 receives an audio stream from microphone 505. This audio stream is synthesized at first client device 510, server 515, or second client device 520. The synthesized audio stream is mixed with one or more other audio streams at server 515 or second client device 520 to form a combined audio stream. The combined audio stream is sent to speaker 525 for playback at first client device 510.

[0100] In addition, server 515 also receives data. For example, server 515 receives a first stream from first client device 510 and a second stream from second client device 520, wherein server 515 synthesizes the first audio stream and mixes the first audio stream with the second audio stream.

[0101] By implementing the metaverse engine 103, the first audio stream of the violin and the second audio stream of the cello on the first client device 510 can be heard simultaneously by the associated user performing the music, as if they were in the same physical space.

[0102] Figure 6 This is an example flowchart 600 for synthesizing an audio stream used for synchronous communication. Task detection 605 is performed on the delayed audio stream to identify performance identifiers and performance types. For example, it identifies whether the delayed audio stream includes instruments, speech, singing, etc., as well as the type of song and the style of performance. Based on this identification, reference audio is output.

[0103] The delayed audio stream is also received by phase-shift analyzer 610 and rate analyzer 615. Phase-shift analyzer 610 identifies the time offset of the delayed audio stream compared to the reference audio, and rate analyzer 615 identifies the rate of the delayed audio stream. The time offset and rate of the delayed audio stream are received by sampler 620, which maps the audio stream to the reference audio. This mapping is received by deep neural network 635 (or other suitable model), which extrapolates the audio stream into the future to complete the synthesis of the audio stream.

[0104] Figure 7 This is an example flowchart of synthesizing an audio stream for synchronous communication using a second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.

[0105] Method 700 may begin at box 702. In box 702, a first audio stream of the performance associated with the first client device 115 is received. Box 702 may be followed by box 704.

[0106] In box 704, during the performance time window, where the time window is less than the total performance time: based on the audio features of the first audio stream, a synthesized first audio stream predicting future portions of the performance is generated; and the synthesized first audio stream and a second audio stream associated with a second client device are mixed to form a combined audio stream of the synchronously synthesized first and second audio streams, wherein the time window is shifted forward and the generation and synchronization are repeated until the performance is complete.

[0107] Figure 8 This is another example flowchart of using a server to synthesize audio for synchronized communication. In this example, the metaverse engine 103 is stored on server 101.

[0108] Method 800 may begin at box 802. In box 802, a first audio stream of the performance associated with the first client device 115 is received. Box 802 may be followed by box 804.

[0109] In box 806, within a performance time window, wherein the time window is less than the total performance time: based on the audio features of the first audio stream, a synthesized first audio stream predicting future portions of the performance is generated; and the synthesized first audio stream and a second audio stream associated with the second client device are blended to form a combined audio stream of the synchronized synthesized first and second audio streams, wherein the time window is shifted forward and generation and synchronization are repeated until the performance is complete. In some embodiments, blending the synthesized first audio stream includes introducing a delay into the combined audio stream to handle delays occurring during the transmission of the combined audio to the second client device. Box 806 may be followed by box 808.

[0110] In box 808, the combined audio stream is sent to the second client device 115.

[0111] The various embodiments described herein include acquiring data from various sensors in the physical environment, analyzing such data, generating recommendations, and providing a user interface. Data collection is conducted only with the specific user's permission and in accordance with applicable regulations. Data storage complies with applicable regulations, including anonymizing or otherwise modifying data to protect user privacy. Users are provided with clear information about data collection, storage, and use, and are given options to select the types of data that can be collected, stored, and used. Furthermore, users control the devices that can store data (e.g., user-only devices, client and server devices, etc.) and the devices that perform data analysis (e.g., user-only devices, client and server devices, etc.). Data is used for the specific purposes described herein. No data will be shared with third parties without the user's explicit permission.

[0112] The methods, blocks, and / or operations described herein may be performed in a different order than those shown or described, and / or, where appropriate, simultaneously (partially or completely) with other blocks or operations. Some blocks or operations may be performed on a portion of data and later, for example, on another portion of data. Not all described blocks and operations need to be performed in all implementations. In some implementations, blocks and operations may be performed multiple times in a different order and / or at different times within a method.

[0113] In the foregoing description, numerous specific details have been set forth for purposes of explanation to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described primarily with reference to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, as well as any external device providing services.

[0114] References to "some embodiments" or "some examples" in the specification mean that a particular feature, structure, or characteristic described in conjunction with an embodiment or example may be included in at least one implementation of the specification. The phrase "in some embodiments" appearing in different places in the specification does not necessarily refer to the same embodiment.

[0115] Some of the parts described in detail above are presented as algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate the substance of their work to others skilled in the art. Algorithms here are generally considered as a series of self-consistent steps that achieve a desired result. These steps are steps that require physical manipulation of physical quantities. Typically, though not always, these quantities appear in the form of electrical or magnetic data that can be stored, transmitted, combined, compared, and otherwise manipulated. It has proven convenient, primarily for general reasons, to sometimes refer to these data as bits, values, elements, symbols, characters, terms, numbers, or similar names.

[0116] However, it should be noted that all these terms and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless explicitly stated in the following discussion, throughout the description, the use of terms such as 'processing,' 'computation,' 'derive,' and 'display' refers to the actions and processes of a computer system or similar electronic computing device. These actions and processes involve the manipulation and transformation of data, represented in the form of physical (electronic) quantities, within the registers and memory of the computer system, into other data, also represented in the form of physical quantities, stored in the computer system's memory or registers or other information storage, transmission, or display devices.

[0117] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including but not limited to any type of disk, including optical discs, ROMs, CD-ROMs, magnetic disks, RAM, EPROMs, EEPROMs, magnetic cards or optical cards, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each medium being coupled to a computer system bus.

[0118] The specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment that includes both hardware and software elements. In some embodiments, the specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0119] Furthermore, this description may also take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by, or associated with, a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium may be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device.

[0120] A data processing system suitable for storing or executing program code will include at least one processor that is directly or indirectly connected to memory elements via a system bus. Memory elements may include local memory used during actual program code execution, mass storage, and cache memory, with the cache memory providing at least some temporary storage for program code to reduce the number of times code must be retrieved from mass storage during execution.

Claims

1. A computer-implemented method, comprising: Receive the first audio stream of the performance associated with the first client device; During the performance's time window, wherein the time window is less than the total performance time: Based on the audio features of the first audio stream, a synthetic first audio stream is generated to predict future portions of the performance; and The synthesized first audio stream and the second audio stream associated with the second client device are mixed to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream; The time window is shifted forward, and the generation and mixing are repeated until the performance is complete.

2. The method of claim 1, further comprising: In response to receiving the first audio stream, a performance identifier for the performance associated with the first audio stream is determined; as well as Reference audio is received based on the performance identifier.

3. The method as described in claim 2, characterized in that: Generating the synthesized first audio stream includes determining the time offset between the first audio stream and the reference audio; as well as The time offset occurs when the first audio stream has a different starting point than the reference audio, and the generation of the synthesized first audio stream is also based on the time offset.

4. The method as described in claim 2, characterized in that: Generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and The generation of the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio.

5. The method of claim 1, wherein, The audio features of the first audio stream are selected from a group consisting of pitch, velocity, phase, or any combination thereof.

6. The method of claim 1, wherein, The audio features of the first audio stream include one or more speaker identifiers detected in the first audio stream.

7. The method of claim 1, further comprising: Determine that the time difference between the first audio stream and the second audio stream exceeds a threshold time difference; as well as Graphical data is generated for displaying a user interface, which includes user instructions and movement indicators for the performance, the movement indicators prompting the performer associated with the second client device to perform in a manner that reduces the time difference between the first audio stream and the second audio stream.

8. The method of claim 1, further comprising modifying the combined audio stream to match the acoustic effects of the environment in which the second client device is located.

9. The method of claim 1, wherein, Generating the synthesized first audio stream includes: Identify that a portion of the performance was skipped in the first audio stream; and The first audio stream is synthesized to correct the skipped portions of the performance.

10. The method of claim 1, further comprising synchronizing the combined audio streams to match the performer’s graphically displayed movements.

11. An apparatus comprising: processor; as well as A memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations including: Receive the first audio stream of the performance associated with the first client device; as well as During the performance's time window, wherein the time window is less than the total performance time: Based on the audio features of the first audio stream, a synthetic first audio stream is generated to predict future portions of the performance; and The synthesized first audio stream and the second audio stream associated with the second client device are mixed to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream; The time window is shifted forward, and the generation and mixing are repeated until the performance is complete.

12. The device as claimed in claim 11, characterized in that: In response to receiving the first audio stream, a performance identifier for the performance associated with the first audio stream is determined; and Reference audio is received based on the performance identifier.

13. The device as described in claim 12, characterized in that: Generating the synthesized first audio stream includes determining the time offset between the first audio stream and the reference audio; as well as The time offset occurs when the first audio stream has a different starting point than the reference audio, and the generation of the synthesized first audio stream is also based on the time offset.

14. The device as claimed in claim 12, characterized in that: Generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and The generation of the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio.

15. The apparatus of claim 11, wherein, The audio features of the first audio stream are selected from a group consisting of pitch, velocity, phase, or any combination thereof.

16. A non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform an operation, the operation comprising: Receive the first audio stream of the performance associated with the first client device; as well as During the performance's time window, wherein the time window is less than the total performance time: Based on the audio features of the first audio stream, a synthetic first audio stream is generated to predict future portions of the performance; and The synthesized first audio stream and the second audio stream associated with the second client device are mixed to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream; The time window is shifted forward, and the generation and mixing are repeated until the performance is complete.

17. The computer-readable medium as claimed in claim 16, characterized in that: In response to receiving the first audio stream, a performance identifier for the performance associated with the first audio stream is determined; and Reference audio is received based on the performance identifier.

18. The computer-readable medium as claimed in claim 17, characterized in that: Generating the synthesized first audio stream includes determining the time offset between the first audio stream and the reference audio; as well as The time offset occurs when the first audio stream has a different starting point than the reference audio, and the generation of the synthesized first audio stream is also based on the time offset.

19. The computer-readable medium as claimed in claim 17, characterized in that: Generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and The generation of the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio.

20. The computer-readable medium as claimed in claim 16, characterized in that, The audio features of the first audio stream are selected from a group consisting of pitch, velocity, phase, or any combination thereof.