Audio synthesis for synchronous communications

By receiving and processing audio streams, generating synthetic audio streams and mixing them with other audio streams, synchronous communication problems caused by network delay are solved, and synchronous performance or speech effects between users are achieved.

CN120077430AActive Publication Date: 2025-05-30ROBLOX CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202380070876.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-10-02
Publication Date
2025-05-30
Estimated Expiration
2043-10-02

AI Technical Summary

Technical Problem

Due to the delay in transmission of computer networks, multiple people in different locations cannot perform music, chant or talk synchronously, and the existing technology is difficult to effectively solve this problem.

Method used

By receiving the performance audio stream associated with the first client device, a synthetic audio stream predicting future portions of the performance is generated based on the audio features and mixed with the audio stream associated with the second client device to form a synchronized combined audio stream until the performance is completed.

Benefits of technology

Synchronous performance or speech on different devices is achieved, hiding delays in streaming, synthesis or mixing, so that users listening to the performance will not perceive any delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077430A_ABST
    Figure CN120077430A_ABST
Patent Text Reader

Abstract

A computer-implemented method includes receiving a first audio stream of a performance associated with a first client device. The method further includes, during a time window of the performance, where the time window is less than a total time of the performance: generating a synthetic first audio stream predicting a future portion of the performance based on audio characteristics of the first audio stream, and mixing the synthetic first audio stream with a second audio stream associated with a second client device, generating and mixing the first audio stream and the second audio stream to form a combined audio stream that synchronously synthesizes the first audio stream and the second audio stream, where the time window is advanced and the generating and mixing are repeated until the performance is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application is an international application and claims priority under 35 U.S.C.§119(e) to a U.S. patent application filed on October 4, 2022, with application number 17 / 959,736 and title "SYNTHESIZING AUDIO FOR SYNCHRONOUS COMMUNICATION", the entire content of which is incorporated herein by reference. Background Art

[0002] For multiple people in different locations but connected by a computer network, it is impossible to perform music, recite, or talk in synchronization. This is because due to the transmission delay of the computer network, the performance of person P0 observed by person P1 is always in the past. If each of the n - 1 performers P1, P2, P3...P(n - 1) delays by an appropriate time precisely relative to performer P0, then P0 will observe that the others are synchronized with each other, but they themselves cannot be synchronized with the others.

[0003] This delay may be caused by one or more of at least the following three types of delays: network delay, input delay, and processing delay. Network delay occurs when the network transmission time is insufficient due to the use of physical devices in the network or the delay in processing time at nodes. Input delay occurs when the user responds late, the user is using a client device with limited functionality, etc. Processing delay occurs when delay is intentionally introduced for adjustment analysis.

[0004] The background art description provided herein is intended to introduce the background of the present disclosure. The work done by the current inventors, to the extent described in this background section, and aspects of the specification that may not constitute prior art at the time of filing, are not expressly or implicitly admitted to be prior art of the present disclosure. Summary of the Invention

[0005] Embodiments of the present application generally relate to a system and method for audio synthesis for synchronous communication. According to one aspect of the present application, a computer - implemented method includes: receiving a first audio stream of a performance associated with a first client device. The method further includes, during a time window of the performance, where the time window is less than the total time of the performance: generating a synthetic first audio stream that predicts a future part of the performance based on audio characteristics of the first audio stream; and mixing the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthetic first audio stream and the second audio stream; wherein the time window is shifted forward, and the generation and mixing are repeated until the performance is completed.

[0006] In some embodiments, the method further includes: in response to receiving a first audio stream, determining a performance identifier of a performance associated with the first audio stream, and receiving a reference audio based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthesized first audio stream is also based on the time offset. In some embodiments, generating a synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and generating the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio characteristics of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof. In some embodiments, the audio characteristics of the first audio stream include one or more speaker identifiers detected in the first audio stream. In some embodiments, the method further includes: determining that a time difference between the first audio stream and a second audio stream exceeds a threshold time difference; and generating graphical data for displaying a user interface, the user interface including user guidance for the performance and a movement indicator that prompts a performer associated with a second client device to perform in a manner that reduces the time difference between the first audio stream and the second audio stream. In some embodiments, the method further includes: modifying the combined audio stream to be consistent with the acoustic effects of the environment where the second client device is located. In some embodiments, generating a synthesized first audio stream includes: identifying that a part of the performance is skipped in the first audio stream; and synthesizing the first audio stream to correct the skipped part of the performance. In some embodiments, the method further includes: synchronizing the combined audio stream to match the graphically displayed actions of the performer.

[0007] In some embodiments, a device includes a processor and a memory coupled to the processor, and instructions are stored on the memory that, when executed by the processor, cause the processor to perform operations including: receiving a first audio stream of a performance associated with a first client device; during a time window of the performance, where the time window is less than the total time of the performance: generating a synthesized first audio stream that predicts a future part of the performance based on the audio characteristics of the first audio stream; and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream; wherein the time window is advanced, and the generating and mixing are repeated until the performance is completed.

[0008] In some embodiments, in response to receiving a first audio stream, a performance identifier of a performance associated with the first audio stream is determined; and reference audio is received based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a starting point different from that of the reference audio, and generating the synthesized first audio stream is also based on the time offset. In some embodiments, generating a synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and generating the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio characteristics of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

[0009] In some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations including: receiving a first audio stream of a performance associated with a first client device; during a time window of the performance, where the time window is less than the total time of the performance: generating a synthesized first audio stream that predicts a future portion of the performance based on the audio characteristics of the first audio stream; and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream; wherein the time window is advanced, and the generating and mixing are repeated until the performance is complete.

[0010] In some embodiments, in response to receiving a first audio stream, a performance identifier of a performance associated with the first audio stream is determined; and reference audio is received based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a starting point different from that of the reference audio, and generating the synthesized first audio stream is also based on the time offset. In some embodiments, generating a synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and generating the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio characteristics of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

[0011] The present application advantageously describes a metaverse engine and / or a metaverse application that generates a synthetic first audio stream predicting a future portion of a performance based on audio features of a first audio stream, and mixes the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthetic first audio stream and the second audio stream, so that the user generating the audio stream is perceived as performing or speaking synchronously. Generating the synthetic first audio stream and mixing the synthetic first audio stream with the second audio stream can be performed on different devices, including a combination of a first audio device, a second audio device, and a server. Accordingly, the method is distributed across multiple devices, and any latency in the streaming, synthesizing, or mixing process is hidden by the mixing step, such that the user listening to the performance does not perceive any latency. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram of an example network environment for audio synthesis for synchronous communication according to some embodiments described herein.

[0013] Figure 2 is a block diagram of an example computing device for audio synthesis for synchronous communication according to some embodiments described herein.

[0014] Figure 3A is a block diagram of an example architecture of a machine learning model according to some embodiments described herein.

[0015] Figure 3B is a block diagram of another example architecture of a machine learning model according to some embodiments described herein.

[0016] Figure 4 is an example user interface for guiding a user to change temporal aspects of a performance according to some embodiments described herein.

[0017] Figure 5 is an example flowchart showing data transfer between a client device and a server according to some embodiments described herein.

[0018] Figure 6 is an example flowchart of synthesizing an audio stream according to some embodiments described herein.

[0019] Figure 7 is an example flowchart of synthesizing an audio stream for synchronous communication according to some embodiments described herein.

[0020] Figure 8 is an example flowchart of using a server to synthesize audio for synchronous communication according to some embodiments described herein DETAILED DESCRIPTION

[0021] Network Environment 100

[0022] Figure 1 FIG. 1 shows a block diagram of an example environment 100 for audio synthesis in synchronous communication. In some embodiments, environment 100 includes a server 101 and client devices 115a...n coupled via a network 105. Users 125a...n may be associated with corresponding client devices 115a...n, respectively. In Figure 1 and the remaining figures, the letter following a reference numeral (e.g., "115a") indicates a reference to an element having that particular reference numeral. A reference numeral without a following letter in this document (e.g., "115") indicates a general reference to embodiments of an element having that reference numeral. In some embodiments, environment 100 may include Figure 1 other servers or devices not shown. For example, server 101 may be multiple servers 101.

[0023] Server 101 includes one or more servers, each server including a processor, a memory, and network communication hardware. In some embodiments, server 101 is a hardware server. Server 101 is communicatively coupled to network 105. In some embodiments, server 101 sends data to and receives data from client devices 115. Server 101 may include a metaverse engine 103 and a database 199.

[0024] In some embodiments, metaverse engine 103 includes code and routines for facilitating communication between two or more client devices 115 associated with users in a virtual metaverse, e.g., between friends at the same location in the metaverse, within the same metaverse experience, or within a metaverse application. In the metaverse, users interact with populations having different characteristics (e.g., different ages, regions, languages, etc.).

[0025] In some embodiments, metaverse engine 103 performs some or all of the following steps: generating a synthetic first audio stream predicting a future portion of a performance based on audio characteristics of a first audio stream, and mixing the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronously synthesizes the first audio stream and the second audio stream. Different embodiments regarding the step division will be discussed in detail below.

[0026] In some embodiments, the metaverse engine 103 is implemented using hardware, which includes a Central Processing Unit (CPU), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), any other type of processor, or any combination thereof. In some embodiments, the metaverse engine 103 is implemented using a combination of hardware and software.

[0027] The database 199 can be a non-transitory computer-readable memory (such as a random access memory), a cache, a drive (such as a hard disk drive), a flash drive, a database system, or another type of component or device capable of storing data. The database 199 can also include multiple storage components (such as multiple drives or multiple databases), which can be distributed across multiple computing devices (such as multiple server computers). The database 199 can store data associated with the metaverse engine 103, such as a dataset for training a machine learning model, reference audio, etc.

[0028] The client device 115 can be a computing device including a memory and a hardware processor. For example, the client device 115 can include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable gaming console, a portable music player, a reader device, or another electronic device capable of accessing the network 105.

[0029] The client device 115a includes the metaverse application 104a, and the client device 115n includes the metaverse application 104b. In some embodiments, the user 125a uses the metaverse application 104a on the client device 115a to generate a communication (such as a first audio stream), which is sent to the metaverse engine 103 on the server 101. The server 101 sends the communication to the metaverse application 104b on the client device 115b of the user 125n.

[0030] In some embodiments, the metaverse application 104 performs some or all of the following steps: generating a synthetic first audio stream that predicts a future portion of a performance based on audio features of a first audio stream, and mixing the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthetic first audio stream and the second audio stream. For example, the metaverse application 104a on the client device 115a may generate the synthetic first audio stream, and the metaverse application 104b on the client device 115n may mix the synthetic first audio stream with the second audio stream. In other embodiments, both the synthesizing and mixing steps may be performed on the server 101, and the metaverse application 104b on the client device 115n outputs the mixed synthetic first audio stream and the second audio stream through a speaker.

[0031] In the illustrated embodiment, entities of the environment 100 are communicatively coupled via the network 105. The network 105 may include a public network (such as the Internet), a private network (such as a Local Area Network (LAN) or a Wide Area Network (WAN)), a wired network (such as Ethernet), a wireless network (such as an 802.11 network, a Wi-Fi network or a Wireless LAN (WLAN)), a cellular network (such as a Long Term Evolution (LTE) network), a router, a hub, a switch, a server computer, or a combination thereof. Although Figure 1 one network 105 is shown coupled to the server 101 and the client device 115, in a practical application, one or more networks 105 may be coupled to these entities.

[0032] Example computing device 200

[0033] Figure 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In some embodiments, the computing device 200 is the server 101. In some embodiments, the computing device 200 is the client device 115.

[0034] In some embodiments, computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a microphone 241, a speaker 243, a display 245, and a storage device 247, each coupled by a bus 218. Depending on whether computing device 200 is a server 101 or a client device 115, some components of computing device 200 may not be present. For example, in an instance where computing device 200 is a server 101, the computing device may not include a microphone 241 and a speaker 243. In some embodiments, computing device 200 includes Figure 2 additional components not shown.

[0035] The processor 235 may be coupled to the bus 218 via a signal line 222, the memory 237 may be coupled to the bus 218 via a signal line 224, the I / O interface 239 may be coupled to the bus 218 via a signal line 226, the microphone 241 may be coupled to the bus 218 via a signal line 228, the speaker 243 may be coupled to the bus 218 via a signal line 230, the display 245 may be coupled to the bus 218 via a signal line 232, and the storage device 247 may be coupled to the bus 218 via a signal line 234.

[0036] The processor 235 includes an arithmetic logic unit, a microprocessor, a general-purpose controller, or other processor arrays for performing calculations and providing instructions to a display device. The processor 235 processes data and may include various computing architectures, including complex instruction set computer (CISC) architecture, reduced instruction set computer (RISC) architecture, or an architecture that implements a combination of instruction sets. Although Figure 2 a single processor 235 is shown, multiple processors 235 may be included. In different embodiments, the processor 235 may be a single-core processor or a multi-core processor. Other processors (such as a graphics processing unit), operating systems, sensors, displays, and / or physical configurations may be part of the computing device 200.

[0037] The memory 237 stores instructions and / or data that can be executed by the processor 235. The instructions can include code and / or routines for performing the techniques described herein. The memory 237 can be a Dynamic Random Access Memory (DRAM) device, a Static RAM, or other memory devices. In some embodiments, the memory 237 further includes non-volatile memory, such as a Static Random Access Memory (SRAM) device or a flash memory, or similar permanent storage devices and media, including hard disk drives, Compact Disc Read Only Memory (CD-ROM) devices, DVD-ROM devices, DVD-RAM devices, DVD-RW devices, flash memory devices, or other mass storage devices for long-term storage of information. The memory 237 includes code and routines that can be used to execute the metaverse engine 103, which is described in more detail below.

[0038] The I / O interface 239 can provide functions that enable the computing device 200 to connect to other systems and devices. The interface device can be part of the computing device 200 or separate from the computing device 200 and communicate with the computing device 200. For example, network communication devices, storage devices (such as the memory 237 and / or the storage device 247), and input / output devices can communicate through the I / O interface 239. In another example, the I / O interface 239 can receive data from the server 101 and transmit the data to the metaverse engine 103 and its components (such as the synthetic machine learning module 204). In some embodiments, the I / O interface 239 can be connected to interface devices, such as input devices (keyboard, pointing device, touch screen, microphone 241, sensors, etc.) and / or output devices (display device, speaker 243, monitor, etc.).

[0039] Examples of interface devices that can be connected to the I / O interface 239 include the display 245, which can be used to display content (such as images, videos, and / or user interfaces of the output applications described herein) and receive touch (or gesture) inputs from users. The display 245 can include any suitable display device, such as a Liquid Crystal Display (LCD), a Light Emitting Diode (LED), or a plasma display screen, a Cathode Ray Tube (CRT), a television, a monitor, a touch screen, a three-dimensional display, or other visual display devices.

[0040] The microphone 241 includes hardware for detecting the audio generated by the user 125. For example, the microphone 241 can detect the user 125 singing, the user 125 playing the violin, etc. The microphone 241 can transmit the audio to the metaverse engine 103 through the I / O interface 239.

[0041] The speaker 243 includes hardware for generating playback audio. For example, the speaker 243 receives instructions from the metaverse engine 103 to generate perceivable playback audio based on the digital combined audio stream generated by the metaverse engine 103. The speaker 233 converts the instructions into audio and generates a combined audio stream for the user.

[0042] The storage device 247 stores data related to the metaverse engine 103. For example, the storage device 247 can store the training data set for training the machine learning model, reference audio, etc. In an embodiment where the computing device 200 is the server 101, the storage device 247 is equivalent to Figure 1 the database 199 in

[0043] the example metaverse engine 103 or the example metaverse application 104

[0044] Figure 2 The computing device 200 that executes the example metaverse engine 103 or the example metaverse application 104 is shown. The example metaverse engine 103 or the example metaverse application 104 includes a performance recognition module 202, a synthesis machine learning module 204, a mixing module 206, a post-processing module 208, and a user interface module 210. Although these modules are exemplified as part of the same metaverse engine 103 or the same metaverse application 104, those of ordinary skill in the art should understand that these modules can be implemented by any computing device 200. For example, the performance recognition module 202 and the synthesis machine learning module 204 may be part of the client device 115, while the mixing module 206 may be part of the server 101 to reduce the computing requirements of the client device 115.

[0045] The performance recognition module 202 determines the performance identifier associated with the audio stream. In some embodiments, the performance recognition module 202 includes a set of instructions executable by the processor 235 to determine the performance identifier associated with the audio stream. In some embodiments, the performance recognition module 202 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0046] An audio stream can be a performance or rendition of known content (e.g., a re-recording of a known song, singing a known melody, reading text material, etc.). In this context, a performance identifier refers to the known content being performed in the audio stream. The performance identifier enables the identification and / or prediction of various attributes of the audio stream. For example, if the performance is based on a written musical score (including symbols for beats and sounds played by various instruments), the beats of the performance and the individual sounds played by each performer can be identified. In another example, if the performance includes singing or reading based on text, the upcoming words / phrases can be identified.

[0047] In some embodiments, after obtaining user permission, the performance recognition module 202 receives an audio stream. For example, when the performance recognition module 202 is part of the client device 115, the performance recognition module 202 receives the audio stream from the microphone 241 via the network 105 through the I / O interface 239. In another example, when the performance recognition module 202 is part of the server 101, the performance recognition module 202 receives the audio stream from the client device 115 via the I / O interface 239. The audio stream is part of a performance. The performance can be a song sung by multiple users, a speech, a recitation, music performed by an instrument, etc. In some embodiments, the user interface module 210 obtains permission to use the audio stream before the performance recognition module 202 receives the audio stream.

[0048] The performance recognition module 202 generates a fingerprint (i.e., an audio fingerprint) from at least a portion of the audio stream by generating a spectrogram of the audio stream that shows how the frequency changes over time and determining the frequency with the highest amplitude in the audio stream, and then generating a hash value of the spectrogram. Since the audio stream is a real-time performance by a person (rather than a pre-recorded performance), the performance recognition module 202 needs to consider pitch differences, including errors in temporal inconsistencies, timbre characteristics, etc. The performance recognition module 202 compares the fingerprint of the audio stream with a set of fingerprints of known songs to identify a match, which is related to the performance identifier of the audio stream. For example, the performance identifier can be "Happy Birthday Song", "Moonlight Sonata, Third Movement", etc.

[0049] The performance recognition module 202 can receive a reference audio associated with the performance identifier. For example, the performance recognition module 202 can retrieve the reference audio from the storage device 247.

[0050] After obtaining the user's permission, the performance recognition module 202 can determine the performance type of the audio stream. For example, the audio stream can include personal singing, personal speech, personal instrument playing, etc. In some embodiments, the performance recognition module 202 determines the type of performance based on the unique aspects of different instruments and human voices, such as by recognizing different frequencies, timings, pitches, speeds, etc. For example, the performance recognition module 202 can determine that in the metaverse, "Happy Birthday" is performed on a kazoo, and the third movement of "Moonlight Sonata" is performed on a piano. In some embodiments, the performance recognition module 202 can also determine the performance style. For example, whether the singer is performing a blues version or a rock version of a song, etc.

[0051] In some embodiments, the performance recognition module 202 obtains the permission of user 125, identifies one or more speakers in the audio stream, and associates the one or more speakers with speaker identifiers. The permission can include permission to use the audio stream for identification purposes, permission to store user information, etc. Information can be provided to user 125 to instruct that it can be stored (e.g., temporarily stored until the performance ends), and options to reject permission and select storage attributes (e.g., only locally stored, stored for X hours, etc.) are provided. After obtaining the permission, the performance recognition module 202 can, for example, determine that the audio stream associated with the first client device 115 includes a solo performance of "Happy Birthday" or a group performance of the song. The performance recognition module 202 can identify different people performing the song based on a mechanism similar to the above, where each performer is associated with a specific way of speaking, singing, etc. based on rhythm, pitch, frequency, etc.

[0052] The synthetic machine learning module 204 trains one or more machine learning models to synthesize the audio stream. In some embodiments, the synthetic machine learning module 204 includes a set of instructions executable by the processor 235 to train a machine learning model to synthesize the audio stream. In some embodiments, the synthetic machine learning module 204 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0053] The synthetic machine learning module 204 can use one or more (e.g., two) different training data sets. In some embodiments, the data set and the corresponding structure can depend on whether the synthetic machine learning module 204 is stored on the client device 115 or the server 101.

[0054] Example machine learning module 204 on the client device 115

[0055] In embodiments where the synthetic machine learning module 204 is stored on the client device 115, the mapping machine learning model 204 trains the machine learning model by using a training data set that manually annotates the audio stream, thereby implementing supervised learning.

[0056] The synthetic machine learning module 204 can be a neural network such as a deep neural network (DNN), which includes layers that identify progressively refined features and patterns in the audio stream, where the output of one layer serves as the input to subsequent layers. The output layer generates a synthetic audio stream. In some embodiments, the first layer (or first set of layers) outputs a mapping of the audio stream to the reference audio of the performance to determine the position of the audio stream relative to the reference audio, the second layer (or second set of layers) outputs a prediction of the future time offset of the audio stream, and the output layer synthesizes the audio stream based on the outputs of the previous layers. In some embodiments, the machine learning model can include long short-term memory (LSTM) nodes in a recurrent neural network (RNN), which is trained for sequential processing of the audio stream. The machine learning model can receive any audio stream as input and perform the same analysis to synthesize an output.

[0057] The synthetic audio stream takes into account network latency, input latency, and / or processing latency that cause audio latency. Network latency occurs when there is a delay in the reception of content from a transmitting node to a receiving node in the network due to an entity device using the network or the processing time at a node. Input latency occurs when there is a user response delay, the user is using a client device with limited capabilities, etc. Processing latency occurs when a delay is intentionally introduced for conditioning analysis or when there is insufficient processing resources. Since network latency is variable, the actual end-to-end latency of each audio stream can only be determined by the client device 115, where the audio stream is played back through the speaker 243 in the client device 115.

[0058] In some embodiments, the synthetic machine learning module 204 advantageously takes into account different types of input latency. For example, the performance may be Vivaldi's "The Four Seasons", the first audio stream is a violin performance with the third note delayed (determined using a performance identifier that supports identifying the notes and beats of "The Four Seasons"). The machine learning model can adjust the first audio stream based on the input latency by outputting a synthetic first audio stream that skips the sixteenth note to compensate for the missed third note. In another example, the synthetic machine learning module 204 can compensate for the processing latency introduced due to conditioning by adding a delay at the end of a speech sentence. Adding a delay at the end of a sentence can be advantageous because it has a lower impact on the listener's perception compared to a delay anywhere within the sentence.

[0059] Turning to Figure 3A , Figure 3A FIG. shows a block diagram of an example architecture of a machine learning model 300. The machine learning model 300 includes a bottleneck backbone 305, a submodel decoder 310, and an audio waveform generator 315. The bottleneck backbone 305 is a general bottleneck backbone for extracting features from the input waveform of the audio stream. The submodel decoder 310 is trained for a specific task: namely, determining the temporal correlation of the audio stream compared to a reference audio (e.g., determined based on a performance identifier).

[0060] The design of the bottleneck backbone 305 and the sub-model decoder 310 is similar to a machine learning model for speech-to-text synthesis. The synthesis machine learning module 204 can train the bottleneck backbone 305 using a large number of, for example, millions of pairs of {audio, text} samples to output bottleneck features. For example, the bottleneck backbone 305 can be trained in a manner similar to the DeepSpeech system, but instead of being trained to predict text, the bottleneck backbone 305 uses previous layers that store embeddings of the mapping of audio features to a reference audio.

[0061] The synthesis machine learning module 204 can train the sub-model decoder 310 by freezing the weights of the bottleneck backbone 305 after the bottleneck backbone 305 is trained, and providing the bottleneck features and audio features (such as pitch, phase, and rate) extracted from the audio stream input waveform to the sub-model decoder 310. In some embodiments, the sub-model decoder 310 also receives a speaker identifier for distinguishing different voices in the audio stream.

[0062] Once the bottleneck backbone 305 and the sub-model decoder 310 are trained, the bottleneck backbone 305 receives the audio stream generated by the client device 115 as input. In some embodiments, the bottleneck backbone 305 includes a fully connected (FC) layer followed by an LSTM layer, which is used to segment the audio stream into increasingly abstract representations of audio data.

[0063] The bottleneck backbone 305 receives the audio stream as an input waveform and Mel Frequency Cepstral Coefficient (MFCC) features sampled from overlapping samples of the audio stream. The MFCC features are from a cepstral representation of the audio stream and are used to represent audio stream information (such as the representation of timbre in the audio stream). The audio stream samples are parameterized by a time window. For example, a time window shorter than 100 milliseconds can be used, which has lower precision than longer time windows but is computationally efficient and also suitable for real-time execution on the client device 115. Although the computational processing power of the client device 115 is lower than that of the server 101, due to the network latency being less than the network latency generated by processing the audio stream on the server 101, the client device 115 can still generate accurate results with a short time window.

[0064] The bottleneck backbone 305 encodes the audio and outputs bottleneck features. The bottleneck features are transmitted to the sub-model decoder 310. In some embodiments, the bottleneck backbone 305 outputs bottleneck features for a subset of audio frames, for example, one out of every five frames in the audio stream. In some embodiments, the sub-model decoder 310 also receives audio-specific features related to the input waveform, such as the pitch, phase, rate of the audio stream, and the speaker identifier.

[0065] In some embodiments, the sub-model decoder 310 includes a fully connected (FC) layer followed by an LSTM layer. The LSTM layer is used to capture the temporal correlation of the audio stream. The sub-model decoder 310 outputs MFCC features, which are used as the input to the audio waveform generator 315.

[0066] The audio waveform generator 315 receives the MFCC features and generates an output waveform that renders the MFCC features as synthesized audio.

[0067] Throughout the process, the machine learning model 300 learns the temporal mapping between positions in the audio stream and synthesizes future frames of the audio stream. For example, if the performance is the song "Happy Birthday", the machine learning model 300 can map the audio stream to a reference audio by determining that the audio stream includes the first two lines of the song and, when the starting point of the audio stream is different from that of the reference audio, determining the difference between the audio stream and the reference audio. This is referred to as the temporal offset between the audio stream and the reference audio. The machine learning model 300 can also output the rate of the audio stream compared to the reference audio. For example, the user may sing the song "Happy Birthday" at a faster rate than the reference audio. The machine learning model 300 can predict the timing of the next frame of the audio stream based on the temporal offset, the rate of the audio stream, and the rate of the future temporal offset of the song based on the rate of the audio stream compared to the reference audio, and thus output synthesized audio that fuses these features.

[0068] In some embodiments, the machine learning model 300 can implement a multi-language bottleneck extractor instead of performing the above-mentioned bottleneck extraction. The multi-language bottleneck extractor is trained to distinguish multi-language senones. The output features are language-independent and are robust to variations caused by language, speaking style, speaking rate, etc. The machine learning model 300 is trained using supervised learning, and the output includes senone posteriors.

[0069] An example machine learning model on the client device 115

[0070] In an embodiment where the synthetic machine learning module 204 is stored in the server 101, the mapping machine learning model 204 can be trained using unsupervised learning with a training data set having unlabeled audio streams.

[0071] Go to Figure 3B , Figure 3B FIG. shows a block diagram of another example architecture of a machine learning model 350. The machine learning model 350 includes a vector quantization variational autoencoder (VQ-VAE) 355, a VQ-VAE codebook 360, a prior model 365, and a VQ-VAE decoder 370.

[0072] The synthetic machine learning module 204 trains the VQ-VAE 355 and the VQ-VAE codebook 360 using a training data set that includes the audio stream waveform. The VQ-VAE 355 and the VQ-VAE codebook 360 train the corresponding modules of the machine learning model 350 by comparing the output waveform with the input waveform. The synthetic machine learning module 204 trains the prior model 365 on the code vector output from the VQ-VAE codebook 360, and the code vector output depends on additional features such as audio features including one or more of pitch, phase, rate, and optionally, a speaker identifier.

[0073] The machine learning model 350 is trained to receive an audio stream parameterized by a time window. In some embodiments, the audio stream has a length of 300 - 500 milliseconds. Longer audio streams require more computation and a larger model size, but the results produced by longer audio streams are more accurate.

[0074] In some embodiments, the VQ-VAE 355 receives the input waveform of the audio stream and converts the input waveform into a latent representation. The VQ-VAE 355 employs an autoregressive network structure that includes a plurality of one-dimensional (1D) convolutional blocks. The VQ-VAE codebook 360 receives the latent representation as input, and the bottleneck quantizes the latent representation into discrete code vectors using a predefined codebook.

[0075] The prior model 365 receives the discrete code vectors from the VQ-VAE codebook 360 and audio features such as the pitch, phase, rate of the input waveform, and (optionally) one or more speaker identifiers of the audio stream. The audio features are calculated based on the input waveform, and the one or more speaker identifiers allow the network to model the codes specific to the user or music genre. The prior model 365 uses a transformation layer to perform temporal extrapolation on the code vectors and outputs the edited and resampled code vectors. The VQ-VAE decoder 370 is an autoregressive audio waveform generator. The VQ-VAE decoder 370 receives the edited and resampled code vectors and synthesizes an output waveform based on the edited and resampled code vectors.

[0076] The mixing module 206 mixes the audio streams from multiple client devices 115 to form a combined audio stream that synchronizes one or more synthetic audio streams with the real audio stream. In some embodiments, the mixing module 206 includes a set of instructions executable by the processor 235 to form the combined audio stream. In some embodiments, the mixing module 206 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0077] In some embodiments, the mixing module 206 receives the output of the synthetic machine learning module 204 and synchronizes the audio streams based on this output. For example, the mixing module 206 may receive the synthetic first audio stream and the second audio stream and generate a combined audio stream. Depending on whether the mixing is performed on the server 101 or on the client device 115, the mixing module 206 may introduce different amounts of delay in different audio streams to compensate for various delays and ensure audio stream synchronization. For example, if the mixing is performed on the server 101, the mixing module 206 needs to consider the delay of the recipient. Different factors of audio stream mixing will be discussed in more detail below.

[0078] The mixing module 206 synchronizes the synthetic first audio stream and the second audio stream to generate a combined audio stream. The time window length of the audio stream may vary depending on the device storing the synthetic machine learning module 204. For example, when the synthetic machine learning module 204 is stored in the client device 115, the time window is less than 100 milliseconds. For another example, when the synthetic machine learning module 204 is stored in the server 101, the time window is 300 - 500 milliseconds.

[0079] The mixing module 206 can achieve synchronization on the same or different devices as the device where the synthetic machine learning module 204 is stored. The devices may include the server 101, the client device 115a that sends the first audio stream, and the client device 115b that receives the first audio stream.

[0080] In the first example, the client device 115b performs the synthesis of the audio stream and the synchronization of the audio stream. The synthetic machine learning module 204 receives the global timestamp data packets from all other client devices 115 through the server 101 and synthesizes the audio streams of the other client devices 115. The mixing module 206 generates a local mix based on the real audio stream of the client device 115b and the synthetic audio of the other client devices 115.

[0081] In some embodiments, the first example is a preferred example. The first example has several advantages. The time required to predict the future part of the audio stream is reduced, thereby improving the quality of the audio stream and reducing the interaction delay. In addition, the client device 115b can recover from high latency. Although the user needs to use a client device with powerful computing capabilities for local synthesis and synchronization of audio, better hardware can provide a better experience. Finally, the advantage of using this architecture is that the synthesized audio remains on the client device 115b.

[0082] Some possible drawbacks of the first example are that for an input stream of n performers, the computational complexity is O(n). Additionally, since all processing is performed on the client device 115b and the processing volume is huge, these processes may cause the battery of the client device 115b to run out. These processes may be somewhat difficult, requiring different implementation codes to be written for different types of client devices 115, and can only run on a subset of client devices 115 with sufficient computing power.

[0083] In the second example, the server 101 performs the synthesis of the audio stream and the synchronization of the audio stream. Since all client streams are received and processed at the server, the advantages of the second example include a lower computational complexity of O(1). In this example, there is no need to perform calculations on the client side to generate the combined audio stream because the server 101 receives the independent audio streams of each client and provides the combined synthesized and synchronized audio stream to each client. The infrastructure is controlled because it is all part of the server 101. Additionally, even in the worst-case scenario, the prediction time is shorter than the prediction time for client-side processing because, even in the worst-case latency, the latency of the audio stream is only the latency between the client and the server. In client-side processing, the latency may be the latency between the client device 115a and the server 101 and then to the client device 115b.

[0084] In the third example, the client device 115a performs the synthesis of the audio stream, and the server 101 performs the synchronization of the audio stream. The client device 115a synthesizes the audio stream for the worst-case latency, and the server 101 synchronously mixes (n - 1) streams to the same playback time point by delaying each audio stream on demand. The advantages of the third example include: the input and output streams of each client device 115 are O(1), and the processing time at the server is O(1). Additionally, the better the hardware on the client device 115a, the better the sound effect of the client device 115a compared to other client devices 115. For example, better hardware enables better processing, which is particularly crucial in situations where it is necessary to quickly process the audio stream to maintain near real-time transmission of the audio stream.

[0085] In the fourth example, the client device 115a performs the synthesis of the audio stream, and the client device 115b performs the synchronization of the audio stream. This is applicable to peer-to-peer streaming without server intervention.

[0086] In instances where post - processing is not performed and the mixing module 206 does not perform synthesis on the client device 115b, the mixing module 206 instructs the I / O interface 239 to send the combined audio stream to the client device 115b, which can play the combined audio stream. In instances where post - processing is not performed and the mixing module 206 is located on the client device 115b, the mixing module 206 instructs the I / O interface 239 to provide the combined audio stream to the speaker 243 for playback.

[0087] The post - processing module 208 processes the combined audio stream. In some embodiments, the post - processing module 208 includes a set of instructions executable by the processor 235 to process the combined audio stream. In some embodiments, the post - processing module 208 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0088] In some embodiments, the post - processing module 208 can modify the combined audio stream by considering echo, background noise, etc., to make it consistent with the acoustic effects of the environment where the client device 115b is located. In some embodiments, the post - processing module 208 performs audio purification or noise suppression on the combined audio stream to avoid situations such as: for example, the first audio stream includes the honking of a car in the background, while the space where the second audio stream is located is completely silent, resulting in disharmony due to the sudden transition of the combined audio stream from having background noise to having no background noise.

[0089] The user interface module 210 generates a user interface. In some embodiments, the user interface module 210 includes a set of instructions executable by the processor 235 to generate the user interface. In some embodiments, the user interface module 210 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0090] The user interface module 210 generates a user interface for the user 125 associated with the client device 115. This user interface can be used to initiate audio communication with other users, participate in games or other experiences in the metaverse, send text to other users, initiate a video call with other users, etc.

[0091] In some embodiments, before a user joins the metaverse, the user interface module 214 generates a user interface that includes information on how user information is collected, stored, and analyzed. For example, the user interface requires the user to provide permission to use any information related to the user. The user is informed that the user can delete the user information, and the user can also choose which types of information to provide for different purposes. The use of the information complies with relevant regulations, and the data is stored securely. In a specific area and for a specific user category (e.g., based on age or other population characteristics), data collection is not performed. The data collection is temporary (i.e., the data is cleared regularly), and the data is not shared with third parties. Some data may be anonymized, aggregated across users, or modified so that a specific user identity cannot be determined.

[0092] In some embodiments, before sending an audio stream to the server 101 or other client device 115, the user interface obtains user permission. The user interface can include different levels of user permission granularity. For example, the user can specify that the synthesized audio can only be generated on the client device 115 and not on the server.

[0093] In some embodiments, the mixing module 206 determines that the time difference between the first audio stream and the second audio stream exceeds a threshold time difference, and sends an instruction to the user interface module 210 to generate a user interface. The user interface can guide the user 125 on how to change the singing rate so that the audio stream can be synchronized with other audio streams.

[0094] Go to Figure 4 , Figure 4 FIG. 400 shows an example user interface that guides a user to change the timing aspects of a performance. In this example, the user interface 400 includes user guidance for the performance and a movement indicator 405 that prompts the performer associated with the client device 115b to perform in a manner that reduces the time difference between the first audio stream and the second audio stream. For example, the movement indicator 405 can move slower than the user's performance, thereby suggesting that the user slow down the performance rate.

[0095] In some embodiments, the user interface module 210 shows the synchronization of the combined audio stream with the actions of the performer graphically displayed. For example, if the combined audio stream is a speech performed by the user graphically displayed as a virtual character, the user interface module 210 can synchronize the combined audio stream with the lip movements, actions, etc. of the virtual character.

[0096] Example method

[0097] Figure 5For example flowchart 500, it shows the data transmission between client device 115 and server 101 according to some embodiments described herein. Flowchart 500 includes a first client device 510, a server 515, and a second client device 520. The thick lines represent the network data transmission between the three devices, and the thin lines represent the data transmission within the first client device 510.

[0098] The first client device 510 receives an audio stream from the microphone 505. The audio stream is synthesized at the first client device 510, the server 515, or the second client device 520. The synthesized audio stream is mixed with one or more other audio streams at the server 515 or the second client device 520 to form a combined audio stream. The combined audio stream is sent to the speaker 525 for playback at the first client device 510.

[0099] In addition, the server 515 also receives data. For example, the server 515 receives a first stream from the first client device 510 and a second stream from the second client device 520, where the server 515 synthesizes the first audio stream and mixes the first audio stream with the second audio stream.

[0100] By implementing the metaverse engine 103, the first audio stream of the violin and the second audio stream of the cello on the first client device 510 can be heard simultaneously by the associated users performing music as if they were in the same physical space.

[0101] Figure 6 It is an example flowchart 600 for synthesizing an audio stream for synchronous communication. Task detection 605 is performed on the delayed audio stream to identify the performance identifier and performance type. For example, it is identified whether the delayed audio stream includes an instrument, voice, singing, etc., as well as the type of the song and the style of the performance. Based on this identification, a reference audio is output.

[0102] The delayed audio stream is also received by the phase shift analyzer 610 and the rate analyzer 615. The phase shift analyzer 610 identifies the time offset of the delayed audio stream compared to the reference audio, and the rate analyzer 615 identifies the rate of the delayed audio stream. The time offset and rate of the delayed audio stream are received by the sampler 620, and the sampler 620 maps the audio stream to the reference audio. This mapping is received by the deep neural network 635 (or other suitable model), and the deep neural network 635 extrapolates the audio stream into the future to complete the synthesis of the audio stream.

[0103] Figure 7 It is an example flowchart for synthesizing an audio stream for synchronous communication using the second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.

[0104] Method 700 may begin at block 702. At block 702, a first audio stream of a performance associated with the first client device 115 is received. After block 702 may be block 704.

[0105] At block 704, during a time window of the performance, where the time window is less than the total time of the performance: generate a combined audio stream that synthesizes the first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronously synthesizes the first audio stream and the second audio stream, and mix the synthesized first audio stream and the second audio stream associated with the second client device to form a combined audio stream that synchronously synthesizes the first audio stream and the second audio stream, where the time window is advanced, and the generation and synchronization are repeated until the performance is complete.

[0106] Figure 7 FIG. is an example flowchart of synthesizing an audio stream for synchronous communication using a second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.

[0107] Figure 8 FIG. is another example flowchart of synthesizing audio for synchronous communication using a server. In this example, the metaverse engine 103 is stored on the server 101.

[0108] Method 800 may begin at block 802. At block 802, a first audio stream of a performance associated with the first client device 115 is received. After block 802 may be block 804.

[0109] At block 804, within a time window of the performance, where the time window is less than the total time of the performance: generate a combined audio stream that synthesizes the first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronously synthesizes the first audio stream and the second audio stream, and mix the synthesized first audio stream and the second audio stream associated with the second client device to form a combined audio stream that synchronously synthesizes the first audio stream and the second audio stream, where the time window is advanced, and the generation and synchronization are repeated until the performance is complete. In some embodiments, mixing the synthesized first audio stream includes introducing a delay into the combined audio stream to handle a delay that occurs during sending the combined audio to the second client device. After block 804 may be block 806.

[0110] At block 806, send the combined audio stream to the second client device 115.

[0111] The various embodiments described herein include obtaining data from various sensors in a physical environment, analyzing such data, generating recommendations, and providing a user interface. Data collection is performed only with the permission of a specific user and in compliance with applicable regulations. The storage of data complies with applicable regulations, including anonymizing the data or otherwise modifying the data to protect user privacy. Clear information about data collection, storage, and use is provided to the user, and options are provided to select the types of data that can be collected, stored, and used. Additionally, the user controls the devices on which data can be stored (e.g., only user devices, client and server devices, etc.) and the devices that perform data analysis (e.g., only user devices, client and server devices, etc.). The data is used for the specific purposes described herein. No data is shared with third parties without the explicit permission of the user.

[0112] The methods, blocks, and / or operations herein can be performed in an order different from that shown or described, and / or, where appropriate, simultaneously (partially or fully) with other blocks or operations. Some blocks or operations can be performed on a portion of the data and later, for example, on another portion of the data again. Not all of the described blocks and operations are required to be performed in various embodiments. In some embodiments, the blocks and operations can be performed multiple times in a different order and / or at different times in the method.

[0113] In the foregoing description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form in order to avoid obscuring the description. For example, the embodiments may be described primarily with reference to a user interface and specific hardware. However, the embodiments may be applied to any type of computing device that can receive data and commands, as well as any external device that provides services.

[0114] References in the specification to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example can be included in at least one implementation of the specification. The phrase "in some embodiments" appearing in different places in the specification does not necessarily refer to the same embodiment.

[0115] Some of the portions detailed above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, considered to be a self-consistent sequence of steps leading to a desired result. These steps are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0116] However, it should be noted that all of these terms and like terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless explicitly stated otherwise in the following discussion, throughout this description, discussions using terms such as 'processing', 'computing', 'deriving', 'displaying', etc., refer to the actions and processes of a computer system or similar electronic computing device. These actions and processes involve the manipulation and transformation of data represented in physical (electronic) quantities within the registers and memories of the computer system and turning them into other data similarly represented in physical quantities, which are stored in the memories or registers of the computer system or other information storage, transmission, or display devices.

[0117] Embodiments of this specification may also relate to a processor for performing one or more steps of the above methods. The processor may be a special purpose processor selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, which includes but is not limited to any type of disk, including optical disks, ROM, CD-ROM, magnetic disks, RAM, EPROM, EEPROM, magnetic or optical cards, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to the computer system bus.

[0118] The specification may take the form of some embodiments that are entirely hardware, some embodiments that are entirely software, or some embodiments that contain both hardware and software elements. In some embodiments, the specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0119] In addition, the present description may take the form of a computer program product, which can be accessed from a computer-usable medium or a computer-readable medium providing program code, the program code being for use by a computer or any instruction execution system, or associated with a computer or any instruction execution system. For the purposes of the present description, a computer-usable or computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or associated with an instruction execution system, apparatus, or device.

[0120] A data processing system suitable for storing or executing program code will include at least one processor, which is directly or indirectly connected to memory elements through a system bus. The memory elements can include local memory, mass storage, and cache memory used during actual execution of the program code, the cache memory at least providing some temporary storage of the program code to reduce the number of times the code must be retrieved from the mass storage during execution.

Claims

1. A computer-implemented method, comprising: receiving a first audio stream of a performance associated with a first client device; during a time window of the performance, wherein the time window is less than the total time of the performance: generating a synthetic first audio stream that predicts a future portion of the performance based on audio characteristics of the first audio stream; and mixing the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthetic first audio stream and the second audio stream; wherein the time window is advanced and the generating and the mixing are repeated until the performance is complete.

2. The method of claim 1, further comprising: responsive to receiving the first audio stream, determining a performance identifier of the performance associated with the first audio stream; and receiving reference audio based on the performance identifier.

3. The method of claim 2, wherein: generating the synthetic first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthetic first audio stream is further based on the time offset.

4. The method of claim 2, wherein: generating the synthetic first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio; and generating the synthetic first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.

5. The method of claim 1, wherein, the audio characteristics of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

6. The method of claim 1, wherein, the audio characteristics of the first audio stream include one or more speaker identifiers detected in the first audio stream.

7. The method of claim 1, further comprising: determining that a time difference between the first audio stream and the second audio stream exceeds a threshold time difference; and generating graphical data for displaying a user interface that includes user guidance for the performance and a movement indicator that prompts a performer associated with the second client device to perform in a manner that reduces the time difference between the first audio stream and the second audio stream.

8. The method of claim 1, further comprising modifying the combined audio stream to be consistent with the acoustic effects of the environment where the second client device is located.

9. The method of claim 1, wherein, generating the synthetic first audio stream includes: identifying that a portion of the performance is skipped in the first audio stream; and synthesizing the first audio stream to correct the skipped portion of the performance.

10. The method of claim 1, further comprising synchronizing the combined audio stream to match the graphically displayed actions of the performer.

11. A device, comprising: a processor; and A memory coupled to the processor, instructions being stored on the memory, which when executed by the processor, cause the processor to perform operations including the following: Receive a first audio stream of a performance associated with a first client device; And During a time window of the performance, wherein the time window is less than the total time of the performance: Based on the audio characteristics of the first audio stream, generate a synthetic first audio stream predicting a future portion of the performance; and Mix the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream synchronizing the synthetic first audio stream and the second audio stream; Wherein the time window is shifted forward, and the generating and the mixing are repeated until the performance is completed.

12. The apparatus according to claim 11, Characterized in that: In response to receiving the first audio stream, determine a performance identifier of the performance associated with the first audio stream; and Receive a reference audio based on the performance identifier.

13. The apparatus according to claim 12, Characterized in that, Generating the synthetic first audio stream includes: Generating the synthetic first audio stream includes determining a time offset between the first audio stream and the reference audio; and The time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthetic first audio stream is also based on the time offset.

14. The apparatus according to claim 12, Characterized in that: Generating the synthetic first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and Generating the synthetic first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio.

15. The apparatus according to claim 11, Characterized in that, The audio characteristics of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

16. A non-transitory computer-readable medium, instructions being stored on the non-transitory computer-readable medium, which when executed by one or more computers, cause the one or more computers to perform operations, the operations Include: Receive a first audio stream of a performance associated with a first client device; And During a time window of the performance, wherein the time window is less than the total time of the performance: Based on the audio characteristics of the first audio stream, generate a synthetic first audio stream predicting a future portion of the performance; and Mix the synthetic first audio stream and a second audio stream associated with a second client device to form a combined audio stream synchronizing the synthetic first audio stream and the second audio stream; Wherein the time window is shifted forward, and the generating and the mixing are repeated until the performance is completed.

17. The computer-readable medium according to claim 16, Characterized in that: In response to receiving the first audio stream, determine a performance identifier of the performance associated with the first audio stream; and Receive a reference audio based on the performance identifier.

18. The computer-readable medium according to claim 17, Characterized in that: Generating the synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio; and the time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthesized first audio stream is also based on the time offset.

19. The computer-readable medium according to claim 17, wherein: generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio; and generating the synthesized first audio stream is also based on the rate of the first audio stream compared to the rate of the reference audio.

20. The computer-readable medium according to claim 16, wherein, the audio features of the first audio stream are selected from the group consisting of pitch, rate, phase, or any combination thereof.

Citation Information

Patent Citations

  • Method, system, and program product for measuring audio video synchronization independent of speaker characteristics

    CA2565758A1

  • Accompaniment method for actively following music signals and related equipment

    CN112669798A

  • System and device for speech analytic synthesis

    JP1989302299A

  • Synchronizing method and system

    US20070188657A1

  • Systems and methods for processing meeting information obtained from multiple sources

    US20190341068A1