Audio synthesis for synchronous communication
A method generates a synthesized audio stream to predict and synchronize performances across client devices, addressing latency issues and ensuring synchronized audio playback.
Patent Information
- Application Number
- JP2025519615
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-04
- Filing Date
- 2023-10-02
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-10-02
AI Technical Summary
Multiple users in different physical locations connected by a computer network cannot play music or speak in sync due to network, input, and processing latencies, causing performance desynchronization.
A computer-implemented method generates a synthesized audio stream that predicts the future of a performance based on audio features, mixing it with other streams to synchronize performances across client devices, accounting for latency and environment.
Synchronizes audio streams across multiple users, hiding latency and ensuring perceived real-time performance alignment.
Smart Images

Figure 2025535711000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is an international application and claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. patent application Ser. No. 17 / 959,736, filed October 4, 2022, entitled SYNTHESIZING AUDIO FOR SYNCHRONOUS COMMUNICATION, the entire contents of which are incorporated herein by reference. [Background technology]
[0002] It is impossible for multiple people in different physical locations but connected by a computer network to play music, sing, or speak in sync. This is because the performance of person P0 observed by person P1 is always in the past due to transmission delays in the computer network. If each of the n-1 performers P1, P2, P3, ...P(n-1) were delayed by exactly the right amount of time relative to performer P0, P0 would observe everyone else in sync with each other, but would be unable to synchronize himself with the other performers.
[0003] The delay may be due to one or more of at least three types of latency: network latency, input latency, and processing latency. Network latency occurs when there is a shortage in network transmission time due to the physical equipment used in the network or a delay in processing time at a node. Input latency occurs when a user delays a response, when a user is using a client device with limited capabilities, etc. Processing latency may occur when a delay is intentionally introduced to perform moderation analysis.
[0004] The background discussion provided herein is for the purpose of presenting a context for the present disclosure. To the extent described in this background section, the work of the presently named inventors, as well as aspects of the body of the specification that may not qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention [Means for solving the problem]
[0005]
[0003] According to one aspect, a computer-implemented method includes receiving a first audio stream of a performance associated with a first client device, generating a synthesized first audio stream that predicts the future of the performance based on audio features of the first audio stream during a time window of the performance that is shorter than a total duration of the performance, and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthesized first audio stream and a second audio stream associated with a second client device, wherein the time window is advanced and the generating and mixing steps are repeated until the performance is completed.
[0006] In some embodiments, the method further includes, in response to receiving the first audio stream, determining a performance identifier of a performance associated with the first audio stream and receiving reference audio based on the performance identifier. In some embodiments, generating the synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, where the time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthesized first audio stream is further based on the time offset. In some embodiments, generating the synthesized first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio, and generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from the group of pitch, rate, phase, or a combination thereof. In some embodiments, the audio features of the first audio stream include one or more speaker identifiers detected in the first audio stream. In some embodiments, the method further includes determining that a time difference between the first audio stream and the second audio stream exceeds a threshold time difference, and generating graphical data for displaying a user interface including user guidance regarding the performance and a movement indicator that prompts a performer associated with the second client device to perform in a manner that reduces the time difference between the first audio stream and the second audio stream. In some embodiments, the method further includes modifying the combined audio stream to match the acoustics of an environment in which the second client device is located.In some embodiments, generating the combined first audio stream includes identifying in the first audio stream that a portion of the performance has been skipped, and synthesizing the first audio stream to correct the skipped portion of the performance. In some embodiments, the method further includes synchronizing the combined audio stream to match actions of the performer as graphically displayed.
[0007] In some embodiments, a device includes a processor and a memory, coupled to the processor, having instructions stored thereon, which, when executed by the processor, cause the processor to perform operations including receiving a first audio stream of a performance associated with a first client device; generating a synthesized first audio stream that predicts the future of the performance based on audio features of the first audio stream during a time window of the performance that is shorter than the total time of the performance; and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device, wherein the time window is advanced and the generating and mixing are repeated until the performance is completed.
[0008] In some embodiments, in response to receiving the first audio stream, a performance identifier of a performance associated with the first audio stream is determined, and reference audio is received based on the performance identifier. In some embodiments, generating the synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, where the time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthesized first audio stream is further based on the time offset. In some embodiments, generating the synthesized first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio, and generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio feature of the first audio stream is selected from the group of pitch, rate, phase, or a combination thereof.
[0009] In some embodiments, a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: receiving a first audio stream of a performance associated with a first client device; generating a synthesized first audio stream that predicts the future of the performance based on audio features of the first audio stream during a time window of the performance that is shorter than a total time of the performance; and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device, wherein the time window is advanced and the generating and mixing are repeated until the performance is completed.
[0010] In some embodiments, in response to receiving the first audio stream, a performance identifier of a performance associated with the first audio stream is determined, and reference audio is received based on the performance identifier. In some embodiments, generating the synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, where the time offset occurs when the first audio stream has a different starting point from the reference audio, and generating the synthesized first audio stream is further based on the time offset. In some embodiments, generating the synthesized first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio, and generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio feature of the first audio stream is selected from the group of pitch, rate, phase, or a combination thereof.
[0011] This application describes a metaverse engine and / or metaverse application that advantageously generates a synthesized first audio stream that predicts the future of a performance based on audio characteristics of the first audio stream and mixes the synthesized first audio stream with a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device so that the users who created the audio streams are perceived as singing or speaking in sync. The generation of the synthesized first audio stream and the mixing of the synthesized first audio stream with the second audio stream may be performed on different devices, including a combination of a first audio device, a second audio device, and a server. As a result, the method is distributed across multiple devices, and any delays in the streaming, synthesis, or mixing processes are hidden by the mixing step so that users listening to the performance do not perceive any latency. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment for synthesizing audio for synchronized communication, according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an exemplary computing device for synthesizing audio for synchronized communication, according to some embodiments described herein. [Figure 3A] FIG. 1 is a block diagram of an example architecture of a machine learning model, according to some embodiments described herein. [Figure 3B] FIG. 1 is a block diagram of another example architecture of a machine learning model, according to some embodiments described herein. [Figure 4]10A-10C are diagrams of exemplary user interfaces that guide a user to change timing aspects of a performance, according to some embodiments described herein. [Figure 5] FIG. 2 is an exemplary flow diagram illustrating the transmission of data between a client device and a server, according to some embodiments described herein. [Figure 6] FIG. 2 is an exemplary flow diagram for synthesizing audio streams according to some embodiments described herein. [Figure 7] FIG. 1 illustrates an exemplary flow diagram for synthesizing audio streams for synchronous communication, according to some embodiments described herein. [Figure 8] FIG. 1 illustrates an exemplary flow diagram for synthesizing audio for synchronous communication using a server, according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Network Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100 for synthesizing audio for synchronized communication. In some embodiments, the environment 100 includes a server 101 and client devices 115a...n coupled via a network 105. Users 125a...n may be associated with respective client devices 115a...n. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," represents a reference to the element with that specific reference number. A reference number in text without a following letter, e.g., "115," represents a general reference to the element with that reference number. In some embodiments, the environment 100 may include other servers or devices not shown in FIG. 1. For example, the server 101 may be multiple servers 101.
[0014] Server 101 includes one or more servers, each including a processor, memory, and network communication hardware. In some embodiments, server 101 is a hardware server. Server 101 is communicatively coupled to network 105. In some embodiments, server 101 transmits data to and receives data from client devices 115. Server 101 may include metaverse engine 103 and database 199.
[0015] In some embodiments, the metaverse engine 103 includes code and routines operable to facilitate communication between client devices 115 associated with two or more users within a virtual metaverse, e.g., communication at the same location within the metaverse, communication within the same metaverse experience, or communication between friends within a metaverse application. Users interact within the metaverse across various demographics (e.g., different ages, regions, languages, etc.).
[0016] In some embodiments, the metaverse engine 103 performs some or all of the steps of generating a synthesized first audio stream that predicts future performance based on audio features of the first audio stream, and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with the second client device. Different embodiments of how the steps are divided are discussed in more detail below.
[0017] In some embodiments, the metaverse engine 103 is implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the metaverse engine 103 is implemented using a combination of hardware and software.
[0018] Database 199 may be non-transitory computer-readable memory (e.g., random access memory), a cache, a drive (e.g., a hard drive), a flash drive, a database system, or another type of component or device capable of storing data. Database 199 may also include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers). Database 199 may store data associated with metaverse engine 103, such as training datasets for trained machine learning models, reference audio, etc.
[0019] The client device 115 may be a computing device that includes a memory and a hardware processor. For example, the client device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.
[0020] Client device 115a includes metaverse application 104a, and client device 115n includes metaverse application 104b. In some embodiments, user 125a uses metaverse application 104a on client device 115a to generate a communication, such as a first audio stream, which is sent to metaverse engine 103 on server 101. Server 101 sends the communication to metaverse application 104b on client device 115b for user 125n.
[0021] In some embodiments, the metaverse application 104 performs some or all of the steps of generating a synthesized first audio stream that predicts future performance based on audio characteristics of the first audio stream and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device. For example, the metaverse application 104a on the client device 115a may generate the synthesized first audio stream, and the metaverse application 104b on the client device 115n may mix the synthesized first audio stream and the second audio stream. In other embodiments, both the synthesis and the mixing may be performed on the server 101, and the metaverse application 104b on the client device 115n outputs the mixed synthesized first audio stream and second audio stream through a speaker.
[0022] In the illustrated embodiment, the entities of environment 100 are communicatively coupled via network 105. Network 105 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, a Wi-Fi network, or a wireless LAN (WLAN)), a cellular network (e.g., a long-term evolution (LTE) network), a router, a hub, a switch, a server computer, or a combination thereof. While FIG. 1 shows one network 105 coupled to server 101 and client device 115, in practice, one or more networks 105 may be coupled to these entities.
[0023] 200 Examples of Computing Devices 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In some embodiments, computing device 200 is server 101. In some embodiments, computing device 200 is client device 115.
[0024] In some embodiments, computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a microphone 241, a speaker 243, a display 245, and a storage device 247, each coupled via a bus 218. Depending on whether computing device 200 is a server 101 or a client device 115, some components of computing device 200 may not be present. For example, if computing device 200 is a server 101, the computing device may not include microphone 241 and speaker 243. In some embodiments, computing device 200 includes additional components not shown in FIG. 2 .
[0025] Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, microphone 241 may be coupled to bus 218 via signal line 228, speaker 243 may be coupled to bus 218 via signal line 230, display 245 may be coupled to bus 218 via signal line 232, and storage device 247 may be coupled to bus 218 via signal line 234.
[0026] Processor 235 includes an arithmetic logic unit, a microprocessor, a general-purpose controller, or some other processor array for performing calculations and providing instructions to a display device. Processor 235 processes data and may include various computing architectures, including a complex instruction set computer (CISC) architecture, a reduced instruction set computer (RISC) architecture, or an architecture implementing a combination of instruction sets. While FIG. 2 shows a single processor 235, multiple processors 235 may be included. In different embodiments, processor 235 may be a single-core processor or a multi-core processor. Other processors (e.g., graphics processing units), operating systems, sensors, displays, and / or physical configurations may be part of computing device 200.
[0027] Memory 237 stores instructions and / or data that may be executed by processor 235. The instructions may include code and / or routines for performing the techniques described herein. Memory 237 may be a dynamic random access memory (DRAM) device, static RAM, or some other memory device. In some embodiments, memory 237 also includes non-volatile memory, such as a static random access memory (SRAM) device or flash memory, or similar persistent storage devices and media, including a hard disk drive, a compact disc read-only memory (CD-ROM) device, a DVD-ROM device, a DVD-RAM device, a DVD-RW device, a flash memory device, or some other mass storage device for more permanently storing information. Memory 237 includes code and routines operable to execute metaverse engine 103, described in more detail below.
[0028] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and communicate with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 247), and input / output devices may communicate through I / O interface 239. In another example, I / O interface 239 may receive data from server 101 and provide the data to components of metaverse engine 103, such as metaverse engine 103 and synthetic machine learning module 204. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone 241, sensor, etc.) and / or output devices (display device, speaker 243, monitor, etc.).
[0029] Some examples of interfaced devices that may be connected to I / O interface 239 may include display 245, which may be used to display content, e.g., images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. Display 245 may include any suitable display device, such as a liquid crystal display (LCD), light emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device.
[0030] Microphone 241 includes hardware for detecting audio performed by user 125. For example, microphone 241 may detect user 125 singing, user 125 playing the violin, etc. Microphone 241 may transmit the audio to metaverse engine 103 via I / O interface 239.
[0031] Speaker 243 includes hardware for generating audio for playback. For example, speaker 243 receives instructions from metaverse engine 103 to generate perceptual audio for playback from the digital combined audio stream generated by metaverse engine 103. Speaker 243 converts the instructions into speech and generates the combined audio stream for the user.
[0032] Storage device 247 stores data related to metaverse engine 103. For example, storage device 247 may store training datasets for trained machine learning models, reference audio, etc. In embodiments in which computing device 200 is server 101, storage device 247 is the same as database 199 in FIG.
[0033] Exemplary Metaverse Engine 103 or Metaverse Application 104 2 illustrates a computing device 200 executing an example metaverse engine 103 or metaverse application 104 that includes a performance recognition module 202, a synthetic machine learning module 204, a blending module 206, a post-processing module 208, and a user interface module 210. While the modules are shown as being part of the same metaverse engine 103 or metaverse application 104, those skilled in the art will recognize that the modules may be implemented by any computing device 200. For example, the performance recognition module 202 and the synthetic machine learning module 204 may be part of the client device 115, while the blending module 206 may be part of the server 101 to reduce the computational requirements of the client device 115.
[0034] The performance recognition module 202 determines a performance identifier associated with the audio stream. In some embodiments, the performance recognition module 202 includes a set of instructions executable by the processor 235 to determine a performance identifier associated with the audio stream. In some embodiments, the performance recognition module 202 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0035] The audio stream may be a performance or rendition of known content, such as a re-recording of a known song, singing a known melody, reading text material, etc. A performance identifier, in this context, refers to the known content being performed within the audio stream. The performance identifier may enable identification and / or prediction of various attributes of the audio stream. For example, if the performance is based on written musical notation (e.g., including tempo and note notation to be played by various instruments), the tempo of the performance and identification of the individual notes played by each performer. In another example, if the performance includes singing or reading from a text, the upcoming word / phrase may be identified.
[0036] In some embodiments, the performance recognition module 202 receives the audio stream after obtaining user permission. For example, if the performance recognition module 202 is part of the client device 115, the performance recognition module 202 receives the audio stream from the microphone 241 via the I / O interface 239 over the network 105. In another example, if the performance recognition module 202 is part of the server 101, the performance recognition module 202 receives the audio stream from the client device 115 via the I / O interface 239. The audio stream is part of a performance. The performance may be a song sung by multiple users, a speech, a chant, music played by an instrument, etc. In some embodiments, the user interface module 210 obtains permission to use the audio stream before the performance recognition module 202 receives the audio stream.
[0037] The performance recognition module 202 generates a fingerprint (i.e., an audio fingerprint) from at least a portion of the audio stream by generating a spectrogram of the audio stream, including frequency as a function of time, determining which frequencies in the audio stream have the highest amplitude, and then generating a hash of the spectrogram. Because the audio stream is being performed live by a person (as opposed to a pre-recorded performance), the performance recognition module 202 considers errors, including key differences, timing inconsistencies, vocal quality, etc. The performance recognition module 202 compares the fingerprint of the audio stream to a set of fingerprints for known songs to identify a match. The match is associated with a performance identifier for the audio stream. For example, the performance identifier may be "Happy Birthday," "Moonlight Sonata, Third Movement," etc.
[0038] The performance recognition module 202 may receive reference audio associated with the performance identifier. For example, the performance recognition module 202 may obtain the reference audio from the storage device 247.
[0039] After obtaining user permission, the performance recognition module 202 may determine a type of performance for the audio stream. For example, the audio stream may include someone singing, giving a speech, playing an instrument, etc. In some embodiments, the performance recognition module 202 determines the type of performance based on unique aspects of various instruments and the human voice, such as by identifying various frequencies, timing, pitch, speed, etc. associated with the instruments or human voice. For example, the performance recognition module 202 may determine that "Happy Birthday" is being played on a kazoo and "Moonlight Sonata, Third Movement" is being played on a piano in the metaverse. In some embodiments, the performance recognition module 202 additionally determines a style of performance, such as whether a singer is performing a blues version of a song or a rock version of a song.
[0040] In some embodiments, the performance recognition module 202 obtains permission from the user 125 to identify one or more speakers in the audio stream and associate the one or more speakers with a speaker identifier. Permission may include permission to use the audio stream for identification purposes, permission to store information about the user, etc. The user 125 is provided with guidance that this information may be stored (e.g., temporarily, until the performance ends), and is provided with the option to deny permission, select attributes of the storage (e.g., local only, X hours, etc.). After obtaining permission, the performance recognition module 202 may determine, for example, that the audio stream associated with the first client device 115 includes one person singing "Happy Birthday," or multiple people performing the song. The performance recognition module 202 may identify the various people performing the song based on mechanisms similar to those described above, in which each performer is associated with a particular speaking style, singing style, etc., based on rhythm, tone, frequency, etc.
[0041] The synthetic machine learning module 204 trains a machine learning model (or models) for synthesizing the audio stream. In some embodiments, the synthetic machine learning module 204 includes a set of instructions executable by the processor 235 to train a machine learning model for synthesizing the audio stream. In some embodiments, the synthetic machine learning module 204 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0042] The synthetic machine learning module 204 may use one or more (e.g., two) different training datasets. In some implementations, the datasets and corresponding structures may depend on whether the synthetic machine learning module 204 is stored on the client device 115 or the server 101.
[0043] Exemplary Machine Learning Module 204 on Client Device 115 In embodiments in which the synthetic machine learning module 204 is stored on the client device 115, the mapping machine learning model 204 implements supervised learning by training the machine learning model using a training dataset having manually labeled audio streams.
[0044] The synthetic machine learning module 204 may be a neural network, such as a deep neural network (DNN), that includes layers that identify more detailed features and patterns within the audio stream, with the output of one layer serving as input to subsequent layers. The output layer generates a synthesized audio stream. In some embodiments, a first layer (or a first set of layers) outputs a mapping of the audio stream to a reference audio of the performance to determine the position of the audio stream relative to the reference audio, and a second layer (or a second set of layers) outputs a prediction of the future time offset of the audio stream, and the output layer synthesizes the audio stream based on the output of the previous layer. In some embodiments, the machine learning model may include a long short-term memory (LSTM) within a recurrent neural network (RNN) trained for sequential processing of the audio stream. The machine learning model can receive any audio stream as input and perform the same analysis to synthesize an output.
[0045] The synthesized audio stream takes into account network latency, input latency, and / or processing latency, which cause audio delays. Network latency occurs when there is a delay in receiving content from a sending node to a receiving node in a network due to the physical equipment used in the network or due to processing time delays at the nodes. Input latency occurs when a user delays their response, when using a client device with limited capabilities, etc. Processing latency can occur when delays are intentionally introduced to perform moderation analysis or due to insufficient processing resources. Because network latency is variable, the true end-to-end latency for each audio stream can only be determined by the client device 115 where the audio stream is played through the speaker 243.
[0046] In some embodiments, the synthetic machine learning module 204 advantageously accounts for different types of input latency. For example, the performance is Vivaldi's "The Four Seasons," and the first audio stream is a violin with a delayed third note (determined using a performance identifier that allows for identification of the notes and tempo of the Four Seasons). The machine learning model may adjust the first audio stream based on the input latency by outputting a synthesized first audio stream that accounts for the missing third note by skipping the 16th note. In another example, the synthetic machine learning module 204 may account for processing latency introduced for moderation by adding a delay to the end of a speech sentence. Adding a delay to the end of a sentence may be advantageous because the perception of the delay has less impact on the listener than a delay at any point in the sentence.
[0047] 3A, a block diagram of an exemplary architecture of a machine learning model 300 is shown. The machine learning model 300 includes a bottleneck trunk 305, a sub-model decoder 310, and an audio waveform generator 315. The bottleneck trunk 305 is a general-purpose bottleneck trunk that extracts features from an input waveform of an audio stream. The sub-model decoder 310 is trained for a specific task, namely, determining temporal correlations within an audio stream compared to reference audio (e.g., determined based on a performance identifier).
[0048] The bottleneck trunk 305 and sub-model decoder 310 are designed similarly to machine learning models that implement speech-to-text synthesis. The machine learning by synthesis module 204 may train the bottleneck trunk 305 using a large number, e.g., millions of pairs of (audio, text) samples to output bottleneck features. For example, the bottleneck trunk 305 may be trained similarly to a DeepSpeech system, but instead of being trained to predict text, the bottleneck trunk 305 uses a previous layer that holds embeddings to map audio features to reference audio.
[0049] The synthetic machine learning module 204 may train the sub-model decoder 310 by fixing the weights of the bottleneck trunk 305 after the bottleneck trunk 305 is trained, and may provide the bottleneck features to the sub-model decoder 310 for training, along with audio features extracted from the input waveform of the audio stream, such as pitch, phase, rate, etc. In some embodiments, the sub-model decoder 310 also receives a speaker identifier used to distinguish between different voices in the audio stream.
[0050] Once the bottleneck trunk 305 and sub-model decoder 310 are trained, the bottleneck trunk 305 receives as input the audio stream generated by the client device 115. In some embodiments, the bottleneck trunk 305 includes a fully connected (FC) layer followed by an LSTM layer used to split the audio stream into increasingly abstract representations of the audio data.
[0051] The bottleneck trunk 305 receives the audio stream as an input waveform along with Mel-Frequency Cepstral Coefficient (MFCC) features sampled from overlapping samples of the audio stream. MFCC features are derived from a type of cepstral representation of the audio stream to represent information about the audio stream, such as a representation of the timbre within the audio stream. The samples of the audio stream are parameterized by a time window. For example, a time window of less than 100 milliseconds, which may be less accurate than a longer time window, may be used for computational efficiency and suitability for real-time execution on the client device 115. A client device 115 with less computational processing power than the server 101 can still generate accuracy using a shorter time window because the network latency is shorter than the network latency incurred when processing the audio stream on the server 101.
[0052] The bottleneck trunk 305 encodes the audio and outputs bottleneck features, which are sent to the sub-model decoder 310. In some embodiments, the bottleneck trunk 305 outputs bottleneck features for a subset of the audio frames, such as one out of every five frames in the audio stream. In some embodiments, the sub-model decoder 310 also receives audio-specific features related to the input waveform, such as the pitch, phase, and rate of the audio stream, as well as a speaker identifier.
[0053] In some embodiments, the sub-model decoder 310 includes a fully connected FC layer followed by an LSTM layer, which is used to capture temporal correlations in the audio stream. The sub-model decoder 310 outputs MFCC features, which are input to the audio waveform generator 315.
[0054] The audio waveform generator 315 receives the MFCC features and generates an output waveform representing the rendered MFCC features as synthesized audio.
[0055] Throughout this process, the machine learning model 300 learns a temporal mapping between positions in the audio stream and synthesizes future frames of the audio stream. For example, if the performance is "Happy Birthday," the machine learning model 300 may map the audio stream to the reference audio by determining that the audio stream contains the first two lines of the song and how the audio stream compares to the reference audio if it has a different starting point than the reference audio. This is referred to as a time offset in the audio between the audio stream and the reference audio. The machine learning model 300 may also output the rate of the audio stream compared to the reference audio. For example, a user singing "Happy Birthday" may sing at a faster rate than the reference audio. The machine learning model 300 may predict the timing of the next frame in the audio stream based on the time offset, the rate of the audio stream, and a future time offset of the song based on the rate of the audio stream compared to the rate of the reference audio, and consequently output synthesized audio that encapsulates those features.
[0056] In some embodiments, instead of performing the bottleneck extraction described above, the machine learning model 300 may implement a multilingual bottleneck extractor. The multilingual bottleneck extractor is trained to distinguish senones from multiple languages. The output features are language-independent and robust to variations in language, speaking style, speaking rate, etc. This machine learning model 300 is trained using supervised learning, and the output includes senone posterior probabilities.
[0057] Exemplary Machine Learning Models on Client Device 115 In embodiments in which the synthetic machine learning module 204 is stored on the server 101, the mapping machine learning model 204 may be trained using unsupervised learning using a training dataset having unlabeled audio streams.
[0058] 3B, a block diagram of another exemplary architecture of a machine learning model is shown. The machine learning model 350 includes a vector quantization variational autoencoder (VQ-VAE) 355, a VQ-VAE codebook 360, a prior model 365, and a VQ-VAE decoder 370.
[0059] The synthetic machine learning module 204 uses a training dataset containing waveforms of audio streams to train the VQ-VAE 355 and the VQ-VAE codebook 360. The VQ-VAE 355 and the VQ-VAE codebook 360 compare output waveforms with input waveforms to train corresponding modules of the machine learning model 350. The synthetic machine learning module 204 trains a prior model 365 on code vector outputs from the VQ-VAE codebook 360 conditioned on additional features, such as audio features including one or more of pitch, phase, and rate, and optionally a speaker identifier.
[0060] The machine learning model 350 is trained to receive an audio stream parameterized by a time window. In some implementations, the audio stream is 300-500 milliseconds long. Longer audio streams are more computationally intensive due to larger model sizes, but longer audio streams result in more accurate results.
[0061] In some embodiments, the VQ-VAE 355 receives an input waveform of an audio stream and converts the input waveform into a latent representation. The VQ-VAE 355 uses an autoregressive network structure that includes several one-dimensional (1D) convolution blocks. The VQ-VAE codebook 360 receives the latent representation as input, and the bottleneck quantizes the latent representation into a discrete code vector using a predefined codebook.
[0062] The a priori model 365 receives discrete code vectors from the VQ-VAE codebook 360 along with the pitch, phase, and rate of the input waveform and (optionally) one or more speaker identifiers for the audio stream. Audio features are computed from the input waveform, and the one or more speaker identifiers enable the network to model codes specific to a user or musical genre. The a priori model 365 uses a transform layer to extrapolate the code vectors over time and output edited and resampled code vectors. The VQ-VAE decoder 370 is an autoregressive audio waveform generator. The VQ-VAE decoder 370 receives the edited and resampled code vectors and synthesizes an output waveform from the edited and resampled code vectors.
[0063] The mixing module 206 mixes audio streams from multiple client devices 115 to form a combined audio stream that synchronizes one or more synthesized audio streams with the actual audio stream. In some embodiments, the mixing module 206 includes a set of instructions executable by the processor 235 to form the combined audio stream. In some embodiments, the mixing module 206 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0064] In some embodiments, the mixing module 206 receives the output of the synthetic machine learning module 204 and synchronizes the audio streams based on the output. For example, the mixing module 206 may receive the synthesized first audio stream and the second audio stream and generate a combined audio stream. Depending on whether the mixing is performed on the server 101 or the client device 115, the mixing module 206 may introduce different amounts of delay in the different audio streams to account for different types of delay and ensure that the audio streams are synchronized. For example, if the mixing is performed on the server 101, the mixing module 206 may account for receiver delay. Details of the various factors for mixing audio streams are discussed in more detail below.
[0065] The mixing module 206 synchronizes the combined first and second audio streams to form a combined audio stream. The time window of the audio streams may have different lengths based on the device that stores the synthetic machine learning module 204. For example, if the synthetic machine learning module 204 is stored on the client device 115, the time window is less than 100 milliseconds. In another example, if the synthetic machine learning module 204 is stored on the server 101, the time window is 300-500 milliseconds.
[0066] The blending module 206 may perform synchronization on the same device or a different device than the device on which the synthetic machine learning module 204 is stored. The devices may include the server 101, a client device 115a that transmits the first audio stream, and a client device 115b that receives the first audio stream.
[0067] In a first example, client device 115b performs both audio stream synthesis and audio stream synchronization. The synthesis machine learning module 204 receives globally time-stamped packets from all other client devices 115 via server 101 and synthesizes the audio streams of the other client devices 115. The mixing module 206 generates a local mix of the actual audio stream of client device 115b and the synthesized audio from the other client devices 115.
[0068] In some embodiments, the first example is the preferred example. There are several advantages to the first example: The time required to predict future portions of the audio stream is reduced, resulting in higher quality audio streams and shorter interaction latency. Additionally, client device 115b can recover from poor latency. While users may have to use client devices with significant computing power to synthesize and synchronize audio locally, better hardware can provide a better experience. Finally, this architecture offers the advantage that the synthesized audio remains on client device 115b.
[0069] Some possible drawbacks of the first example are that the computational complexity is O(n) for an incoming stream for n performers. Furthermore, because all processing is on the client device 115b and is processing-intensive, the processing may drain the battery of the client device 115b, the processing may be difficult, require different implementation code for different types of client devices 115, and may only work on a subset of client devices 115 with sufficient computing power.
[0070] In the second example, the server 101 performs both audio stream synthesis and audio stream synchronization. Advantages of the second example include a lower computational complexity of O(1) because all client streams are received and processed at the server. In this example, the server 101 receives individual audio streams from each client and provides a synthesized, synchronized, combined audio stream to each client, so no client-side computation is required to generate the combined audio stream. The infrastructure is all part of the server 101 and therefore controlled. Furthermore, the worst-case latency of the audio stream is only the latency between the client and the server, whereas with client-side processing, it may be between client device 115a and the server 101 and then between client device 115b, so the worst-case expected time is shorter than when client-side processing is used.
[0071] In a third example, the client device 115a performs audio stream synthesis, and the server 101 performs audio stream synchronization. The client device 115a synthesizes the audio streams for worst-case latency, and the server 101 mixes (n-1) streams in lockstep at the same time for the player by delaying each audio stream as needed. Advantages of the third example include O(1) incoming and outgoing streams per client device 115 and O(1) processing time at the server. Additionally, better hardware at the client device 115a makes the client device 115a sound better to other client devices 115. For example, better hardware results in better processing, which is essential when audio streams must be processed quickly to maintain near-real-time transmission of the audio streams.
[0072] In a fourth example, client device 115a performs the composition of the audio streams and client device 115b performs the synchronization of the audio streams, which may be suitable for peer-to-peer streaming without a server.
[0073] If no post-processing occurs and the mixing module 206 does not perform the combining on the client device 115b, the mixing module 206 instructs the I / O interface 239 to send the combined audio stream to the client device 115b, which can play the combined audio stream. If no post-processing occurs and the mixing module 206 is on the client device 115b, the mixing module 206 can instruct the I / O interface 239 to provide the combined audio stream to the speaker 243 for playback.
[0074] The post-processing module 208 processes the combined audio stream. In some embodiments, the post-processing module 208 includes a set of instructions executable by the processor 235 to process the combined audio stream. In some embodiments, the post-processing module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0075] In some embodiments, the post-processing module 208 may modify the combined audio stream to match the acoustics of the environment in which the client device 115b is located by taking into account echoes, background noise, etc. In some embodiments, the post-processing module 208 performs audio cleaning or noise suppression on the combined audio stream to avoid a situation where, for example, if a first audio stream includes a car honking in the background and a second audio stream includes a completely quiet room, a cacophony may result when the combined audio stream goes from having background noise to not having background noise.
[0076] The user interface module 210 generates the user interface. In some embodiments, the user interface module 210 includes a set of instructions executable by the processor 235 to generate the user interface. In some embodiments, the user interface module 210 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0077] The user interface module 210 generates a user interface for a user 125 associated with a client device 115. The user interface may be used to initiate audio communications with other users, participate in games or other experiences within the metaverse, send text to other users, initiate video communications with other users, etc.
[0078] In some embodiments, before a user joins the metaverse, the user interface module 214 generates a user interface that includes information about how the user's information will be collected, stored, and analyzed. For example, the user interface requests the user to provide permission to use any information associated with the user. The user is notified that user information may be deleted by the user, and the user may have the option to select which types of information are provided for different uses. Use of the information complies with applicable regulations, and the data is securely stored. Data collection may not be performed in certain locations and for certain user categories (e.g., based on age or other demographics), data collection is temporary (i.e., the data is destroyed after a certain period of time), and the data is not shared with third parties. Some of the data may be anonymized, aggregated across users, or otherwise altered so that the identity of a particular user cannot be determined.
[0079] In some embodiments, a user interface obtains a user's permission before any audio stream is sent to the server 101 or another client device 115. The user interface may include different levels of granularity for user permissions. For example, a user may specify that synthesized audio can be generated only if the synthesized audio is generated on the client device 115 and not on the server.
[0080] In some embodiments, the mixing module 206 determines that the time difference between the first audio stream and the second audio stream exceeds a threshold time difference and sends instructions to the user interface module 210 to generate a user interface. The user interface may provide guidance to the user 125 on how to change the singing speed so that the audio stream can be synchronized with another audio stream.
[0081] 4, an exemplary user interface 400 is shown that guides a user to modify timing aspects of a performance. In this example, the user interface 400 includes user guidance regarding the performance and a movement indicator 405 that prompts a performer associated with the client device 115b to perform in a manner that reduces the time difference between the first audio stream and the second audio stream. For example, the movement indicator 405 may move slower than the user is performing to suggest a slowdown in the user's performance.
[0082] In some embodiments, the user interface module 210 indicates synchronization of the combined audio stream with the actions of a graphically displayed performer. For example, if the combined audio stream is speech that a performer is performing while graphically displayed as an avatar, the user interface module 210 may synchronize the combined audio stream with the avatar's mouth, movements, etc.
[0083] Exemplary Methods 5 is an exemplary flow diagram 500 illustrating the transmission of data between a client device 115 and a server 101, according to some embodiments described herein. The flow diagram 500 includes a first client device 510, a server 515, and a second client device 520. Bold lines indicate network data transmission between the three devices, and thin lines indicate data transmission within the first client device 510.
[0084] A first client device 510 receives an audio stream from a microphone 505. The audio stream is combined at the first client device 510, the server 515, or the second client device 520. The combined audio stream is mixed with one or more other audio streams at the server 515 or the second client device 520 to form a combined audio stream. The combined audio stream is sent to a speaker 525 for playback at the first client device 510.
[0085] In addition to the above, data is also received by the server 515. For example, the server 515 receives a first stream from the first client device 510 and a second stream from the second client device 520, and the server 515 synthesizes the first audio stream and mixes it with the second audio stream.
[0086] As a result of implementing the metaverse engine 103, a first audio stream of a violin and a second audio stream of a cello at a first client device 510 are heard simultaneously by associated users playing the music as if they were in the same physical room.
[0087] 6 is an example flow diagram 600 for synthesizing an audio stream for synchronous communication. Task detection 605 is performed on the delayed audio stream to identify a performance identifier and a type of performance. For example, identification is made of whether the delayed audio stream contains instruments, speech, singing, etc., as well as what type of song and style of performance it is. Based on the identification, reference audio is output.
[0088] The delayed audio stream is also received by a phase shift analyzer 610, which identifies the time offset of the delayed audio stream compared to the reference audio, and a rate analyzer 615, which identifies the rate of the delayed audio stream. The time offset and rate of the delayed audio stream are received by a sampler 620, which maps the audio stream to the reference audio. The mapping is received by a deep neural network 635 (or other suitable model), which extrapolates the audio stream into the future to complete the synthesis of the audio stream.
[0089] 7 is an example flow diagram for synthesizing an audio stream for synchronous communication using a second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.
[0090] The method 700 may begin at block 702. At block 702, a first audio stream of a performance associated with a first client device 115 is received. Block 702 may be followed by block 704.
[0091] In block 704, during a time window of the performance that is shorter than the total time of the performance, a combined audio stream is formed that synchronizes the combined first audio stream with the second audio stream associated with the second client device, and the first audio stream and the second audio stream are mixed to form a combined audio stream that synchronizes the combined first audio stream with the second audio stream associated with the second client device, the time window is advanced and the generating and mixing is repeated until the performance is complete.
[0092] 7 is an example flow diagram for synthesizing an audio stream for synchronous communication using a second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.
[0093] 8 is another exemplary flow diagram for synthesizing audio for synchronous communication using a server. In this example, the metaverse engine 103 is stored on the server 101.
[0094] The method 800 may begin at block 802. At block 802, a first audio stream of a performance associated with a first client device 115 is received. Block 802 may be followed by block 804.
[0095] In block 804, during a time window of the performance that is shorter than the total time of the performance, the combined first audio stream and the second audio stream are formed to form a combined audio stream that synchronizes the combined first audio stream and the second audio stream associated with the second client device, and the first audio stream and the second audio stream are mixed to form a combined audio stream that synchronizes the combined first audio stream and the second audio stream associated with the second client device, the time window is advanced, and the generating and mixing are repeated until the performance is complete. In some embodiments, mixing the combined first audio stream includes introducing a delay into the combined audio stream to account for latency incurred by transmitting the combined audio to the second client device. Block 804 may be followed by block 806.
[0096] In block 806 , the combined audio stream is transmitted to the second client device 115 .
[0097] Various embodiments described herein involve acquiring data from various sensors in a physical environment, analyzing such data, generating recommendations, and providing a user interface. Data collection is performed only with specific user permission and in accordance with applicable regulations. Data is stored in accordance with applicable regulations, including anonymizing or otherwise modifying the data to protect the user's privacy. Users are provided with clear information regarding data collection, storage, and use and are given the option to select the types of data that may be collected, stored, and utilized. Furthermore, users control the devices on which data may be stored (e.g., client device only, client device + server device, etc.) and the devices on which data analysis is performed (e.g., client device only, client device + server device, etc.). The data is utilized for the specific purposes described herein. The data is not shared with third parties without explicit user permission.
[0098] The methods, blocks, and / or operations described herein may be performed in a different order than illustrated or described, and / or may be performed concurrently (partially or completely) with other blocks or operations as appropriate. Some blocks or operations may be performed on a portion of the data and, for example, performed again later on another portion of the data. Not all of the blocks and operations described need be performed in various implementations. In some implementations, blocks and operations may be performed multiple times in a method, in a different order, and / or at different times.
[0099] In the above description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described above primarily with reference to a user interface and specific hardware. However, embodiments may apply to any type of computing device capable of receiving data and commands and any peripheral device that provides services.
[0100] References herein to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with the embodiments or examples may be included in at least one implementation of the description. The appearances of the phrase "in some embodiments" in various places in this specification do not necessarily all refer to the same embodiments.
[0101] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0102] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. As will become apparent from the discussion that follows, unless specifically stated otherwise, throughout the description, discussions utilizing terms including "processing" or "computing" or "calculating" or "determining" or "displaying," etc., will be understood to refer to the actions and processing of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display device.
[0103]
[0013] Embodiments herein may also relate to a processor for performing one or more steps of the methods described above. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including, but not limited to, any type of disk, including an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0104] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0105] Furthermore, the descriptions may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any device that can contain, store, communicate, propagate, or transfer a program for use by or in connection with an instruction execution system, device, or drive.
[0106] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly via a system bus to memory elements that may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution. [Explanation of symbols]
[0107] 100 Network environment, environment 101 Server 103 Metaverse Engine 104 Metaverse Applications 104a Metaverse Applications 104b Metaverse Applications 105 Network 115 client devices 115a...n client device 125 users 125a...n users 199 databases 200 computing devices 202 Performance Recognition Module 204 Synthetic Machine Learning Module, Machine Learning Module, Mapping Machine Learning Model 206 Mixed Module 208 Post-processing module 210 User Interface Module 218 Bus 222 signal line 224 signal line 226 Signal Line 228 Signal Line 230 Signal Line 232 signal line 234 Signal Line 235 processor 237 memory 239 Input / Output (I / O) Interface, I / O Interface 241 Microphone 243 Speaker 245 display 247 Storage Devices 300 machine learning models 305 Bottleneck Trunk 310 Submodel Decoder 315 Audio Waveform Generator 350 machine learning models 355 Vector Quantized Variational Autoencoder (VQ-VAE), VQ-VAE 360 VQ-VAE Codebook 365 Advance Model 370 VQ-VAE decoder 400 User Interface 405 Movement Indicator 500 Flow Diagram 505 Microphone 510 First Client Device 515 Server 520 Second Client Device 525 Speaker 600 Flow Diagram 605 Task Discovery 610 Phase Shift Analyzer 615 Rate Analyzer 620 Sampler 635 Deep Neural Networks
Claims
1. 1. A computer-implemented method comprising: receiving a first audio stream of a performance associated with a first client device; During a time window of the performance that is shorter than the total time of the performance, generating a synthesized first audio stream that predicts a future of the performance based on audio features of the first audio stream; mixing the combined first audio stream and the second audio stream to form a combined audio stream that synchronizes the combined first audio stream and a second audio stream associated with a second client device; Including, The time window is advanced and the generating and mixing steps are repeated until the performance is complete.
2. in response to receiving the first audio stream; determining a performance identifier for the performance associated with the first audio stream; receiving reference audio based on the performance identifier; The method of claim 1 further comprising:
3. generating the synthesized primary audio stream includes determining a time offset between the primary audio stream and the reference audio; 3. The method of claim 2, wherein the time offset occurs when the primary audio stream has a different starting point than the reference audio, and wherein generating the synthesized primary audio stream is further based on the time offset.
4. generating the synthesized first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio; The method of claim 2 , wherein generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.
5. The method of claim 1 , wherein the audio features of the first audio stream are selected from the group of pitch, rate, phase, or combinations thereof.
6. The method of claim 1 , wherein the audio features of the first audio stream include one or more speaker identifiers detected in the first audio stream.
7. determining that a time difference between the first audio stream and the second audio stream exceeds a threshold time difference; generating graphical data for displaying a user interface including user guidance regarding the performance and a movement indicator prompting a performer associated with the second client device to perform in a manner that reduces the time difference between the first audio stream and the second audio stream; The method of claim 1 further comprising:
8. The method of claim 1 , further comprising modifying the combined audio stream to match the acoustics of an environment in which the second client device is located.
9. generating the synthesized first audio stream comprises: identifying in the first audio stream that a portion of the performance has been skipped; synthesizing the first audio stream to correct the portion of the performance that was skipped; 2. The method of claim 1, comprising:
10. The method of claim 1 , further comprising synchronizing the combined audio stream to coincide with graphically displayed actions of a performer.
11. A device, a processor; a memory coupled to the processor and having instructions stored thereon, the instructions, when executed by the processor, causing the processor to: receiving a first audio stream of a performance associated with a first client device; During a time window of the performance that is shorter than the total time of the performance, generating a synthesized first audio stream that predicts a future of the performance based on audio features of the first audio stream; mixing the combined first audio stream and the second audio stream to form a combined audio stream that synchronizes the combined first audio stream and a second audio stream associated with a second client device; Execute an operation including The time window is advanced and the generating and mixing are repeated until the performance is complete.
12. In response to receiving the first audio stream, determining a performance identifier for the performance associated with the first audio stream; The device of claim 11 , further comprising: a device configured to receive reference audio based on the performance identifier.
13. generating the synthesized primary audio stream includes determining a time offset between the primary audio stream and the reference audio; 13. The device of claim 12, wherein the time offset occurs when the primary audio stream has a different starting point than the reference audio, and generating the synthesized primary audio stream is further based on the time offset.
14. generating the synthesized first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio; 13. The device of claim 12, wherein generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.
15. The device of claim 11 , wherein the audio features of the first audio stream are selected from the group of pitch, rate, phase, or combinations thereof.
16. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations including: receiving a first audio stream of a performance associated with a first client device; During a time window of the performance that is shorter than the total time of the performance, generating a synthesized first audio stream that predicts a future of the performance based on audio features of the first audio stream; mixing the combined first audio stream and the second audio stream to form a combined audio stream that synchronizes the combined first audio stream and a second audio stream associated with a second client device; Including, The time window is advanced and the generating and mixing are repeated until the performance is complete.
17. In response to receiving the first audio stream, determining a performance identifier for the performance associated with the first audio stream; The computer-readable medium of claim 16 , further comprising receiving reference audio based on the performance identifier.
18. generating the synthesized primary audio stream includes determining a time offset between the primary audio stream and the reference audio; 18. The computer-readable medium of claim 17, wherein the time offset occurs when the first audio stream has a different starting point than the reference audio, and generating the synthesized first audio stream is further based on the time offset.
19. generating the synthesized first audio stream includes determining a rate of the first audio stream compared to a rate of the reference audio; 20. The computer-readable medium of claim 17, wherein generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.
20. 17. The computer-readable medium of claim 16, wherein the audio features of the first audio stream are selected from the group of pitch, rate, phase, or combinations thereof.
Citation Information
Patent Citations
Accompaniment method for actively following music signals and related equipment
CN112669798A
Audiovisual synchronously compositing and distributing method, device for player'S terminal, program for the device and recording medium where the program for the device is recorded, service providing device, and program for the service providing device and recording medium recorded with the program for the device
JP2003167575A
Remote multipoint concert system using network
JP2007041320A
Remote music performance system
JP2008089849A
Device and method for distributing data, and program
JP2009005012A