Audio synthesis for synchronous communication
The system generates and synchronizes audio streams across networks by predicting future performance and mixing them to overcome latency, ensuring synchronized audio perception across different locations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-10-02
- Publication Date
- 2026-04-15
AI Technical Summary
Multiple individuals connected via a computer network cannot perform music, sing, or talk in synchronization due to network, input, and processing latencies, causing desynchronization.
A method and system that generates a synthesized audio stream predicting future performance based on audio characteristics, mixing it with other streams to synchronize audio across different client devices, accounting for latency and environment.
Synchronizes audio performances across different locations by masking delays, allowing users to perceive synchronized audio streams without latency issues.
Smart Images

Figure 0007846831000001 
Figure 0007846831000002 
Figure 0007846831000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application is an international application and claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Patent Application No. 17 / 959,736, filed on October 4, 2022, titled SYNTHESIZING AUDIO FOR SYNCHRONOUS COMMUNICATION, the entire content of which is incorporated herein by reference.
Background Art
[0002] It is impossible for multiple people who are physically in different locations but connected by a computer network to play music, sing, or talk in synchronization. This is because the performance of person P0 observed by person P1 is always in the past due to the transmission delay of the computer network. If each of the n - 1 performers P1, P2, P3,...P(n - 1) is delayed by exactly the appropriate amount of time relative to performer P0, P0 will observe that all others are synchronized with each other, but cannot synchronize with the other performances themselves.
[0003] The delay may be due to one or more of at least three types of latency: network latency, input latency, and processing latency. Network latency occurs when there is a shortage in the network transmission time due to the physical devices used in the network or a delay in the processing time at a node. Input latency occurs when the user delays a response, or when the user is using a client device with limited capabilities. Processing latency may occur when a delay is intentionally introduced to perform moderation analysis.
[0004] The background information provided herein is for the purpose of presenting the context of this disclosure. The research of the currently named inventors up to the extent described in this background section, as well as aspects of the specification that may not qualify as prior art at the time of filing, are not expressly or implicitly recognized as prior art to this disclosure. [Overview of the project] [Means for solving the problem]
[0005] The embodiments generally relate to systems and methods for synthesizing audio for synchronous communication. According to one embodiment, a method implemented by a computer includes the step of receiving a first audio stream of performance associated with a first client device. The method further includes the step of generating a synthesized first audio stream that predicts the future of the performance based on the audio characteristics of the first audio stream during a performance time window shorter than the total duration of the performance, and the step of mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device, the time window advancing and the generation and mixing steps being repeated until the performance is complete.
[0006] In some embodiments, the method further includes the steps of determining a performance identifier associated with a performance in response to the reception of a first audio stream, and receiving a reference audio based on the performance identifier. In some embodiments, the step of generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, the time offset occurring if the first audio stream has a different start point than the reference audio, and the step of generating the synthesized first audio stream is further based on the time offset. In some embodiments, the step of generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio, and the step of generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from a group of pitch, rate, phase, or combinations thereof. In some embodiments, the audio features of the first audio stream include one or more speaker identifiers detected in the first audio stream. In some embodiments, the method further includes the steps of determining whether the time difference between a first audio stream and a second audio stream exceeds a threshold time difference, and generating graphical data for displaying a user interface that includes user guidance regarding performance and movement indicators prompting a performer associated with a second client device to perform in a way that reduces the time difference between the first audio stream and the second audio stream. In some embodiments, the method further includes the step of modifying the combined audio streams to match the acoustics of the environment in which the second client device is located.In some embodiments, the step of generating a synthesized first audio stream includes identifying that a portion of the performance has been skipped in the first audio stream, and synthesizing the first audio stream to correct the skipped portion of the performance. In some embodiments, the method further includes synchronizing the combined audio streams to match the actions of a graphically displayed performer.
[0007] In some embodiments, the device includes a processor and a memory coupled to the processor where instructions are stored. When an instruction is executed by the processor, the device causes the processor to perform operations including: receiving a first audio stream of performance associated with a first client device; generating a synthesized first audio stream that predicts the future of the performance based on the audio characteristics of the first audio stream during a performance time window shorter than the total performance time; and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device, with the time window advancing and the generation and mixing being repeated until the performance is complete.
[0008] In some embodiments, in response to receiving a first audio stream, a performance identifier for the performance associated with the first audio stream is determined, and a reference audio is received based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, where the time offset occurs if the first audio stream has a different starting point than the reference audio, and generating the synthesized first audio stream is further based on the time offset. In some embodiments, generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio, and generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from a group of pitch, rate, phase, or combinations thereof.
[0009] In some embodiments, when executed by one or more processors, a non-temporary computer-readable medium storing instructions causing one or more processors to perform an operation, the operation comprising: receiving a first audio stream of performance associated with a first client device; generating a synthesized first audio stream that predicts the future of the performance based on the audio characteristics of the first audio stream during a performance time window shorter than the total performance time; and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream and a second audio stream associated with a second client device, the time window advancing and the generation and mixing being repeated until the performance is complete.
[0010] In some embodiments, in response to receiving a first audio stream, a performance identifier for the performance associated with the first audio stream is determined, and a reference audio is received based on the performance identifier. In some embodiments, generating a synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, where the time offset occurs if the first audio stream has a different starting point than the reference audio, and generating the synthesized first audio stream is further based on the time offset. In some embodiments, generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio, and generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio. In some embodiments, the audio features of the first audio stream are selected from a group of pitch, rate, phase, or combinations thereof.
[0011] This application describes a metaverse engine and / or metaverse application that advantageously generates a synthesized first audio stream that predicts future performance based on the audio characteristics of the first audio stream, and mixes the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device so that the user who created the audio stream is perceived as singing or speaking in sync. The generation of the synthesized first audio stream and the mixing of the synthesized first audio stream and the second audio stream may be performed on different devices, including a combination of a first audio device, a second audio device, and a server. As a result, the method is distributed across multiple devices, and any delays in the streaming process, synthesis process, or mixing process are masked by the mixing step so that the user listening to the performance does not perceive any latency. [Brief explanation of the drawing]
[0012] [Figure 1] This is a block diagram of an exemplary network environment for synthesizing audio for synchronous communication, according to some embodiments described herein. [Figure 2] This is a block diagram of an exemplary computing device for synthesizing audio for synchronous communication, according to some embodiments described herein. [Figure 3A] This is a block diagram of an exemplary architecture of a machine learning model according to some embodiments described herein. [Figure 3B] This is a block diagram of another exemplary architecture of a machine learning model, according to some embodiments described herein. [Figure 4]This is a diagram of an exemplary user interface that guides the user to change the timing of performance according to some embodiments described herein. [Figure 5] This is an exemplary flowchart illustrating data transmission between a client device and a server according to some embodiments described herein. [Figure 6] This is an exemplary flowchart for synthesizing audio streams according to some embodiments described herein. [Figure 7] This is an exemplary flowchart for synthesizing audio streams for synchronous communication, according to some embodiments described herein. [Figure 8] This is an exemplary flowchart for synthesizing audio for synchronous communication using a server, according to some embodiments described herein. [Modes for carrying out the invention]
[0013] Network environment 100 Figure 1 shows a block diagram of an exemplary environment 100 for synthesizing audio for synchronous communication. In some embodiments, the environment 100 includes a server 101 connected via a network 105 and client devices 115a...n. Users 125a...n may be associated with each client device 115a...n. In Figure 1 and the remaining figures, a letter following a reference number, e.g., "115a", represents a reference to an element having that particular reference number. A reference number in text without a following letter, e.g., "115", represents a general reference to an element having that reference number. In some embodiments, the environment 100 may include other servers or devices not shown in Figure 1. For example, server 101 could be multiple servers 101.
[0014] Server 101 includes one or more servers, each including a processor, memory, and network communication hardware. In some embodiments, Server 101 is a hardware server. Server 101 is communicably coupled to Network 105. In some embodiments, Server 101 sends data to and receives data from Client Devices 115. Server 101 may include a metaverse engine 103 and a database 199.
[0015] In some embodiments, the metaverse engine 103 includes code and routines capable of facilitating communication between client devices 115 associated with two or more users in a virtual metaverse, such as communication in the same location within the metaverse, communication within the same metaverse experience, or communication between friends within a metaverse application. Users interact across diverse demographics (e.g., different ages, regions, languages, etc.) within the metaverse.
[0016] In some embodiments, the metaverse engine 103 performs some or all of the following steps: generating a synthesized first audio stream that predicts future performance based on the audio characteristics of the first audio stream, and mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream with a second audio stream associated with a second client device. Different embodiments regarding how the steps are divided are discussed in more detail below.
[0017] In some embodiments, the metaverse engine 103 is implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the metaverse engine 103 is implemented using a combination of hardware and software.
[0018] The database 199 can be a non-transitory computer-readable memory (e.g., random access memory), a cache, a drive (e.g., hard drive), a flash drive, a database system, or another type of component or device capable of storing data. The database 199 can also include multiple storage components (e.g., multiple drives or multiple databases) that can span multiple computing devices (e.g., multiple server computers). The database 199 can store data associated with the metaverse engine 103, such as a training dataset for a trained machine learning model, reference audio, etc.
[0019] The client device 115 can be a computing device including a memory and a hardware processor. For example, the client device 115 can include a mobile device, a tablet computer, a cellular phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device capable of accessing the network 105.
[0020] Client device 115a includes a metaverse application 104a, and client device 115n includes a metaverse application 104b. In some embodiments, user 125a uses the metaverse application 104a on client device 115a to generate a communication such as a first audio stream, and the communication is transmitted to the metaverse engine 103 on server 101. Server 101 transmits the communication to the metaverse application 104b on client device 115b for user 125n.
[0021] In some embodiments, the metaverse application 104 performs some or all of the steps of generating a synthesized first audio stream that predicts the future performance based on the audio characteristics of the first audio stream, and mixing the synthesized first audio stream and a second audio stream associated with a second client device to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream. For example, the metaverse application 104a on client device 115a may generate a synthesized first audio stream, and the metaverse application 104b on client device 115n may mix the synthesized first audio stream and the second audio stream. In other embodiments, both synthesis and mixing may be performed at server 101, and the metaverse application 104b on client device 115n outputs the mixed synthesized first audio stream and the second audio stream via a speaker.
[0022] In the illustrated embodiment, entities in environment 100 are connected communicably via a network 105. Network 105 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, a Wi-Fi® network, or a wireless LAN (WLAN)), a cellular network (e.g., a Long-Term Evolution (LTE) network), a router, a hub, a switch, a server computer, or a combination thereof. Figure 1 shows one network 105 connected to a server 101 and a client device 115, but in practice, one or more networks 105 may be connected to these entities.
[0023] 200 Examples of Computing Devices Figure 2 is a block diagram of an exemplary computing device 200 that may be used to implement one or more features described herein. The computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In some embodiments, the computing device 200 is a server 101. In some embodiments, the computing device 200 is a client device 115.
[0024] In some embodiments, the computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, microphone 241, speaker 243, display 245, and storage device 247, each coupled via a bus 218. Depending on whether the computing device 200 is a server 101 or a client device 115, some components of the computing device 200 may be absent. For example, if the computing device 200 is a server 101, the computing device may not include the microphone 241 and speaker 243. In some embodiments, the computing device 200 includes additional components not shown in Figure 2.
[0025] The processor 235 may be connected to the bus 218 via signal line 222, the memory 237 may be connected to the bus 218 via signal line 224, the I / O interface 239 may be connected to the bus 218 via signal line 226, the microphone 241 may be connected to the bus 218 via signal line 228, the speaker 243 may be connected to the bus 218 via signal line 230, the display 245 may be connected to the bus 218 via signal line 232, and the storage device 247 may be connected to the bus 218 via signal line 234.
[0026] The processor 235 includes an arithmetic logic unit, microprocessor, general-purpose controller, or some other processor array for performing calculations and providing instructions to a display device. The processor 235 may include various computing architectures, including composite instruction set computer (CISC) architectures, reduced instruction set computer (RISC) architectures, or architectures that process data and implement combinations of instruction sets. Figure 2 shows a single processor 235, but multiple processors 235 may be included. In different embodiments, the processor 235 may be a single-core processor or a multi-core processor. Other processors (e.g., graphics processing units), operating systems, sensors, displays, and / or physical configurations may be part of the computing device 200.
[0027] Memory 237 stores instructions and / or data that can be executed by processor 235. Instructions may include code and / or routines for performing the techniques described herein. Memory 237 may be a dynamic random access memory (DRAM) device, static RAM, or any other memory device. In some embodiments, memory 237 also includes a static random access memory (SRAM) device or non-volatile memory such as flash memory, or similar persistent storage devices and media, including a hard disk drive, a compact disc read-only memory (CD-ROM) device, a DVD-ROM device, a DVD-RAM device, a DVD-RW device, a flash memory device, or any other mass storage device for storing information more persistently. Memory 237 includes code and routines that can operate to run the metaverse engine 103, which will be described in more detail below.
[0028] The I / O interface 239 can provide the functionality to interface the computing device 200 with other systems and devices. Interfaced devices can be included as part of the computing device 200 or communicate with the computing device 200 separately. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 247), and input / output devices can communicate via the I / O interface 239. In another example, the I / O interface 239 can receive data from server 101 and supply the data to the metaverse engine 103 and components of the metaverse engine 103 such as synthetic machine learning module 204. In some embodiments, the I / O interface 239 can connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone 241, sensor, etc.) and / or output devices (display device, speaker 243, monitor, etc.).
[0029] Some examples of interfaced devices that can be connected to the I / O interface 239 include a display 245 that can be used to display content, such as images, videos, and / or the user interface of the output application described herein, and to receive touch (or gesture) input from the user. The display 245 may include any suitable display device such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, 3D display screen, or other visual display device.
[0030] Microphone 241 includes hardware for detecting audio performed by user 125. For example, microphone 241 may detect user 125 singing, user 125 playing the violin, etc. Microphone 241 may transmit audio to metaverse engine 103 via I / O interface 239.
[0031] Speaker 243 includes hardware for generating audio for playback. For example, speaker 243 receives commands from the metaverse engine 103 to generate perceptual audio for playback from a digitally combined audio stream generated by the metaverse engine 103. Speaker 243 converts the commands into audio and generates a combined audio stream for the user.
[0032] The storage device 247 stores data related to the metaverse engine 103. For example, the storage device 247 may store training datasets for trained machine learning models, reference audio, etc. In embodiments where the computing device 200 is the server 101, the storage device 247 is the same as the database 199 in Figure 1.
[0033] Example: Metaverse engine 103 or metaverse application 104 Figure 2 shows a computing device 200 running an exemplary metaverse engine 103 or metaverse application 104, which includes a performance awareness module 202, a synthetic machine learning module 204, a hybrid module 206, a post-processing module 208, and a user interface module 210. Although the modules are shown as being part of the same metaverse engine 103 or metaverse application 104, those skilled in the art will recognize that the modules may be implemented by any computing device 200. For example, the performance awareness module 202 and the synthetic machine learning module 204 may be part of a client device 115, while the hybrid module 206 may be part of a server 101 to reduce the computational requirements of the client device 115.
[0034] The performance awareness module 202 determines a performance identifier associated with an audio stream. In some embodiments, the performance awareness module 202 includes a set of instructions that can be executed by the processor 235 to determine the performance identifier associated with an audio stream. In some embodiments, the performance awareness module 202 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0035] An audio stream may be a performance or rendition of known content, such as a re-recording of a known song, singing a known melody, or reading text material aloud. In this context, a performance identifier refers to the known content being performed within the audio stream. Performance identifiers can enable the identification and / or prediction of various attributes of the audio stream. For example, if the performance is based on a written musical score (including, for example, the tempo and note notation to be played by various instruments), the tempo of the performance and the identification of individual notes played by each performer may be identified. In another example, if the performance involves singing or reading from text, the next word / phrase may be identified.
[0036] In some embodiments, after obtaining user permission, the performance awareness module 202 receives the audio stream. For example, if the performance awareness module 202 is part of a client device 115, the performance awareness module 202 receives the audio stream from the microphone 241 via the I / O interface 239 over the network 105. In another example, if the performance awareness module 202 is part of a server 101, the performance awareness module 202 receives the audio stream from the client device 115 via the I / O interface 239. The audio stream is part of the performance. The performance could be songs sung by multiple users, speeches, hymns, music played by instruments, etc. In some embodiments, the user interface module 210 obtains permission to use the audio stream before the performance awareness module 202 receives the audio stream.
[0037] The performance recognition module 202 generates a spectrogram of the audio stream, including frequency as a function of time, determines which frequencies in the audio stream have the highest amplitude, and then generates a hash of the spectrogram to produce a fingerprint (i.e., an audio fingerprint) from at least a portion of the audio stream. Since the audio stream is performed live by a person (as opposed to a pre-recorded performance), the performance recognition module 202 takes into account errors, including key differences and timing discrepancies, as well as vocal quality. To identify matches, the performance recognition module 202 compares the audio stream fingerprint to a set of fingerprints for known songs. Matches are associated with a performance identifier for the audio stream. For example, the performance identifier could be "Happy Birthday" or "Moonlight Sonata, 3rd Movement."
[0038] The performance awareness module 202 may receive reference audio associated with a performance identifier. For example, the performance awareness module 202 may obtain reference audio from the storage device 247.
[0039] After obtaining user permission, the performance recognition module 202 may determine the type of performance in the audio stream. For example, the audio stream may include a person singing, a person speaking, a person playing a musical instrument, and so on. In some embodiments, the performance recognition module 202 determines the type of performance based on the unique characteristics of various instruments and human voices, such as by identifying various frequencies, timings, pitches, speeds, etc., associated with the instruments or human voices. For example, the performance recognition module 202 may determine that "Happy Birthday" is being played on a kazoo and "Moonlight Sonata, 3rd Movement" is being played on a piano in the metaverse. In some embodiments, the performance recognition module 202 may further determine the style of the performance, such as whether the singer is performing a blues version of the song or a rock version.
[0040] In some embodiments, the performance recognition module 202 obtains permission from the user 125 to identify one or more speakers in the audio stream and associate one or more speakers with a speaker identifier. Permission may include permission to use the audio stream for identification purposes, permission to store information about the user, etc. The user 125 is given guidance that this information may be stored (e.g., temporarily, until the performance ends) and is given the option to deny permission, select a storage attribute (e.g., local only, X hours, etc.). After obtaining permission, the performance recognition module 202 may determine, for example, that the audio stream associated with the first client device 115 contains one person singing "Happy Birthday" or multiple people performing the song. The performance recognition module 202 may identify the different people performing the song based on a mechanism similar to those described above, where each performer is associated with a specific way of speaking, singing, etc., based on rhythm, tone, frequency, etc.
[0041] The synthetic machine learning module 204 trains a machine learning model (or multiple models) for synthesizing an audio stream. In some embodiments, the synthetic machine learning module 204 includes a set of instructions that can be executed by the processor 235 to train the machine learning model for synthesizing an audio stream. In some embodiments, the synthetic machine learning module 204 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0042] The synthetic machine learning module 204 may use one or more (e.g., two) different training datasets. In some implementations, the dataset and corresponding structure may depend on whether the synthetic machine learning module 204 is stored on the client device 115 or on the server 101.
[0043] Exemplary machine learning module 204 on client device 115 In an embodiment in which the synthetic machine learning module 204 is stored on the client device 115, the mapping machine learning model 204 implements supervised learning by training the machine learning model using a training dataset having manually labeled audio streams.
[0044] The synthetic machine learning module 204 may be a neural network, such as a deep neural network (DNN), and includes layers that identify more detailed features and patterns within an audio stream, with the output of one layer serving as input to a subsequent layer. The output layer generates a synthesized audio stream. In some embodiments, a first layer (or a first set of layers) outputs a mapping of the audio stream to a reference audio to determine the position of the audio stream compared to the reference audio, a second layer (or a second set of layers) outputs a prediction of future time offsets of the audio stream, and the output layer synthesizes the audio stream based on the outputs of the previous layers. In some embodiments, the machine learning model may include a long-short-term memory (LSTM) within a recurrent neural network (RNN) trained for sequential processing of audio streams. The machine learning model can receive any audio stream as input and perform the same analysis to synthesize an output.
[0045] The synthesized audio stream takes into account network latency, input latency, and / or processing latency, which cause audio delays. Network latency occurs when there is a delay in the reception of content from a transmitting node to a receiving node in the network, due to the physical equipment used in the network or processing time delays at the nodes. Input latency occurs when a user delays their response or when using a client device with limited capabilities. Processing latency can occur when delays are intentionally introduced to perform moderation analysis or due to insufficient processing resources. Since network latency is variable, the true end-to-end latency for each audio stream can only be determined by the client device 115 on which the audio stream is played via speaker 243.
[0046] In some embodiments, the synthetic machine learning module 204 may advantageously consider different types of input latency. For example, the performance is Vivaldi's "The Four Seasons," and the first audio stream is a violin with a delayed third note (determined using a performance identifier that allows for the identification of the notes and tempo of The Four Seasons). The machine learning model may adjust the first audio stream based on the input latency by outputting a synthesized first audio stream that takes into account the missing third note by skipping the sixteenth note. In another example, the synthetic machine learning module 204 may consider processing latency introduced for moderation by adding a delay to the end of sentences in the speech. Adding a delay to the end of sentences may be advantageous because the perception of the delay has less impact on the listener than a delay at any point in the sentence.
[0047] Looking at Figure 3A, a block diagram of an exemplary architecture of the machine learning model 300 is shown. The machine learning model 300 includes a bottleneck trunk 305, a submodel decoder 310, and an audio waveform generator 315. The bottleneck trunk 305 is a general-purpose bottleneck trunk that extracts features from the input waveform of an audio stream. The submodel decoder 310 is trained to perform a specific task, namely, determining temporal correlations within an audio stream compared to a reference audio (determined, for example, based on a performance identifier).
[0048] The bottleneck trunk 305 and submodel decoder 310 are designed similarly to machine learning models that implement speech-to-text synthesis. The synthetic machine learning module 204 can train the bottleneck trunk 305 using a large number of samples, e.g., millions of pairs of (audio, text) to output bottleneck features. For example, the bottleneck trunk 305 can be trained similarly to a DeepSpeech system, but instead of being trained to predict text, the bottleneck trunk 305 uses a previous layer that holds embeddings for mapping audio characteristics to reference audio.
[0049] The synthetic machine learning module 204 can train the submodel decoder 310 by fixing the weights of the bottleneck trunk 305 after the bottleneck trunk 305 has been trained, and for training, it can provide the submodel decoder 310 with bottleneck features along with audio features extracted from the input waveform of the audio stream, such as pitch, phase, and rate. In some embodiments, the submodel decoder 310 also receives speaker identifiers used to distinguish different voices in the audio stream.
[0050] Once the bottleneck trunk 305 and the submodel decoder 310 are trained, the bottleneck trunk 305 receives an audio stream generated by the client device 115 as input. In some embodiments, the bottleneck trunk 305 includes a fully connected (FC) layer followed by an LSTM layer used to decompose the audio stream into increasingly abstract representations of the audio data.
[0051] The bottleneck trunk 305 receives the audio stream as an input waveform, along with Mel-frequency cepstrum coefficient (MFCC) features sampled from duplicate samples of the audio stream. The MFCC features are derived from a kind of cepstrum representation of the audio stream to represent information about the audio stream, such as timbre representations within the audio stream. The sampling of the audio stream is parameterized by a time window. For example, a time window of less than 100 milliseconds, which may result in lower accuracy than longer time windows, can be used for computational efficiency and suitability for real-time execution on the client device 115. The client device 115, having less computing power than the server 101, can still achieve accuracy using a shorter time window because its network latency is shorter than the network latency that occurs when the server 101 processes the audio stream.
[0052] The bottleneck trunk 305 encodes the audio and outputs bottleneck features. These bottleneck features are sent to the submodel decoder 310. In some embodiments, the bottleneck trunk 305 outputs bottleneck features for a subset of audio frames, such as one every five frames in the audio stream. In some embodiments, the submodel decoder 310 also receives audio-specific features related to the input waveform, such as the pitch, phase, and rate of the audio stream, as well as a speaker identifier.
[0053] In some embodiments, the submodel decoder 310 includes a fully connected FC layer followed by an LSTM layer. The LSTM layer is used to capture temporal correlations within the audio stream. The submodel decoder 310 outputs MFCC features, which are input as input to the audio waveform generator 315.
[0054] The audio waveform generator 315 receives MFCC features and generates an output waveform that represents the rendered MFCC features as synthesized audio.
[0055] Throughout this process, the machine learning model 300 learns the temporal mapping between positions within the audio stream and synthesizes future frames of the audio stream. For example, if the performance is "Happy Birthday," the machine learning model 300 may map the audio stream to the reference audio by determining that the audio stream contains the first two lines of the song and how to compare it to the reference audio if the audio stream has a different starting point than the reference audio. This is called the time offset in the audio between the audio stream and the reference audio. The machine learning model 300 may also output the rate of the audio stream compared to the reference audio. For example, a user singing "Happy Birthday" might sing at a faster rate than the reference audio. Based on the time offset, the rate of the audio stream, and the future time offset of the song based on the rate of the audio stream compared to the rate of the reference audio, the machine learning model 300 may predict the timing of the next frame in the audio stream and, as a result, output synthesized audio that encapsulates those features.
[0056] In some embodiments, instead of performing the bottleneck extraction described above, the machine learning model 300 may implement a multilingual bottleneck extractor. The multilingual bottleneck extractor is trained to distinguish senons from multiple languages. The output features are language-independent and robust to variations due to language, speech patterns, speaking speed, etc. This machine learning model 300 is trained using supervised learning, and its output includes senon posterior probabilities.
[0057] Exemplary machine learning model on client device 115 In an embodiment where the synthetic machine learning module 204 is stored on the server 101, the mapping machine learning model 204 can be trained using unsupervised learning with a training dataset having unlabeled audio streams.
[0058] Turning to Figure 3B, a block diagram of another exemplary architecture of a machine learning model is shown. The machine learning model 350 includes a vector quantization variational autoencoder (VQ-VAE) 355, a VQ-VAE codebook 360, a pre-model 365, and a VQ-VAE decoder 370.
[0059] The synthetic machine learning module 204 trains VQ-VAE355 and VQ-VAE codebook 360 using a training dataset containing audio stream waveforms. VQ-VAE355 and VQ-VAE codebook 360 compare their output waveforms to the input waveforms to train the corresponding modules of the machine learning model 350. The synthetic machine learning module 204 trains the pre-model 365 on the code vector output from VQ-VAE codebook 360, which is conditioned on additional features such as audio features including pitch, phase, rate, and optionally one or more speaker identifiers.
[0060] The machine learning model 350 is trained to receive an audio stream parameterized by a time window. In some implementations, the audio stream is 300-500 milliseconds long. Longer audio streams are more computationally intensive due to the larger model size, but they also result in more accurate results.
[0061] In some embodiments, the VQ-VAE355 receives an input waveform of an audio stream and converts the input waveform into a latent representation. The VQ-VAE355 uses an autoregressive network structure that includes several one-dimensional (1D) convolutional blocks. The VQ-VAE codebook 360 receives the latent representation as input, and the bottleneck quantizes the latent representation into discrete code vectors using a predefined codebook.
[0062] The pre-model 365 receives discrete code vectors from the VQ-VAE codebook 360, along with the pitch, phase, and rate of the input waveform, and (optionally) one or more speaker identifiers related to the audio stream. Audio features are computed from the input waveform, and the one or more speaker identifiers allow the network to model codes specific to a user or music genre. The pre-model 365 uses a transformation layer to extrapolate the code vectors over time, edit them, and output resampled code vectors. The VQ-VAE decoder 370 is an autoregressive audio waveform generator. The VQ-VAE decoder 370 receives the edited and resampled code vectors and synthesizes an output waveform from the edited and resampled code vectors.
[0063] The mixing module 206 mixes audio streams from multiple client devices 115 to form a combined audio stream that synchronizes one or more synthesized audio streams with the actual audio streams. In some embodiments, the mixing module 206 includes a set of instructions that can be executed by the processor 235 to form the combined audio stream. In some embodiments, the mixing module 206 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0064] In some embodiments, the mixing module 206 receives the output of the synthesis machine learning module 204 and synchronizes the audio streams based on the output. For example, the mixing module 206 may receive a synthesized first audio stream and a second audio stream and generate a combined audio stream. Depending on whether the mixing is performed on the server 101 or the client device 115, the mixing module 206 may introduce different amounts of delay into different audio streams to account for different types of delay and ensure that the audio streams are synchronized. For example, if the mixing is performed on the server 101, the mixing module 206 may account for receiver delay. The various factors for mixing audio streams are discussed in more detail below.
[0065] The mixing module 206 synchronizes the synthesized first audio stream and the second audio stream to form a combined audio stream. The time window of the audio stream may have different lengths depending on the device storing the synthetic machine learning module 204. For example, if the synthetic machine learning module 204 is stored on the client device 115, the time window is less than 100 milliseconds. In another example, if the synthetic machine learning module 204 is stored on the server 101, the time window is 300 to 500 milliseconds.
[0066] The mixed module 206 may perform synchronization on the same device as the device in which the synthetic machine learning module 204 is stored, or on a different device. The device may include a server 101, a client device 115a that transmits a first audio stream, and a client device 115b that receives the first audio stream.
[0067] In the first example, client device 115b performs both audio stream synthesis and audio stream synchronization. Synthesis machine learning module 204 receives globally time-stamped packets from all other client devices 115 via server 101 and synthesizes the audio streams of the other client devices 115. Mixing module 206 generates a local mix of the actual audio stream of client device 115b and the synthesized audio from the other client devices 115.
[0068] In some embodiments, the first example is a preferred example. Several advantages of the first example exist. The time required to predict future portions of the audio stream is reduced, which results in higher quality audio stream and shorter interaction latency. In addition, the client device 115b can recover from poor latency. Users may have to use a client device with considerable computing power to synthesize and synchronize audio locally, but better hardware can provide a better experience. Finally, this architecture offers the advantage that the synthesized audio remains on the client device 115b.
[0069] Some possible drawbacks of the first example are that the computational complexity is O(n) for incoming streams with respect to n performers. Furthermore, all processing is on the client device 115b, and since the processing is critical, it may drain the battery of the client device 115b, the processing may be difficult, it may require different implementation code for different types of client devices 115, and it may only work on a subset of client devices 115 that have sufficient computational power.
[0070] In the second example, server 101 performs both the synthesis and synchronization of audio streams. The advantage of the second example is that all client streams are received and processed at the server, resulting in a lower computational complexity of O(1). In this example, server 101 receives separate audio streams from each client, synthesizes them, and provides a synchronized combined audio stream to each client, so no client-side computation is required to generate the combined audio stream. The infrastructure is all part of server 101 and therefore controlled. Furthermore, the worst-case latency of the audio stream is only the latency between the client and the server, but in client-side processing, it may be between client device 115a and server 101, and then between client device 115b, so the worst-case predicted time is shorter than when client-side processing is used.
[0071] In the third example, client device 115a performs the synthesis of audio streams, and server 101 performs the synchronization of audio streams. Client device 115a synthesizes the audio streams for worst-case latency, and server 101 mixes (n-1) streams for the player in locksteps at the same time by delaying each audio stream as needed. The advantages of the third example include that incoming and outgoing streams are O(1) per client device 115 and processing time at the server is O(1). In addition, better hardware in client device 115a makes the sound of client device 115a better for other client devices 115. For example, better hardware results in better processing, which is essential when audio streams must be processed quickly to maintain near real-time transmission of the audio streams.
[0072] In the fourth example, client device 115a performs audio stream synthesis, and client device 115b performs audio stream synchronization. This may be suitable for peer-to-peer streaming that does not involve a server.
[0073] If no post-processing occurs and the mixing module 206 does not perform synthesis on the client device 115b, the mixing module 206 instructs the I / O interface 239 to send the combined audio stream to the client device 115b that can play the combined audio stream. If no post-processing occurs and the mixing module 206 is on the client device 115b, the mixing module 206 may instruct the I / O interface 239 to provide the combined audio stream to the speaker 243 for playback.
[0074] The post-processing module 208 processes the combined audio stream. In some embodiments, the post-processing module 208 includes a set of instructions that can be executed by the processor 235 to process the combined audio stream. In some embodiments, the post-processing module 208 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0075] In some embodiments, the post-processing module 208 may modify the combined audio stream to match the acoustics of the environment in which the client device 115b is located, by taking into account echoes, background noise, etc. In some embodiments, the post-processing module 208 may perform audio cleaning or noise suppression on the combined audio stream to avoid situations where dissonance may occur when the combined audio stream goes from having background noise to being noise-free, for example, if the first audio stream includes a car honking in the background and the second audio stream is in a completely quiet room.
[0076] The user interface module 210 generates the user interface. In some embodiments, the user interface module 210 includes a set of instructions that can be executed by the processor 235 to generate the user interface. In some embodiments, the user interface module 210 can be stored in the memory 237 of the computing device 200 and made accessible and executable by the processor 235.
[0077] The user interface module 210 generates a user interface for user 125 associated with client device 115. The user interface can be used to initiate audio communication with other users, participate in games or other experiences within the metaverse, send texts to other users, initiate video communication with other users, and so on.
[0078] In some embodiments, before a user participates in the metaverse, the user interface module 214 generates a user interface that includes information about how the user's information will be collected, stored, and analyzed. For example, the user interface may ask the user to provide permission to use any information associated with the user. The user is informed that user information may be deleted by the user and that the user may have the option to choose which types of information are provided for different uses. The use of information is subject to applicable regulations, and the data is stored securely. Data collection is not performed in specific locations and for specific user categories (e.g., based on age or other demographics), data collection is temporary (i.e., the data is discarded after a certain period), and the data is not shared with third parties. Some of the data may be anonymized, aggregated across users, or otherwise modified in a way that makes it impossible to determine the identity of a particular user.
[0079] In some embodiments, the user interface obtains user permission before any audio stream is sent to the server 101 or another client device 115. The user interface may include different levels of granularity for user permissions. For example, the user may specify that synthesized audio can only be generated if it is generated on the client device 115 and not on the server.
[0080] In some embodiments, the mixing module 206 determines that the time difference between the first audio stream and the second audio stream exceeds a threshold time difference and sends a command to the user interface module 210 to generate a user interface. The user interface may provide guidance to the user 125 on how to change the singing speed so that the audio stream can be synchronized with the other audio stream.
[0081] Looking at Figure 4, an exemplary user interface 400 is shown that guides the user to change the timing of the performance. In this example, the user interface 400 includes user guidance regarding performance and a movement indicator 405 that prompts the performer associated with the client device 115b to perform in a way that reduces the time difference between the first audio stream and the second audio stream. For example, the movement indicator 405 may move slower than the user is performing to indicate a decrease in the speed of the user's performance.
[0082] In some embodiments, the user interface module 210 indicates the synchronization of a combined audio stream with the actions of a graphically displayed performer. For example, if the combined audio stream is speech being performed by a performer while they are graphically displayed as an avatar, the user interface module 210 may synchronize the combined audio stream with the avatar's mouth, movements, etc.
[0083] Exemplary Method Figure 5 is an exemplary flowchart 500 illustrating data transmission between a client device 115 and a server 101 according to several embodiments described herein. The flowchart 500 includes a first client device 510, a server 515, and a second client device 520. Thick lines represent network data transmission between the three devices, and thin lines represent data transmission within the first client device 510.
[0084] The first client device 510 receives an audio stream from the microphone 505. The audio stream is combined in the first client device 510, the server 515, or the second client device 520. The combined audio stream is mixed with one or more other audio streams in the server 515 or the second client device 520 to form a combined audio stream. The combined audio stream is transmitted to the speaker 525 for playback in the first client device 510.
[0085] In addition to the above, data is also received by server 515. For example, server 515 receives a first stream from the first client device 510 and a second stream from the second client device 520, and server 515 synthesizes the first audio stream and mixes it with the second audio stream.
[0086] As a result of implementing the metaverse engine 103, the first audio stream of the violin and the second audio stream of the cello on the first client device 510 are heard simultaneously by the relevant users playing the music, as if they were in the same physical room.
[0087] Figure 6 is an exemplary flowchart 600 for synthesizing audio streams for synchronous communication. Task detection 605 is performed on the delayed audio stream to identify a performance identifier and the type of performance. For example, it is determined whether the delayed audio stream contains instruments, speech, singing, etc., as well as what type of song it is and what style of performance it is. Based on the identification, a reference audio is output.
[0088] The delayed audio stream is also received by a phase shift analyzer 610, which identifies the time offset of the delayed audio stream compared to the reference audio, and a rate analyzer 615, which identifies the rate of the delayed audio stream. The time offset and rate of the delayed audio stream are received by a sampler 620, which maps the audio stream to the reference audio. The mapping is received by a deep neural network 635 (or another suitable model) which extrapolates the audio stream into the future to complete the synthesis of the audio stream.
[0089] Figure 7 is an illustrative flowchart for synthesizing an audio stream for synchronous communication using a second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.
[0090] Method 700 may begin in block 702, where a first audio stream of performance associated with the first client device 115 is received. Block 702 may be followed by block 704.
[0091] In block 704, the first and second audio streams are combined to form a combined audio stream that synchronizes the first and second audio streams associated with the second client device during a performance time window shorter than the total performance time. The first and second audio streams are then mixed to form a combined audio stream that synchronizes the first and second audio streams associated with the second client device. The time window is advanced and the generation and mixing are repeated until the performance is complete.
[0092] Figure 7 is an illustrative flowchart for synthesizing an audio stream for synchronous communication using a second client device 115. In this example, the metaverse application 104 is stored on the second client device 115.
[0093] Figure 8 is another exemplary flowchart for synthesizing audio for synchronous communication using a server. In this example, the metaverse engine 103 is stored on server 101.
[0094] Method 800 may begin in block 802, where a first audio stream of performance associated with the first client device 115 is received. Block 802 may be followed by block 804.
[0095] In block 804, the first and second audio streams are combined to form a combined audio stream that synchronizes the combined first audio stream with the second audio stream associated with the second client device during a performance time window shorter than the total performance time. The first and second audio streams are then mixed to form a combined audio stream that synchronizes the combined first audio stream with the second audio stream associated with the second client device, and the time window is advanced and generation and mixing are repeated until the performance is complete. In some embodiments, mixing the combined first audio stream involves introducing a delay into the combined audio stream to account for the latency that occurs by sending the combined audio to the second client device. Block 804 may be followed by block 806.
[0096] In block 806, the combined audio stream is sent to the second client device 115.
[0097] The various embodiments described herein include acquiring data from various sensors in a physical environment, analyzing such data, generating recommendations, and providing a user interface. Data acquisition is performed only with the permission of a specific user and in accordance with applicable regulations. The data is stored in accordance with applicable regulations, including anonymizing the data or otherwise modifying the data to protect the user's privacy. Users are provided with clear information regarding data collection, storage, and use, and are given the option to select the types of data that may be collected, stored, and used. Furthermore, users control the devices on which data may be stored (e.g., client devices only, client devices + server devices, etc.) and the devices on which data analysis is performed (e.g., client devices only, client devices + server devices, etc.). The data is used for the specific purposes described herein. The data is not shared with third parties without the explicit permission of the user.
[0098] The methods, blocks, and / or operations described herein may be executed in an order different from the order illustrated or described, and / or may be executed concurrently (partially or completely) with other blocks or operations as necessary. Some blocks or operations may be executed on a portion of the data and then executed again later on another portion of the data, for example. Not all of the described blocks and operations must be executed in various implementations. In some implementations, blocks and operations may be executed multiple times in a method, in different orders, and / or at different times.
[0099] In the above description, numerous specific details are provided for explanatory purposes and to provide a complete understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In some cases, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands and any peripheral device providing services.
[0100] Any reference in this specification to “some embodiments” or “some examples” means that certain features, structures, or characteristics described in relation to the embodiments or examples may be included in at least one implementation of the description. The phrase “in some embodiments” appearing in various places within this specification does not necessarily refer to the same embodiment.
[0101] Some parts of the detailed explanation above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic explanations and representations are the means used by those skilled in data processing techniques to most effectively communicate the content of their work to others skilled in the art. An algorithm is generally considered here as a self-consistent set of steps leading to a desired result. The steps require the physical manipulation of physical quantities. Usually, but not always, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and other manipulated. For reasons of general use, it has sometimes proven convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.
[0102] However, it should be kept in mind that all of these terms and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. As will be evident from the following discussion, unless otherwise noted, discussions throughout this explanation using terms such as “processing,” “computing,” “calculating,” “decision,” or “display” are understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data expressed as physical (electronic) quantities in the registers and memory of a computer system into other data similarly expressed as physical quantities in the computer system memory or registers, or other such information storage, transmission, or display devices.
[0103] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in the computer. Such computer programs may be stored in non-temporary computer-readable storage media, including, but not limited to, any type of disk including optical disks, ROM, CD-ROM, magnetic disk, RAM, EPROM, EEPROM, magnetic card or optical card, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0104] This specification may take the form of several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including, but not limited to, firmware, resident software, microcode, etc.
[0105] Furthermore, the description may take the form of a computer program product accessible from a computer-usable medium or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, the computer-usable medium or computer-readable medium may be any device that can store, communicate, propagate or transfer a program for use by or in connection with an instruction execution system, device, or drive.
[0106] A data processing system suitable for storing or executing program code includes at least one processor directly or indirectly coupled to a memory element via a system bus. The memory element may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage for at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution. [Explanation of Symbols]
[0107] 100 Network environment, environment 101 Server 103 Metaverse Engine 104 Metaverse Applications 104a Metaverse applications 104b Metaverse Applications 105 Network 115 Client Devices 115a...n Client Device 125 users 125a...n User 199 Databases 200 computing devices 202 Performance Awareness Module 204 Synthetic Machine Learning Modules, Machine Learning Modules, Mapping Machine Learning Models 206 Mixing Module 208 Post-processing module 210 User Interface Module 218 Bus 222 signal line 224 signal line 226 signal line 228 signal line 230 signal line 232 signal line 234 signal line 235 processors 237 memory 239 Input / Output (I / O) Interface, I / O Interface 241 Microphone 243 speakers 245 displays 247 Storage Devices 300 Machine Learning Models 305 Bottleneck Trunk 310 Submodel Decoder 315 Audio Waveform Generator 350 Machine Learning Models 355 Vector Quantization Variational Autoencoder (VQ-VAE), VQ-VAE 360 VQ-VAE Codebook 365 Pre-model 370 VQ-VAE Decoder 400 User Interfaces 405 Moving Indicator 500 Flowcharts 505 Microphone 510 First client device 515 Servers 520 Second client device 525 Speakers 600 Flowchart 605 Task detection 610 Phase Shift Analyzer 615 Rate Analyzer 620 Sample 635 Deep Neural Networks
Claims
1. A method performed by a computer, The steps include receiving a first audio stream of performance associated with a first client device, During the performance time window which is shorter than the total performance time, The steps include generating a synthesized first audio stream that predicts the future performance based on the audio characteristics of the first audio stream, The steps include mixing the synthesized first audio stream and the second audio stream to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream associated with the second client device. Includes, A method in which the time window is advanced and the generating step and the mixing step are repeated until the performance described above is completed.
2. In response to receiving the first audio stream described above, The steps include determining the performance identifier of the performance associated with the first audio stream, The steps include receiving reference audio based on the aforementioned performance identifier, and The method according to claim 1, further comprising:
3. The step of generating the synthesized first audio stream includes the step of determining a time offset between the first audio stream and the reference audio, The method according to claim 2, wherein the time offset occurs when the first audio stream has a different starting point than the reference audio, and the step of generating the synthesized first audio stream is further based on the time offset.
4. The step of generating the synthesized first audio stream includes the step of determining the rate of the first audio stream compared to the rate of the reference audio, The method of claim 2, wherein the step of generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.
5. The method according to claim 1, wherein the audio features of the first audio stream are selected from a group of pitch, rate, phase, or combinations thereof.
6. The method according to claim 1, wherein the audio features of the first audio stream include one or more speaker identifiers detected in the first audio stream.
7. A step of determining whether the time difference between the first audio stream and the second audio stream exceeds a threshold time difference, The steps include generating graphical data for displaying a user interface that includes user guidance regarding the performance and movement indicators that prompt the performer associated with the second client device to perform in a way that reduces the time difference between the first audio stream and the second audio stream; The method according to claim 1, further comprising:
8. The method according to claim 1, further comprising the step of modifying the coupled audio stream to match the acoustics of the environment in which the second client device is located.
9. The step of generating the synthesized first audio stream is: The steps include identifying that a portion of the performance was skipped within the first audio stream, The steps include: synthesizing the first audio stream to correct the portion of the performance that was skipped; The method according to claim 1, including the method described in claim 1.
10. The method according to claim 1, further comprising the step of synchronizing the combined audio streams to match the actions of a graphically displayed performer.
11. It is a device, Processor and The system comprises a memory coupled to the processor, in which instructions are stored, and when an instruction is executed by the processor, the processor: Receiving a first audio stream of performance associated with a first client device, During the performance time window which is shorter than the total performance time, To generate a synthesized first audio stream that predicts the future performance based on the audio characteristics of the first audio stream, The synthesized first audio stream and the second audio stream are mixed to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream associated with the second client device. Perform an action that includes this, The time window is advanced until the aforementioned performance is completed, and the generation and mixing processes are repeated in the device.
12. In response to receiving the aforementioned first audio stream, determine the performance identifier of the performance associated with the aforementioned first audio stream. The device according to claim 11, which receives reference audio based on the performance identifier.
13. Generating the synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, The device according to claim 12, wherein the time offset occurs when the first audio stream has a different starting point than the reference audio, and generating the synthesized first audio stream is further based on the time offset.
14. Generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio, The device according to claim 12, wherein generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.
15. The device according to claim 11, wherein the audio features of the first audio stream are selected from a group of pitch, rate, phase, or combinations thereof.
16. A non-temporary computer-readable medium storing instructions, wherein, when an instruction is executed by one or more computers, the following operations are performed on the one or more computers: Receiving a first audio stream of performance associated with a first client device, During the performance time window which is shorter than the total performance time, A synthesized first audio stream is generated to predict the future performance based on the audio characteristics of the first audio stream. In order to form a combined audio stream that synchronizes the synthesized first audio stream and the second audio stream associated with the second client device, the synthesized first audio stream and the second audio stream are mixed, The time window is advanced until the aforementioned performance is completed, and the generation and mixing are repeated. A non-temporary computer-readable medium that enables the operation of [the process].
17. The operation described above is: In response to receiving the aforementioned first audio stream, determine the performance identifier of the performance associated with the aforementioned first audio stream. Based on the aforementioned performance identifier, receive the reference audio. The computer-readable medium according to claim 16, further comprising the following:
18. Generating the synthesized first audio stream includes determining a time offset between the first audio stream and the reference audio, The computer-readable medium according to claim 17, wherein the time offset occurs when the first audio stream has a different starting point than the reference audio, and generating the synthesized first audio stream is further based on the time offset.
19. Generating the synthesized first audio stream includes determining the rate of the first audio stream compared to the rate of the reference audio, The computer-readable medium according to claim 17, wherein generating the synthesized first audio stream is further based on the rate of the first audio stream compared to the rate of the reference audio.
20. The computer-readable medium according to claim 16, wherein the audio features of the first audio stream are selected from a group of pitch, rate, phase, or combinations thereof.
Citation Information
Patent Citations
Accompaniment method for actively following music signals and related equipment
CN112669798A
Audiovisual synchronously compositing and distributing method, device for player'S terminal, program for the device and recording medium where the program for the device is recorded, service providing device, and program for the service providing device and recording medium recorded with the program for the device
JP2003167575A
Remote multipoint concert system using network
JP2007041320A
Remote music performance system
JP2008089849A
Device and method for distributing data, and program
JP2009005012A