Electronic device and method
By generating spatial audio signals through an artificial neural network system, the problem of traditional audio systems being unable to provide immersive 3D audio is solved, and automatic conversion from text to high-quality 3D audio is achieved, enhancing the realism and immersion of the sound experience.
Patent Information
- Application Number
- CN202480044094.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-07
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-30
AI Technical Summary
Traditional audio systems cannot provide an immersive 3D sound experience, and existing 3D audio technologies have complex generation processes and their effects need improvement.
Using an artificial neural network (ANN) system, an output spatial audio signal corresponding to the input text is generated through spatial trajectories. This includes a generative ANN and an encoder/decoder system. By utilizing a diffusion model and audio rendering technology, the conversion from text to 3D audio is achieved.
It achieves automatic conversion from text to high-quality 3D audio signals, avoiding the manual search and editing process and providing a more realistic 3D sound experience.
Smart Images

Figure CN121444484A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to apparatus and methods for generating audio signals. Background Technology
[0002] Traditional audio systems use stereo or multi-channel speakers to deliver a sound experience. These speakers can provide two-dimensional (stereo) sound, or to some extent, a surround sound experience. However, the realism of the sound experience in such systems is limited, and listeners may not be fully immersed in the sound environment. Recently, 3D audio technology has emerged as a promising solution to this problem, offering a more realistic sound experience that simulates three-dimensional space.
[0003] In recent years, driven by the rise of virtual reality, gaming, and other applications requiring immersive audio experiences, the demand for 3D audio technology has grown rapidly. Unlike traditional audio systems, 3D audio technology can create a sound environment that simulates three-dimensional space, allowing listeners to experience sounds from different directions and distances. Besides VR and gaming, 3D audio technology also has potential applications in other fields, such as audiobooks, music production, and the film industry. For example, in the film industry, 3D audio technology can be used to create more realistic and immersive sound environments, enhancing the audience's viewing experience.
[0004] Current methods for creating 3D audio require the use of specialized microphones and complex software to capture and render sound.
[0005] Although there are existing technologies for generating 3D sound, the generation process and the resulting 3D sound effects still need improvement. Summary of the Invention
[0006] According to a first aspect, this disclosure provides an electronic device, comprising: a circuit configured to generate an output spatial audio signal corresponding to input text based on a spatial trajectory; and a spatial trajectory generated by a second artificial neural network (ANN) system based on the input text.
[0007] According to a second aspect, this disclosure provides a method for generating an output spatial audio signal corresponding to an input text, based on an artificial neural network (ANN) system.
[0008] Further aspects are set forth in the dependent claims, the following description, and the accompanying drawings. Attached Figure Description
[0009] The embodiments are illustrated by way of example with respect to the accompanying drawings, wherein: Figure 1 An exemplary process for text-to-3D sound generation is illustrated schematically; Figure 2An exemplary training process for implementing a diffusion model of a text-to-speech unit is illustrated schematically; Figure 3 An exemplary encoder / decoder ANN is illustrated schematically; Figure 4 An exemplary training process for a text-to-speech unit, including a diffusion generator and an encoder / decoder model, is shown. Figure 5 An exemplary reasoning process for a text-to-speech unit 103, including a diffusion model and a decoder ANN, is shown. Figure 6 An exemplary reasoning process for a text-to-3D trajectory unit 105, including a diffusion model, is shown; Figure 7 An implementation method for 3D audio rendering based on a digital monopole synthesis algorithm is provided; Figure 8 A flowchart illustrating the generation of 3D sound based on input text is shown; and Figure 9 A block diagram illustrating an implementation of an electronic device capable of generating 3D sound based on input text is shown. Detailed Implementation
[0010] In reference Figures 1 to 9 Before providing a detailed description of the implementation methods, some general information will be provided first.
[0011] The embodiment discloses an electronic device including a circuit configured to generate an output spatial audio signal corresponding to an input text based on a spatial trajectory, and the spatial trajectory is generated by a second artificial neural network (ANN) system based on the input text.
[0012] The circuit may include a processor, memory (RAM, ROM, etc.), storage devices, input devices (mouse, keyboard, camera, etc.), output devices (displays (e.g., liquid crystal displays, (organic) light-emitting diode displays, etc.), speakers, etc.), and (wireless) interfaces, etc., a common design in electronic devices (computers, smartphones, etc.). Furthermore, it may include sensors for sensing still images or video image data (image sensors, camera sensors, video sensors, etc.), sensors for sensing fingerprints, and sensors for sensing environmental parameters (e.g., radar, humidity, light, temperature), etc.
[0013] Spatial audio signals can be 3D sound. Spatial audio refers to any type of audio signal designed to create a sense of space and a three-dimensional sound field. This is achieved using specialized audio processing techniques that take the spatial characteristics of sound into account. Several audio rendering techniques exist for creating 3D audio, including binaural audio, immersive sound, surround sound, Dolby Atmos, DTS:X, and object-based audio. These techniques employ various methods to create a sense of sound direction and spatial awareness, allowing listeners to experience audio in a more immersive way. 3D audio is commonly used in applications such as virtual reality, gaming, and audio production.
[0014] Spatial trajectories can be, for example, 3D trajectories. A 3D trajectory can be encoded using a series of 2D or 3D coordinates. In another embodiment, the spatial trajectory may have more than three dimensions.
[0015] Artificial neural networks (ANNs) can include multilayer perceptrons (MLPs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTM) networks, deep belief networks (DBNs), autoencoders, generative adversarial networks (GANs), Hopfield networks, Boltzmann machines, self-organizing maps (SOMs), feedforward neural networks (FNNs), modular neural networks (MNNs), radial basis function networks (RBFNs), cross-transfer neural networks (CPNs), cascaded correlation neural networks (CCNs), and so on.
[0016] The input text may include spatial information. Spatial information in text form may include terms such as "in front of", "behind", "above", "below", "around", "next to", "adjacent to", "near", "far away", "left", "right", "center", "edge", "corner", "inside", "outside", "periphery", "boundary", "surface", "depth", "height", etc.
[0017] Generating output spatial audio signals corresponding to input text using artificial neural network systems can be used, for example, in film post-production. The generation process described in these embodiments allows creators to perform this task with a single input text (e.g., with the aid of text prompts). This avoids manually searching for matching sounds from a sound dataset (a process that requires significant time and effort from technicians) and also avoids the need to edit the sound to generate high-quality 3D sound effects for moving objects.
[0018] In some implementations, the circuitry can be configured to generate a mono audio signal based on the input text via a first ANN system. Mono audio, also known as single-channel audio, refers to sound recorded, processed, or played back using only one channel. In other words, it is an audio type where all sound is mixed into a single channel, typically played through a single speaker or earpiece. A mono audio signal can be an audio signal with no distinction between the left and right channels, and the listener will perceive the sound as a single, uniform sound source. This type of audio cannot provide the spatial or directional information that can be obtained from stereo, surround sound, or 3D audio.
[0019] In some implementations, the circuitry may be further configured to render an output spatial audio signal corresponding to the input text based on a 3D trajectory and a mono audio signal.
[0020] In some implementations, the first ANN system may include a generative ANN for generating mono audio signals. For example, the first ANN system may include a text encoder and / or a generative diffusion model and / or an audio decoder for decompression.
[0021] Probabilistic generative models are a class of machine learning models designed to learn the underlying probability distribution of a dataset and generate new samples from that distribution. These models define a joint probability distribution for the input data and the corresponding output variables, and generate new samples using this distribution by sampling from latent variables. Probabilistic generative models can be used for tasks such as image generation, text generation, anomaly detection, and data augmentation. These models can be trained using maximum likelihood estimation, variational inference, or adversarial training, and typically employ some form of regularization to avoid overfitting and improve generalization. Several types of generative models exist, such as variational autoencoders (VAEs), generative adversarial networks (GANs), autoregressive models, normalized flow models, Boltzmann machines, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), and diffusion models.
[0022] In some implementations, the generative ANN used to generate mono audio signals is trained based on multiple pairs of input text and corresponding mono audio signals.
[0023] In some implementations, the generative ANN used to generate mono audio signals may include a diffusion model. The diffusion model may be a deep generative model that learns the data distribution by slowly adding random noise to the input data and then learning to reverse this process. A diffusion model can be implemented as described in the following paper: “Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion.”, Schneider, Flavio, Zhijing Jin, Bernhard Schölkopf, arXiv preprint, arXiv:2301.11757 (2023). Alternatively, the diffusion model can be implemented as described in the following paper: “Make-an-audio: Text-to-audiogeneration with prompt-enhanced diffusion models,” Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X. and Zhao, Z., 2023, arXiv preprint, arXiv:2301.12661. Alternatively, the diffusion model can be implemented as described in the following paper: “Msanii: HighFidelity Music Synthesis on a Shoestring Budget,” Maina, Kinyug, arXiv preprint, arXiv:2301.06468 (2023).
[0024] In some implementations, the first ANN system may output a compressed audio signal, which is then decompressed to receive a mono audio signal.
[0025] Audio signal compression / decompression can be based on ANNs, such as encoder / decoder systems. Alternatively, compression / decompression can be based on known technologies in the field of audio compression / decompression, such as MP3 (MPEG Audio Layer-3), AAC (Advanced Audio Coding), FLAC (Free Lossless Audio Codec), ALAC (Apple Lossless Audio Codec), WAV (Waveform Audio File Format), AIFF (Audio Exchange File Format), OGG (OggVorbis), Opus, WMA (Windows Media Audio), ATRAC (Adaptive Transform Acoustic Coding), Dolby Digital, DTS (Digital Cinema System), AC-3 (Audio Codec 3), MPEG-4 audio, etc.
[0026] In some implementations, the first ANN system may include a decoder ANN for decompressing the compressed audio signal to receive a mono audio signal.
[0027] In some implementations, the decoder model used to decompress the compressed audio signal is part of an encoder-decoder ANN. An encoder-decoder artificial neural network (ANN) is a neural network architecture commonly used for tasks such as image or speech recognition, natural language processing, and data compression / decompression. For example, an encoder-decoder ANN can be trained, for instance, to compress an audio signal into a smaller representation, or "latent code," which can be transmitted or stored more efficiently, and then decompress it back to its original form. The network is designed to learn the most important features of the audio signal and encode them into a compact representation that can then be decoded by a decoder to reconstruct (i.e., decompress) the original signal with minimal or no quality loss.
[0028] Other ANN types exist that can be used for audio file compression / decompression. For example, feedforward neural networks, convolutional neural networks, and recurrent neural networks have all been used for audio compression and decompression tasks.
[0029] In some implementations, the second ANN system may include a generative ANN for generating spatial trajectories based on input text. The second ANN may, for example, include a text encoder and / or a generative ANN. Several types of generative models exist, such as variational autoencoders (VAEs), generative adversarial networks (GANs), autoregressive models, normalized flow, Boltzmann machines, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), diffusion models, etc.
[0030] In some implementations, generative ANNs used to generate spatial trajectories based on input text include diffusion models.
[0031] In some implementations, the generative ANN used to generate spatial trajectories is trained based on multiple pairs of input text and their corresponding spatial trajectories. This training data can be obtained through manual annotation or extracted from movie clip metadata.
[0032] In some implementations, the spatial trajectory may be encoded as a series of 3D coordinates of a first predetermined length, and wherein the mono audio signal is encoded as a series of audio frames of a second predetermined length. The term audio frame may, for example, refer to a block of audio samples that is processed or analyzed as a single unit. In digital audio, an audio frame typically consists of a fixed number of consecutive audio samples sampled at a constant rate (e.g., 44.1 kHz).
[0033] In some implementations, the first predetermined length of 3D coordinates and the second predetermined length of audio frames may have a predetermined ratio to each other.
[0034] This ratio could be, for example,
[0035] in, It is the number of 3D trajectory elements generated. It is the number of audio frames generated by the diffusion model. It is the time interval between trajectory elements and It is the length of the time frame.
[0036] In some implementations, generating an output spatial audio signal corresponding to the input text may include a (3D) audio rendering process.
[0037] The term (3D) audio rendering process can refer, for example, to the process of generating a spatial audio output from digital audio signals and spatial information (such as spatial trajectories) that simulates a three-dimensional sound field around the listener. This can involve using well-known audio processing techniques such as spatial audio rendering algorithms, equalization (EQ), filtering, spatialization, binaural processing, head-related transfer function (HRTF) processing, surround sound, wavefield synthesis (WFS), distance attenuation, Doppler shift, reverberation, room modeling, directivity, sound field analysis and synthesis (SFAS), waveguide modeling, diffraction modeling, reflection modeling, absorption modeling, convolution processing, psychoacoustic models, etc. Thus, it creates an auditory spatial sense of the listener's surroundings, which includes not only the direction and distance of the sound source but also the spatial characteristics of the listening environment. The resulting 3D audio output signal can be output through dedicated or non-dedicated audio playback systems, such as headphones, surround sound speakers, or other immersive audio systems designed to provide listeners with a more realistic and immersive audio experience.
[0038] In some implementations, the rendering process may include monopole synthesis. The term "audio monopole synthesis" refers to a technique used to generate a sound field in which sound appears to originate from a single point, called a monopole. This is achieved by combining loudspeakers with signal processing algorithms to synthesize a sound field that produces the effect of a monopole sound source.
[0039] In some implementations, the circuit may also be configured to obtain input text via text prompts or sound prompts.
[0040] The embodiments also disclose a method for generating an output spatial audio signal corresponding to input text, based on an artificial neural network (ANN) system. This method may include any aspects described in the embodiments above and below.
[0041] The embodiments also disclose a computer implementation method for generating an output spatial audio signal corresponding to input text, based on an artificial neural network (ANN) system. This computer implementation method may include any aspects described in the embodiments above and below.
[0042] The embodiments also disclose a computer program containing instructions that, when executed by a processor, cause the processor to generate an output spatial audio signal corresponding to the input text based on an artificial neural network (ANN) system. The computer program may include any aspects described in the embodiments above and below.
[0043] The embodiments also disclose a machine-readable medium containing instructions that, when executed by a processor, cause the processor to generate an output spatial audio signal corresponding to the input text based on an artificial neural network (ANN) system. The computer program may include any aspects described in the embodiments above and below.
[0044] Figure 1 An exemplary process for text-to-3D sound generation is schematically illustrated. System 100 generates and outputs 3D audio signals based on text input. Input text 101 “A train is passing from back to front on the left” is input to text-to-sound unit 103 via text prompt 102 (see...). Figure 5 The sound unit 103 generates and outputs a mono audio signal 104 corresponding to the input text 101. Furthermore, the input text 101 is also input to the text-to-3D trajectory unit 105 via text prompts 102 (see...). Figure 6 The text to 3D trajectory unit 105 generates a 3D trajectory 106 based on the input text 101. The mono audio signal 104 and the 3D trajectory 106 are input to the 3D sound rendering unit 106 (see...). Figure 7 The 3D sound rendering unit 106 generates and outputs an output 3D audio signal 107 corresponding to the input text 101. The output sound is sent to an audio output system 108 (e.g., headphones), which plays the 3D sound to the audience.
[0045] Generative models for text-to-sound generation
[0046] The sound generation unit 103 can be implemented, for example, using a probabilistic generative model. A probabilistic generative model is a type of machine learning model designed to learn the underlying probability distribution of a dataset and generate new samples from that distribution. These models define a joint probability distribution for the input data and the corresponding output variables, and generate new samples using this distribution by sampling from latent variables. Probabilistic generative models can be used for tasks such as image generation, text generation, anomaly detection, and data augmentation. These models can be trained using maximum likelihood estimation, variational inference, or adversarial training, and typically employ some form of regularization to avoid overfitting and improve generalization ability.
[0047] There are several types of generative models, such as variational autoencoders (VAE), generative adversarial networks (GAN), autoregressive models, normalized flow, Boltzmann machines, restricted Boltzmann machines (RBM), deep belief networks (DBN), diffusion models, etc.
[0048] The text-to-sound unit 103 can be implemented using a diffusion model. The diffusion model can be implemented as described in the following paper: “Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion.”, Schneider, Flavio, Zhijing Jin, Bernhard Schölkopf, arXiv preprint, arXiv:2301.11757 (2023). Alternatively, the diffusion model can be implemented as described in the following paper: “Make-an-audio: Text-to-audiogeneration with prompt-enhanced diffusion models”, Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X. and Zhao, Z., 2023, arXiv preprint, arXiv:2301.12661. Alternatively, the diffusion model can be implemented as described in the following paper: “Msanii: HighFidelity Music Synthesis on a Shoestring Budget.”, Maina, Kinyug, arXiv preprint, arXiv:2301.06468 (2023).
[0049] Figure 2 This schematically illustrates the implementation of a text-to-sound unit (TSU). Figure 1The following is an exemplary training process for the diffusion model 201 (103 in the example). The diffusion model 201 can be a Long Contextual Latent Diffusion (LCLD) model. The text encoder 203 receives input text 202 via text prompts, such as "A train is passing from back to front on the left-hand side". The text encoder 203 is an ANN. The text encoder 203 converts the input text 202 into a sequence of embeddings, which is represented as a sequence of embedding vectors. These embeddings capture the meaning of the input text 202 with context and are incorporated into the backdiffusion generator at each step. The text is used as conditional input (see below). The sequence of embedding vectors of the text encoder 203 is a representation of the input text 202. Each vector has a predefined size (often called the embedding dimension) and is designed to capture semantic and syntactic information in the text. This vector representation is obtained by processing the input text through a sequence of neural network layers, typically using a recurrent neural network (RNN) or a Transformer-based architecture. For example, this is described in the paper "Attention is all you need" by Ashish Vaswani et al. (published in Advances in Neural Information Processing Systems, Volume 30 (2017)), in which the so-called Transformer decoder, as a standard model architecture, takes the text embeddings as input and outputs a latent representation for the diffusion model. Pre-trained models can be used here. For example, the embedding of the word "cat" can be represented as a vector of length 100 or 300, where each element of the vector represents the meaning of the word or a different feature or aspect of the context. The actual values of the vectors are learned during training and will vary depending on the specific model architecture and the training data.
[0050] The diffusion model 201 further includes a diffusion generator 204. The diffusion generator comprises a forward diffusion generator (upper layer) and a backward diffusion generator (lower layer). Both the forward and backward diffusion generators include several stages of an ANN. (These are different in forward diffusion and backward diffusion). During the training of the ANN, the forward diffusion generator receives input vector 205 as input. Input vector 205 can be a latent space representation of the audio signal (corresponding to input text 202) of the encoder / decoder system (see [link to ANN training]). Figure 3 and Figure 5 (That is, training is based on text and corresponding sound signal pairs.) The input vector 205 is generated by the first-level ANN. The transformation is then performed. The output of the first-level ANN (called the intermediate noise vector) is then forwarded to the second-level ANN. The second-level ANN then transforms it again. This process is repeated until the T-level ANN... The final transformation is performed at level T. During each stage of the forward diffusion, the intermediate noise vector contains more noise. The final intermediate noise vector at level T is a noise vector with a uniform distribution of discrete random variables. This intermediate noise vector is then transformed at level T. (The ANNs used for forward and backward diffusion are typically different.) The input vector at the T-th stage ANN is fed into the backward diffusion generation. Furthermore, the input vector at the T-th stage ANN is conditional on the input text 202 provided by the text encoder 203. This conditionalization is performed by the text encoder network 202, which processes the input text 202 to generate a fixed-length vector, which is concatenated with an intermediate noise vector and fed into the ANN stage. The output of the T-th level ANN is again conditional on the input text 202 and fed into the (T-1)-th level ANN. This process is repeated until the output level. At this point, the output vector 206 is received, and the initial latent representation (i.e., the input vector 205) should be obtained again.
[0051] Input vector 205 and output vector 206 can be encoded into a series of lengths of The audio frames, that is, both vectors have a predefined size.
[0052] Output vector 206 should be consistent with input vector 205, which is the training target of diffusion generator 204. Therefore, the diffusion generator loss function is calculated in the backdiffusion time step t (which can be randomly sampled). The diffusion generator loss function is determined to adjust the weights of the forward and backward diffusion neural networks. The stochastic gradient descent method can be used to minimize the negative log-likelihood of the target sound signal given a noise vector. To do this, the gradient of the negative log-likelihood with respect to the diffusion generator parameters is calculated using the backpropagation algorithm. Then, the parameters of the diffusion generator are updated using this gradient along the direction of minimizing the negative log-likelihood.
[0053] In another implementation, the training process is stabilized using a technique called noise conditionation, in which the diffusion generator is conditional on noise variables added to the intermediate noise vector at each diffusion step. This noise conditionation enables the diffusion generator to learn a smoother, more continuous mapping between the noise vector and the musical notation, and has been shown to improve the quality of the generated music.
[0054] In another implementation, the diffusion generator 204 can also be trained to directly output uncompressed music files without using a decoder. However, this requires a different architecture for the diffusion generator and the use of a different loss function. Decoders are typically used to map the intermediate representations generated by the diffusion process to compressed audio signals, such as MIDI. If the diffusion generator is trained to directly output uncompressed music files, then music in its original audio format needs to be generated, which requires a different approach.
[0055] Overall, employing both forward and backward diffusion during training allows the model to better capture the complex dependencies between the input text, intermediate noise vectors, and corresponding audio signals, resulting in higher-quality music. Once the text encoder and diffusion generator are pre-trained, they are fine-tuned on a smaller dataset of paired text and audio signals. During fine-tuning, a multi-task learning objective is used to jointly train the text encoder and diffusion generator, aiming to minimize the negative log-likelihood of the target audio signal given the input text and noise vectors.
[0056] The diffusion model is trained using multiple pairs of audio signals and input text.
[0057] The text encoder 203 can be trained end-to-end with the diffusion model, or it can be trained alone using text data or paired data (with the aim of finding matching features between text and sound).
[0058] In the field of neural networks, latent representations or latent vectors refer to a set of hidden features or variables learned through the training process of an ANN. These features are not explicitly specified by the inputs or outputs of the ANN, but are inferred from the hidden layers of the network. A latent representation can be viewed as a compressed and refined version of the input data, capturing the key information needed for the network to make accurate predictions or classifications. These representations can then be used for various tasks, such as clustering, dimensionality reduction, and visualization of high-dimensional data. The audio signal used as input to the diffusion generator during training, as well as the audio signal that the diffusion generator might generate during inference, can both be compressed forms and / or latent representations of the audio signal. Compressed latent representations can be generated based on encoder / decoder ANNs, as described, for example, in the paper "High fidelity neural audiocompression" by Défossez, Alexandre et al. (arXiv preprint, arXiv:2210.13438, 2022).
[0059] Figure 3An exemplary encoder / decoder ANN is schematically illustrated. This ANN is based on the paper "High fidelity neural audio compression" by Défossez, Alexandre et al., arXiv preprint, arXiv:2210.13438, 2022. Encoder 302 receives an input audio signal 301, such as a sound signal. Encoder 302 includes an ANN composed of multiple layers of convolutional neural networks (CNNs) and extracts features, for example, by using a short-time Fourier transform (STFT). Encoder 302 aggregates the extracted features into a compressed latent representation that captures key features of the input music signal. Optionally, a quantizer 303 is applied to the latent representation, which performs vector quantization on the latent representation, discretizing it into a finite set of codebook vectors, thereby obtaining a quantized latent representation 304 of the input music signal (i.e., the latent vectors contain discrete values, rather than continuous values). The discrete / quantized latent representation is used to further reduce the bit rate of the latent representation of the input audio signal, thereby reducing its storage and transmission costs. The (quantized) latent representation 304 of the input music signal is then passed to the decoder 305, which takes the compressed latent representation 304 as input and maps it back to the time domain using several layers of transposed convolutional neural networks to generate an audio output signal 306 that optimally reconstructs the input audio signal 301. During encoder / decoder training, the output audio signal 306 is then fed into a discriminator 307 along with the input audio signal 301. The discriminator 307 is a binary classifier that evaluates the perceptual quality of the reconstructed audio output signal 306 relative to the input audio signal 301. The discriminator 307 receives the original and reconstructed audio signals as input and compares them to evaluate whether the reconstructed signal has high perceptual quality. A perceptual loss term in the loss function is used to adjust and train the weights of the encoder 302 and decoder 305 to generate a compressed audio signal with high perceptual quality, thereby reducing the discriminator's ability to distinguish it.
[0060] The neural audio compression system is trained separately using multiple audio data sets.
[0061] The aforementioned diffusion model can be a discrete diffusion model, such as the one described by Austin, Jacob, et al. in “Structureddenoising diffusion models in discrete state-spaces” (published in Advances in Neural Information Processing Systems, Vol. 34 (2021): 17981-17993).
[0062] Figure 4An exemplary training flow for a text-to-sound unit 104, including a diffusion generator and an encoder / decoder model, is shown. Encoder / decoder ANNs 302 / 305 have been trained. Encoder / decoder ANNs 302 / 305 receive an audio input signal 301 and determine a quantized latent representation 304 of the input audio signal. The quantized latent representation 304 of the input audio signal is used as the input vector 205 of the diffusion generator 204. The output vector 206 of the diffusion generator (i.e., the output vector 206 of backdiffusion) is input to the decoder 305, which generates an audio output signal 306 (i.e., a mono audio signal 104). A diffusion generator loss function is computed at the backdiffusion time step t (which may be randomly sampled). This loss function is calculated based on the difference between the output signal (based on the output vector 206) and the input audio signal 301. The diffusion generator loss function determines the adjustment of the diffusion ANN weights, as described above. Figure 2 As stated above.
[0063] In other implementations, other compression techniques such as mp3 can be used instead, or compression can be omitted entirely.
[0064] Figure 5 An exemplary inference process for a text-to-speech unit 103, including a diffusion model and a decoder ANN, is illustrated. The text encoder ANN 203, the diffusion model 204 (which uses only backdiffusion during inference, also known as a denoising model), and the decoder ANN 305 have all been trained. Input text 202 (e.g., “A train is passing from back to front on the left-hand side”) is fed into the text encoder 203, which converts it into an embedding vector, which is used as the conditional input to the backdiffusion model, as described above. Figure 2 The random input vector 401 (e.g., sampled from a Gaussian distribution (e.g., mean 0, variance 1)) is input into the backdiffusion model. The backdiffusion model outputs an output vector 402, which is an encoded audio signal in the latent representation corresponding to the input text 202. This output vector is input to a decoder 305, which outputs an output audio signal 403 corresponding to the input text 202.
[0065] Output vector 402 and output audio signal 403 may have the same frame length. Alternatively, they may have different frame lengths, depending on the level of compression.
[0066] Text to 3D trajectory unit
[0067] The text-to-3D trajectory unit is configured to generate 3D trajectories from input text using generative machine learning. Generative machine learning can be implemented, for example, through the diffusion model described above or other generative models, such as variational autoencoders (VAEs), generative adversarial networks (GANs), autoregressive models, normalized flow, Boltzmann machines, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), diffusion models, etc.
[0068] The text-to-3D trajectory unit 105 can be implemented as a diffusion model, which is similar to the model described above. Figure 2 , Figure 4 , Figure 5 The described text is translated to sound unit 103. However, since 3D trajectories can be represented by vectors with dimensions lower than the sound signal, the above-mentioned... Figure 3 and Figure 5 The encoder / decoder described may not be applicable.
[0069] The trajectory vector may consist of continuous values, and therefore the continuous diffusion process described by Rombach, Robin et al. in their paper “High-resolution image synthesis with latent diffusion models” (published in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022) can be employed.
[0070] The training of the text-to-3D trajectory unit 105 implemented as a diffusion model may be related to the above. Figure 2 Similar to the above. However, in this case, a large amount of training data containing pairs of input text and corresponding 3D trajectories is available. The 3D trajectory can be encoded as a series of lengths... 3D coordinates :
[0071] That is, for training, training sample pairs It is available; the sample pair contains input text. With the corresponding 3D trajectory, the 3D trajectory is encoded as a series of lengths. 3D coordinates The diffusion model learns from training sample pairs. The length of the input received in the middle is 3D coordinates The sequence is forwarded through forward diffusion and then through back diffusion, where the input text... This is then used as a condition to further input into the backdivergence. The output of the backdivergence (i.e., the output 3D trajectory, encoded as a series of lengths...) 3D coordinates ) and a series of input 3D coordinates All of these are fed into the diffusion generator loss function, which determines how to adjust the weights of forward and backward diffusion to minimize the difference between the diffusion generator's input and output.
[0072] Figure 6 An exemplary inference process for a text-to-3D trajectory unit 105, including a diffusion model, is illustrated. The text encoder ANN 602 and the diffusion model 603 (using only backdiffusion during inference) have both been trained. Input text 601 (e.g., “A train is passing from back to front on the left-hand side”) is fed into the text encoder 602, which converts it into an embedding vector, which is used as the conditional input to the backdiffusion model, as described above. Figure 2 The random input vector 604 is input into the backdiffusion model, wherein the random input vector includes a series of lengths... 3D coordinates This is, for example, sampled from a Gaussian distribution (e.g., mean 0, variance 1). The backdiffusion model outputs a 3D trajectory output vector 605, which is encoded as a series of lengths corresponding to the input text 601. 3D coordinates 3D trajectory.
[0073] The 3D trajectory output vector 605 of the text-to-3D trajectory unit is coordinated with the output audio signal 403 of the text-to-sound unit in such a way that an output 3D audio signal 107 can be generated (see the section below on duration considerations).
[0074] In another implementation, the 3D trajectory can be generated directly by the operator.
[0075] In another implementation, the 3D trajectory can be generated by a rule-based system that establishes an association between a set of defined text commands and the output 3D trajectory. In this case, the text may be constrained by a predefined list of text commands that defines the trajectory. The trajectory is then stretched or shortened to fit a specific waveform length.
[0076] In another implementation, the trajectory can be extracted through a process of morphological processing and semantic analysis.
[0077] In another implementation, the machine learning system can be trained using a large number of text examples with corresponding trajectories (e.g., manually labeled or extracted from movie clip metadata). In this case, the trajectories can be generated with explicit indications in the text or without explicit indications. For example, a train typically moves along a straight line at a constant speed, while a dog barking is more likely to move around the "audience."
[0078] In another implementation, this can also be achieved by using the metadata of the generated audio object, which contains the corresponding timeline information, and this information is available in many audio editing systems. For example, such an audio object standard can be found in the MPEG-H standard (see J. Herre, J. Hilpert, A. Kuntz and J. Plogsties, “MPEG-H Audio - The new standard for universal spatial / 3D audio coding”, Journal of the Audio Engineering Society, Vol. 62, pp. 821-830, December 2014).
[0079] Duration consideration
[0080] The aforementioned diffusion model with Transformer denoising model (for text-to-sound generation and for text-to-3D trajectory generation) generates a fixed-length audio signal.
[0081] In the first embodiment, the variable length of different outputs is handled by training two diffusion models (i.e., text-to-sound generation and text-to-3D trajectory generation) based on training data including long sound signals and cropping the sound signals to the desired length (and possibly post-processing them), while downsampling the generated trajectories.
[0082] In another implementation, the variable length of different outputs is handled by training two diffusion models (i.e., text-to-sound generation and text-to-3D trajectory generation), which are based on splicing sound signals (which may require post-processing to smooth the transition) to achieve the desired length, while upsampling the trajectory.
[0083] In another implementation, the variable length of different outputs is handled by training two diffusion models (i.e., text-to-sound generation and text-to-3D trajectory generation), the training being based on training different diffusion models, each of which corresponds to a predefined output length (given by the user).
[0084] In another implementation, the variable length of the different outputs is handled by training two diffusion models (i.e., text-to-speech generation and text-to-3D trajectory generation), based on selecting different types of models for the denoising DNN (i.e., the backdiffusion model) in the diffusion process. For example, a fully convolutional network, such as U-net. The output length is then determined by the length of the input vector (i.e., a random / noise vector set by the user).
[0085] In all the above embodiments, the audio signal is encoded as a series of discrete values. The text-to-speech generator encodes one value per time frame, where the length of the time frame is determined during training. For example, the frame length could be 1 second, 5 seconds, or 10 seconds, etc.
[0086] For a single audio frame, one or more points can be generated as 3D trajectory coordinates. This is to determine the coordinates of the 3D trajectory elements. The number of audio frames, thus making the audio frames The number of 3D trajectory elements For quantity matching, every two trajectory elements must be defined. The time interval between them. That is, the length is Each element of the 3D trajectory sequence corresponds to a specific time interval. This time interval can be 10 ms, 50 ms, 100 ms, or 500 ms, etc. Then, in the 3D trajectory elements... The quantity is determined as follows:
[0087] in, It is the number of trajectory elements generated by the diffusion model (it can be fixed or set by the noise length). It is the number of audio frames generated by the diffusion model (it can be fixed or set by the noise length). It is the time interval between trajectory elements and It is the length of the time frame.
[0088] In the equation above, a 3D trajectory element ("+1") is added to obtain the 3D position of the audio object at the beginning and end of the audio signal. In another implementation, this "+1" can be omitted.
[0089] In another embodiment, to obtain improved synchronized output audio, a sound is selected. and This allows it to remain: .
[0090] 3D sound rendering
[0091] The 3D trajectory output vector corresponding to the input text is 605 (i.e., of length ). 3D coordinates The sequence and output audio signal 403 are rendered by the 3D sound rendering unit 106 to generate an output 3D audio signal 107 corresponding to the input text 101.
[0092] In another implementation, 3D sound rendering can be achieved using a tool such as the 360-walkmix creator (see https: / / 360ra.com / 360-walkmix-creator / ): the generated output audio signal 403 and 3D trajectory output vector 605 are input into the tool as a series of appropriate command sequences to control the automation parameters of the rendering algorithm.
[0093] In another embodiment, an audio workstation (such as Reaper) providing a defined programming interface is used to automatically generate and render the output 3D audio signal 107. This interface allows input of the generated output audio signal 403 and the 3D trajectory output vector 605. For this application, the generated sound and trajectory metadata information can be input or converted.
[0094] In another implementation, the rendering process can be based on monopole synthesis. Audio monopole synthesis refers to a technique used to generate sound fields where sound appears to originate from a single point, called a monopole. This is achieved by combining loudspeakers with signal processing algorithms to synthesize a sound field that produces the monopole source effect.
[0095] Figure 7 An implementation method for 3D audio rendering based on a digital monopole synthesis algorithm is provided. The theoretical background of this technology is described in more detail in patent application US2016 / 0037282A1, which is incorporated herein by reference.
[0096] This technique is implemented in the embodiment of US2016 / 0037282A1, which is conceptually similar to wave field synthesis, which uses a finite number of acoustic enclosures to generate a defined sound field. However, the fundamental basis of the generation principle in these embodiments is specific, because the synthesis does not attempt to accurately model the sound field, but is based on the least squares method.
[0097] The target sound field is modeled as at least one target monopole located at a predetermined target position defined by the 3D trajectory output vector 605. In one embodiment, the target sound field is modeled as a single target monopole. In other embodiments, the target sound field is modeled as multiple target monopoles, each located at a defined (moving) target position defined by the 3D trajectory output vector 605. For example, each target monopole may represent one of a set of noise cancellation sources arranged at a specific location in space. The position of the target monopole can be moved according to the indication of the 3D trajectory output vector 605. For example, the target monopole may adapt to the movement of the noise source to be attenuated. If multiple target monopoles are used to represent the target sound field, the method for synthesizing target monopole sound based on a set of defined synthesized monopoles, as described below, can be applied independently to each target monopole, and the contributions of the synthesized monopoles corresponding to each target monopole can be summed to reconstruct the target sound field.
[0098] Input signal (This could be the output audio signal 403) being fed to... The marked delay unit is then fed to the amplification unit. ,in, This is the index of the corresponding synthesized monopole used to synthesize the target monopole signal. The resulting signal used to synthesize the target monopole signal can be calculated using equation (117) of reference US2016 / 0037282A1, based on the delay and amplification unit of this embodiment. Resulting signal Amplified by power and fed to the speaker .
[0099] Therefore, in this embodiment, the source signal is synthesized. It is performed in the form of delayed and amplified components.
[0100] According to this implementation method, the index is... Delay of synthetic monopole Corresponding to sound in the target monopole With the sound generator Euclidean distance between The propagation time on the surface. Furthermore, according to this embodiment, the amplification factor... With distance Inversely proportional. In an alternative implementation of the system, a correction amplification factor can be used in equation (118) according to reference US2016 / 0037282A1.
[0101] Figure 8A flowchart illustrating the generation of 3D sound based on input text is shown. In step 800, input text is received. In step 801, a mono audio signal is generated based on the received input text. In step 802, a 3D trajectory is generated based on the input text. In step 803, a 3D audio signal corresponding to the input text is rendered based on the mono audio signal and the 3D trajectory.
[0102] Implementation
[0103] Figure 9 A block diagram illustrating an implementation of an electronic device capable of generating 3D sound based on input text is shown. The electronic device 1200 includes a CPU 1201 serving as a processor. The electronic device 1200 also includes a microphone array 1210, a speaker array 1211, and a convolutional neural network unit 1220 connected to the processor 1201. The CNN unit may be, for example, an artificial neural network in hardware, such as a neural network on a GPU or any other hardware specifically designed to implement an artificial neural network. The speaker array 1211 consists of one or more speakers distributed in a predefined space and configured to render 3D audio. The electronic device 1200 also includes a user interface 1212 connected to the processor 1201. This user interface 1212 acts as a human-machine interface and enables dialogue between an administrator and the electronic system. The user interface 1212 may be a graphical user interface (GUI). Additionally, the administrator can use the user interface 1212 to configure the system. The electronic device 1200 also includes a Bluetooth interface 1204 and a WLAN interface 1205. These units 1204 and 1205 serve as I / O interfaces for data communication with external devices. For example, additional speakers, microphones, and video cameras with Ethernet, WLAN, or Bluetooth connectivity can be coupled to processor 1201 via these interfaces 1204 and 1205. Electronic system 1200 also includes a data storage device 1202 and a data memory 1203 (here, RAM). Data memory 1203 is arranged for temporary storage or caching of data or computer instructions processed by processor 1201. Data storage device 1202 is arranged for long-term storage, for example, for recording sensor data acquired from or retrieved from microphone array 1210 and provided to or from CNN unit 1220. Data storage device 1202 may also store audio data representing audio messages, which a notification system may transmit to people moving within a predefined space.
[0104]
[0105] It should be noted that the above description is only an exemplary configuration. Alternative configurations can be implemented using additional or other sensors, storage devices, interfaces, etc.
[0106] It should be further noted that, alternatively, the electronic device 1200 may be implemented using a digital signal processor (DSP) or a graphics processing unit (GPU), without limiting the invention in this respect.
[0107] It should also be noted that... Figure 9 The division of electronic devices into units is for illustrative purposes only, and this disclosure is not limited to any specific division of functions within a particular unit. For example, at least a portion of the circuit may be implemented by a separately programmed processor, a field-programmable gate array (FPGA), a dedicated circuit, etc.
[0108] It should be recognized that the embodiments describe a method having an exemplary order of method steps. However, the specific order of method steps is given for illustrative purposes only and should not be construed as a constraint.
[0109] Unless otherwise stated, all units and entities described in this specification and claimed in the appended claims may be implemented as integrated circuit logic, for example, on a chip, and unless otherwise stated, the functionality provided by these units and entities may be implemented by software.
[0110] With regard to the implementation of the embodiments of the present disclosure at least in part using a software-controlled data processing apparatus, it should be understood that providing such a software-controlled computer program and the transmission, storage or other medium through which such a computer program is provided are conceived as aspects of the present disclosure.
[0111] It should be noted that this technology can also be configured as follows.
[0112] (1) An electronic device, including a circuit configured to: The system generates an output spatial audio signal corresponding to the input text based on a spatial trajectory; and generates a spatial trajectory based on the input text using a second artificial neural network (ANN) system.
[0113] (2) According to the electronic device described in (1), the circuit is further configured to generate a mono audio signal (104) based on input text (101; 202; 601) through a first ANN system (203; 204; 302; 305).
[0114] (3) According to the electronic device described in (1) or (2), the circuit is further configured to render an output spatial audio signal corresponding to the input text (101; 202; 601) based on the 3D trajectory (106; 605) and the mono audio signal (104).
[0115] (4) The electronic device according to (2) or (3), wherein the first ANN system (203; 204; 302; 305) includes a generative ANN (204) for generating a mono audio signal (104).
[0116] (5) The electronic device according to (4), wherein the generative ANN (204) for generating mono audio signals is trained based on multiple pairs of input text and corresponding mono audio signals (104).
[0117] (6) The electronic device according to (4) or (5), wherein the generative ANN (204) for generating a mono audio signal (104) includes a diffusion model (204).
[0118] (7) An electronic device according to any one of (2) to (6), wherein the first ANN system (203; 204; 302; 305) outputs a compressed audio signal (206), which is decompressed to receive a mono audio signal (104).
[0119] (8) An electronic device according to any one of (2) to (7), wherein the first ANN system (203; 204; 302; 305) includes a decoder ANN (305) for decompressing a compressed audio signal (402) to receive a mono audio signal (104).
[0120] (9) The electronic device according to (8), wherein the decoder model (305) for decompressing the compressed audio signal is part of an encoder-decoder ANN (302; 305).
[0121] (10) An electronic device according to any one of (1) to (9), wherein the second ANN system (602; 603) includes a generative ANN (603) for generating a spatial trajectory (106; 605) based on input text (101; 202; 601).
[0122] (11) The electronic device according to (10), wherein the generative ANN (603) for generating spatial trajectories (106; 605) based on input text (101; 202; 601) includes a diffusion model (603).
[0123] (12) The electronic device according to (10) or (11), wherein the generative ANN for generating spatial trajectories (106; 605) is trained based on multiple pairs of input texts (101; 202; 601) and the corresponding spatial trajectories (106; 605).
[0124] (13) An electronic device according to any one of (2) to (12), wherein the spatial trajectory (106; 605) is encoded as a series of 3D coordinates (605) of a first predetermined length, wherein the mono audio signal (104) is encoded as a series of audio frames of a second predetermined length.
[0125] (14) The electronic device according to (13), wherein the first predetermined length of 3D coordinates (605) and the second predetermined length of audio frames have a predetermined ratio to each other.
[0126] (15) The electronic device according to any one of (1) to (14), wherein generating an output spatial audio signal (107) corresponding to the input text (101; 202; 601) includes a rendering process.
[0127] (16) The electronic device according to (15), wherein the rendering process includes monopole synthesis.
[0128] (17) An electronic device according to any one of (1) to (16), wherein the circuit is configured to obtain input text (101; 202; 601) via text prompt (102) or sound prompt (102).
[0129] (18) A method comprising the following steps: Generate an output spatial audio signal corresponding to the input text (101; 202; 601) based on the spatial trajectory (106; 605); and Spatial trajectories (106; 605) are generated based on the input text (101; 202; 601) by a second artificial neural network (ANN) system (602; 603).
Claims
1. An electronic device comprising circuitry configured for: generating, based on a spatial trajectory (106; 605), an output spatial audio signal corresponding to an input text; and generating, by a second artificial neural network system, the spatial trajectory based on the input text.
2. The electronic device of claim 1, the circuitry being further configured for generating, by a first artificial neural network system, a mono audio signal based on the input text.
3. The electronic device of claim 2, the circuitry being further configured for rendering the output spatial audio signal corresponding to an input text based on a 3D trajectory and a mono audio signal.
4. The electronic device of claim 2, wherein, the first artificial neural network system comprises a generative artificial neural network for generating the mono audio signal.
5. The electronic device of claim 4, wherein, the generative artificial neural network for generating the mono audio signal is trained based on pairs of input texts and corresponding mono audio signals.
6. The electronic device of claim 4, wherein, the generative artificial neural network for generating the mono audio signal comprises a diffusion model.
7. The electronic device of claim 2, wherein, the first artificial neural network system outputs a compressed audio signal, which is decompressed to receive the mono audio signal.
8. The electronic device of claim 2, wherein, the first artificial neural network system comprises a decoder artificial neural network for decompressing a compressed audio signal to receive the mono audio signal.
9. The electronic device of claim 8, wherein, the decoder model for decompressing a compressed audio signal is part of an encoder-decoder artificial neural network. 10.The electronic device of claim 1, wherein, the second artificial neural network system comprises a generative artificial neural network for generating the spatial trajectory based on the input text.
11. The electronic device of claim 10, wherein, the generative artificial neural network for generating the spatial trajectory based on the input text comprises a diffusion model.
12. The electronic device of claim 10, wherein, the generative artificial neural network for generating the spatial trajectory is trained based on pairs of input texts and corresponding spatial trajectories.
13. The electronic device of claim 2, wherein, the spatial trajectory is encoded as a series of 3D coordinates of a first predetermined length, and wherein the mono audio signal is encoded as a series of audio frames of a second predetermined length.
14. The electronic device of claim 13, wherein, the 3D coordinates of the first predetermined length and the audio frames of the second predetermined length have a predetermined ratio to each other.
15. The electronic device of claim 1, wherein, generating an output spatial audio signal corresponding to an input text comprises a rendering process.
16. The electronic device of claim 15, wherein, the rendering process comprises a monopole synthesis.
17. The electronic device of claim 1, wherein, the circuitry is configured for obtaining the input text by a textual cue or a vocal cue.
18. A method comprising the steps of: generating, based on a spatial trajectory (106; 605), an output spatial audio signal corresponding to an input text; and generating, by a second artificial neural network system, the spatial trajectory based on the input text.
Citation Information
Patent Citations
Method, device and system
US20160037282A1