Deep learning solution for virtual rotation of binaural audio signals

A deep learning-based neural network adjusts binaural audio to simulate stationary sound fields during head rotations, addressing the limitations of fixed-head recordings and HRTF methods.

DE102025100313A1Pending Publication Date: 2025-08-07INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025100313
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2025-01-07
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Binaural recordings generate sound that rotates with the listener's head, failing to adapt to head rotations, and existing methods like HRTF require costly and time-consuming re-recording or sophisticated microphone arrays.

Method used

A deep learning-based audio regression method using a neural network processes binaural audio signals and a rotation angle to generate new signals that conform to head rotations, effectively replicating a stationary sound field.

Benefits of technology

Enables dynamic adjustment of binaural audio to match head movements without the need for re-recording or complex setups, enhancing immersive listening experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Techniques are provided here for providing binaural audio signals that are virtually rotated to correspond to head rotation, so that the audio output to headphones is perceived to maintain its location relative to the user when a user rotates their head. Specifically, techniques are presented for extracting spherical location information already embedded in binaural signals to generate binaural audio signals that change to correspond to head rotation. A deep learning-based audio regression method can take a two-channel binaural audio signal and a selected rotation angle as input and generate a new binaural audio output signal with the rotated environment corresponding to the rotation angle.The deep learning-based audio regression method can be implemented as a neural network and can include deep learning operations such as convolution, pooling, element-wise operations, linear operations, and nonlinear operations. A deep learning operation can be performed on internal parameters of the DNN and one or more activations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical area

[0001] This disclosure relates generally to binaural audio signals and more particularly to virtual rotation of binaural audio signals. background

[0002] Binaural sound is audio heard by both ears of a listener. In a natural environment, sound waves arrive at each of a listener's ears from a direction other than the center of the listener's head at slightly different times and at slightly different volumes. By using the difference in the arrival time of the sound waves and the difference in the volume of the sound waves, a listener's brain can calculate the origin of the sound. The perception of a sound originating from a selected location relative to a listener's head can be replicated by using binaural recordings and producing a binaural sound with a slightly different output to that speaker of a stereo headset (or earphone). When binaural sound is played through headphones, the listener can hear the binaural sound as three-dimensional (3D).However, binaural recordings used to generate binaural sound are recorded and / or generated with a fixed head orientation. A fixed head orientation for binaural sounds leads to a phenomenon where, when the listener turns their head, the binaural sound rotates along with the head and does not automatically change with the head rotation. Short description of the drawings

[0003] Embodiments will be readily understood from the following detailed description taken in conjunction with the accompanying drawings. To facilitate this description, like reference numerals indicate like structural elements. Embodiments are shown in the figures of the accompanying drawings by way of example and not by way of limitation. Figure 1 is a block diagram of an example of the sound localization process using binaural sound perception. Fig. Figure 2 is a block diagram of a system for generating binaural sound with a selected perceived location-source angle using HRTF. Fig. Figure 3A shows four fixed noise sources around a person's head according to various embodiments. Fig. Figure 3B shows rotation of the four fixed sound sources around a person's head according to various embodiments. Fig. 4 is an illustration of an exemplary virtual rotation system according to various embodiments. Fig. 5 is a block diagram of a deep neural network (DNN) module that may be used for a virtual rotation system, according to various embodiments. Fig. 6 is a block diagram of an example neural network architecture 600 that can perform virtual rotation of input sound locations, according to various embodiments. Fig. 7 is a flowchart of an exemplary method for virtual rotation according to various embodiments. Fig. 8 is a block diagram of an example computing device according to various embodiments. Detailed descriptionOverview

[0004] Listeners can generally determine where a sound in the environment is coming from because sounds arrive at a listener's ears differently depending on where the sound is located in relation to the head. For example, a sound coming from directly on the right side of a person's head will reach the right ear clearly and distinctly, while the left ear will receive a version of the sound that is somewhat attenuated by the presence of the head and body in the sound's acoustic path. The difference in sound at each ear is used by the human brain to estimate sound location. However, it should be noted that even for very short and impulsive sounds (like clapping or a gunshot), a person will not notice any delay or difference between the sounds at each ear; the process is fast and organic.

[0005] Binaural audio can be recorded using specific recording techniques that attempt to emulate how audio is heard from an actual human head. The recording techniques may include the use of a special mannequin head and the use of binaural head-worn microphones placed in a person's ears for the duration of the recording. In some examples, rather than recording binaural audio directly, binaural audio may be created using special digital filters based on a head-relative transfer function (HRTF). An HRTF filter is based on a selected direction from which the sound source is to be perceived, with different filter parameters for each selected direction.

[0006] The perception of a sound originating from a selected location relative to a listener's head can be replicated by using binaural recordings and generating a binaural sound with a slightly different output to that speaker of a stereo headset (or earphone). Alternatively, the perception of a sound originating from a selected location relative to a listener's head can be replicated by continuously filtering audio recordings using an HRTF filter to produce binaural audio that evokes the selected sound directionality. When binaural sound is played through headphones, the listener can hear the binaural sound as three-dimensional (3D).

[0007] Generally, binaural recordings used to generate binaural sound are created with a fixed head orientation, so that when the listener turns their head, the binaural sound rotates along with the head and does not automatically change with head rotation. However, head rotation and the associated change in the sound field are an important listening and immersion characteristic.

[0008] To implement a change in binaural sound corresponding to head rotation, binaural audio recordings are generally re-obtained with different head positions or recordings are made using sophisticated and expensive directional microphone arrays. However, these expensive and time-consuming strategies may not be available for sounds that cannot be re-recorded. Another method that can be used to implement a change in binaural sound corresponding to head rotations is to filter the audio sources with head-related transfer functions (HRTFs) designed to emulate the binaural experience. However, using HRTFs implies filtering the audio signals and adding multiple audio channels to generate a corresponding HRTF for each source location. Additionally, HRTFs require multiple microphone recording arrays and a known source location for each signal.Thus, a simple binaural signal cannot be used to generate a binaural signal rotation using HRTF.

[0009] Systems and methods for providing binaural audio signals that change to coincide with head rotation are presented here. Specifically, techniques are presented for extracting spherical spatial information already embedded in binaural signals to generate binaural audio signals that change to coincide with head rotation. A deep learning-based audio regression method is discussed here that can take a two-channel binaural audio signal and a selected rotation angle as input and generate a new binaural audio output signal with the rotated environment corresponding to the selected rotation angle.

[0010] According to various implementations, a virtual audio rotation module may include a deep learning-based audio regression method, which may be implemented as a neural network, such as a DNN (deep neural network). As described here, a DNN layer may include one or more deep learning operations, such as convolution, pooling, element-wise operation, linear operation, nonlinear operation, and so on. A deep learning operation in a DNN may be performed on one or more internal parameters of the DNN (e.g., weights) determined during the training phase and one or more activations. An activation may become a data point (also referred to as "data items" or "items"). Activations or weights of a DNN layer may be elements of a tensor of the DNN layer. A tensor is a data structure with multiple elements across one or more dimensions.Example tensors are a vector, which is a one-dimensional tensor, and a matrix, which is a two-dimensional tensor. There can also be three-dimensional tensors and even higher-dimensional tensors. A DNN layer may include an input tensor (generated during feature extraction) that includes one or more input activations (also called "input elements") and a weight tensor that has one or more weights. A weight is an element in the weight tensor. A weight tensor of a convolution can be a kernel, a filter, or a group of filters. The output data of the DNN layer may be an output tensor that includes one or more output activations (also called "output elements").

[0011] For purposes of explanation, specific numbers, materials, and configurations are set forth to provide a thorough understanding of the example implementations. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without the specific details and / or that the present disclosure may be practiced with only some of the described aspects. In other instances, well-known features are omitted or simplified to avoid obscuring the example implementations.

[0012] Further reference is made to the accompanying drawings, which form a part hereof, and in which is shown by way of illustration embodiments which may be practiced. It should be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.

[0013] Various operations may be described as multiple discrete actions or operations in sequence in a manner most helpful for understanding the claimed subject matter. However, the order of description should not be construed to imply that these operations are necessarily order-dependent. In particular, these operations need not be performed in the described order. Described operations may be performed in a different order than that of the described embodiment. In additional embodiments, various additional operations may be performed, or described operations may be omitted.

[0014] For the purposes of this disclosure, the phrase "A and / or B" or the phrase "A or B" means (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B and / or C" or the phrase "A, B, or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). The term "between," when used with respect to measurement ranges, includes the ends of the measurement ranges.

[0015] The description uses the phrases "in one embodiment" or "in embodiments," each of which may refer to one or more of the same or different embodiments. The terms "comprising," "having," "with," and the like, as used with respect to embodiments of the present disclosure, are synonymous. The disclosure may use perspective-based descriptions such as "above," "below," "top," "bottom," and "side" to explain various features of the drawings, but these terms are provided for convenience of discussion only and do not imply any desired or required orientation. The accompanying drawings are not necessarily drawn to scale. Unless otherwise indicated, the use of the ordinal adjectives "first," "second," and "third," etc., indicates a preferred embodiment.to describe a common object merely indicates that reference is made to different instances of the same objects and is not intended to imply that the objects so described must be in a given order, whether temporal, spatial, ranked or in any other way.

[0016] In the following detailed description, various aspects of the example implementations are described using terms commonly used by those skilled in the art to convey the nature of their work to others skilled in the art.

[0017] The terms "substantially," "near," "approximately," "close to," and "about" generally refer to being within + / - 20% of a target value based on the input operand of a particular value, as described herein or as known in the art. Likewise, terms indicating an orientation of various elements, e.g., "coplanar," "perpendicular," "orthogonal," "parallel," or any other angle between the elements, generally refer to being within + / - 5-20% of a target value based on the input operand of a particular value, as described herein or as known in the art.

[0018] Additionally, the terms "comprise," "comprising," "include," "including," "having," "with," or any other variation thereof are intended to cover non-exclusive inclusion. For example, a method, process, apparatus, or system comprising a list of elements is not necessarily limited to those elements, but may also include other elements not expressly listed or inherent in such method, process, apparatus, or system. Also, the term "or" refers to an inclusive "or" and not an exclusive "or."

[0019] The systems, methods, and devices of this disclosure each have several novel aspects, none of which alone accounts for all of the desirable attributes disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the following description and the accompanying drawings. Example binaural sound location

[0020] Figure 1 is a block diagram of an example of the sound localization process using binaural sound perception. Specifically, as shown in the leftmost panel, a sound originates from a selected location (φ,θ) 104 around the head 102, where φ is the angle from the central axis in the horizontal plane and θ is the angle from the central axis in the vertical plane, with the central axis just in front of the face, directly in front of the head. As shown in the middle panel, the sound from location 104 reaches each ear via a different acoustic path. Specifically, the sound reaches the left ear via a direct path 106 and the right ear via one or more indirect paths 108a, 108b. The acoustic paths vary depending on head size, head shape, pinna, and environment.The right panel shows that the brain uses the difference between the sound received at each ear to estimate the sound position (φ,θ) 104.

[0021] Using the diagram of Fig. 1, binaural sound with a selected perceived location source angle can be generated by selecting digital filters known as HRTF to filter a monaural sound using selected sound source position angles and sending the output signals to a stereo headphone. Fig. Figure 2 is a block diagram of a system for generating binaural sound with a selected perceived source location angle using HRTF. Virtual sound source location angles 202 indicate the desired perceived position of the sound source. In block 204, an HRTF is selected for each ear based on the virtual sound source location angles 202. The original monaural signal 206, along with the selected HRTF from block 204, is input to a processing block. The original monaural signal 206 is processed with the right HRTF 208a to generate a right-ear signal to be played back in the right headphone speaker 210a, and the original monaural signal 206 is processed with the left HRTF 208b to generate a left-ear signal to be played back in the left headphone speaker 210b.If a listener simultaneously listens to the signal in the right headphone speaker 210a and the left headphone speaker 201b, the listener will perceive the signal as originating from the noise source designated by the virtual noise source position angles 202. Example virtual rotation of binaural audio signals

[0022] Fig. Figure 3A shows four fixed noise sources around a person's head according to various embodiments. In particular, as shown in Fig. 3A shows the sound sources (a telephone, an airplane, a car, and a cat) fixed on the example sphere. As the person rotates their head, the location of the sources relative to the person's ears changes. Thus, the person can perceive an approximate location in their environment for each fixed source. When a person rotates their head by a selected angle, the spherical angular position of each sound source changes by an equivalent angle in the opposite direction of the head rotation.

[0023] Fig. Figure 3B shows rotation of the four fixed sound sources around a person's head according to various embodiments. In particular, Fig. Figure 3B shows a virtual environment in which the person's head remains fixed and instead the source locations move relative to the head, changing the approximate position of the source location relative to the head in the same way as if the head were rotated. Thus, a counterclockwise rotation of a person's head by a selected angle is equivalent to a clockwise rotation of each sound source around that person's head by the same selected angle. If a listener is wearing headphones (or earbuds) and the position of the user's head relative to the headphone speakers is fixed, the equivalence can be used to change the perceived sound source positions.In particular, using equivalence (a counterclockwise rotation of a person's head with the clockwise rotation of each sound source), when a user wears headphones playing binaural signals to create a binaural sound environment, the perceived sound source positions can move with the person's head movement by rotating the perceived audio source positions in the binaural signal by an equal but opposite angle from the head movement, as shown in . Fig. 3B shown. Example overview of the virtual rotation system

[0024] Fig. 4 is an illustration of an exemplary virtual rotation system 400 according to various embodiments. In particular, as shown in Fig. 4, the system 400 includes a neural network 406 that receives a binaural audio signal 402 and a rotation angle 404. The neural network 406 may be a regression neural network that uses the binaural audio signal 402 and the rotation angle 404 as input and generates a new binaural audio output signal 408 with the rotated environment indicated by the rotation angle 404. According to various implementations, the neural network 406 may be a deep neural network, as described in more detail below. Example DNN system

[0025] Fig. 5 is a block diagram of a DNN (Deep Neural Network) module 501 that may be used for a virtual rotation system, according to various embodiments. In the embodiments of Fig. 5, the DNN module 501 includes an interface module 511, a training module 521, a validation module 531, a convolution module 541, and a data store 551. In other embodiments, alternative configurations, other, or additional components may be included in the DNN module 501. Furthermore, functionality associated with a component of the DNN module 501 may be achieved by another component included in the DNN module 501 or another module or system, such as any of the neural networks and / or deep learning systems described herein.

[0026] The interface module 511 enables communication of the DNN module 501 with other modules or systems. For example, the interface module 511 establishes communication between the DNN module 501 and an external database to receive data that can be used to train the DNN or that can be input into the DNN to perform tasks. As another example, the interface module 511 supports the DNN module 501 in distributing DNN to other systems, e.g., data processing devices configured to apply DNN to perform tasks.

[0027] The training module 521 trains the DNN using a training dataset. In some examples, the training dataset may be generated using synthetic audio samples using HRTF. Multiple datasets may be used to provide a variety of audio source types (e.g., human speech, instruments, animals, ambient noise, etc.). For each sample in a dataset, a number of random audio sources may be selected, and each selected audio source may be assigned a random angular direction. Each of the selected audio sources may be played from a respective direction based on the assigned random angle using the corresponding HRTF, and audio sources may be blended to generate training input data.To generate the training output, a second random direction may be determined to represent the new head orientation, and based on the new head orientation and the input direction of each of the audio sources, new relative audio source directions with respect to the head may be determined for each of the audio sources.

[0028] In one embodiment where training module 521 trains a DNN to generate binaural signals with selected perceived source locations, the training dataset includes binaural training signals and training labels. The training labels describe ground truth locations of sound sources in training signals. In some embodiments, each label in the training dataset corresponds to an angle (relative to a centerline and plane) in a training stereo sound space. In some embodiments, each label in the training dataset corresponds to a location in a training stereo sound space. In some embodiments, a portion of the training dataset may be used to initially train the DNN, and the remainder of the training dataset may be retained as a validation subset used by validation module 531 to validate the performance of a trained DNN.The part of the training dataset that does not contain the voting subset and the validation subset can be used to train the DNN.

[0029] The training module 521 also determines hyperparameters for training the DNN. Hyperparameters are variables that specify the DNN training process. Hyperparameters differ from parameters within the DNN (e.g., filter weights). In some embodiments, hyperparameters include variables that determine the architecture of the DNN, such as the number of hidden layers, etc. Hyperparameters also include variables that determine how the DNN is trained, such as batch size, number of epochs, etc. A batch size defines the number of training samples to be processed before the DNN parameters are updated. The batch size is equal to or less than the number of samples in the training dataset. The training dataset can be divided into one or more batches. The number of epochs defines how many times the entire training dataset is run forward and backward through the entire network.The number of epochs defines the number of times the deep learning algorithm processes the entire training dataset. An epoch means that each training sample in the training dataset has had one opportunity to update the parameters within the DNN. An epoch can consist of one or more batches. The number of epochs can be 3, 30, 300, 500, or even more.

[0030] The training module 521 defines the architecture of the DNN, e.g., based on some of the hyperparameters. The architecture of the DNN includes an input layer, an output layer, and a plurality of hidden layers. The input layer of a DNN may include tensors (e.g., a multidimensional array) that specify attributes of the input signal, such as frequency, loudness, and other spectral properties. The output layer includes labels of angles and / or locations of noise sources in the input layer. The hidden layers are layers between the input layer and the output layer. The hidden layers include one or more convolutional layers and one or more other types of layers, such as pooling layers, fully connected layers, normalization layers, softmax or logistic layers, and so on. The convolutional layers of the DNN abstract the input signals to perform feature extraction.In some examples, feature extraction is based on a spectrogram of an input audio signal. A pooling layer is used to reduce the volume of the input signal after convolution. It is used between two convolutional layers. A fully connected layer includes weights, biases, and neurons. It connects neurons in one layer to neurons in another layer. It is used to classify signals between different categories through training. It should be noted that training a DNN is different from using the DNN in real time, and when using a DNN to process data received in real time, latency can become an issue that is not present during training, where the dataset can be pre-loaded.

[0031] During the process of defining the DNN architecture, the training module 521 also adds an activation function to a hidden layer or the output layer. A layer's activation function transforms the weighted sum of the layer's input into a layer's output. The activation function can be, for example, a rectified linear unit activation function, a tangent activation function, or other types of activation functions.

[0032] After the training module 521 defines the architecture of the DNN, the training module 521 inputs a training dataset to the DNN. The training dataset includes multiple training samples. An example of a training sample would be a source position of a feature in an audio sample and a ground truth position of the feature. The training module 521 modifies the parameters within the DNN ("internal parameters of the DNN") to minimize the error between labels of the training objects generated by the DNN and the ground truth labels of the objects. The internal parameters include weights of filters in the convolutional layers of the DNN. In some embodiments, the training module 521 uses a cost function to minimize the error.

[0033] The training module 521 may train the DNN for a predetermined number of epochs. The number of epochs is a hyperparameter that defines the number of times the deep learning algorithm will process the entire training dataset. An epoch means that each sample in the training dataset has had an opportunity to update the DNN's internal parameters. After the training module 521 has completed the predetermined number of epochs, the training module 521 may stop updating the parameters in the DNN. The DNN with the updated parameters is referred to as the trained DNN.

[0034] The validation module 531 verifies the accuracy of trained or compressed DNNs. In some embodiments, the validation module 531 inputs samples in a validation dataset to a trained DNN and uses the outputs of the DNN to determine the model accuracy. In some embodiments, a validation dataset may be formed from some or all of the samples in the training dataset. Additionally or alternatively, the validation dataset includes additional samples besides those in the training sets. In some embodiments, the validation module 531 may determine an accuracy score that measures precision, recall, or a combination of precision and recall of the DNN. The validation module 531 may use the following metrics to determine the accuracy score: Precision = TP / (TP + FP) and Recall = TP / (TP + FN), where Precision may be how many of the total number it predicted (TP + FP or FN, respectively) are correct.false positives), the reference classification model correctly predicted (TP or true positives), and recall can be how many of the total number of objects that exhibited the property in question (TP + FN or false negatives) the reference classification model correctly predicted (TP). The F-score (F-score = 2 * PR / (P + R)) unifies precision and recall into a single measure.

[0035] The validation module 531 may compare the accuracy score to a threshold score. In one example, where the validation module 531 determines that the accuracy score of the augmented model is less than the threshold score, the validation module 531 instructs the training module 521 to retrain the DNN. In one embodiment, the training module 521 may iteratively retrain the DNN until a stopping condition occurs, such as an accuracy measure indicating that the DNN is sufficiently accurate or that a number of training rounds have occurred.

[0036] The convolution module 541 performs real-time data processing, such as speech enhancement, dynamic noise reduction, blind source separation, and / or self-noise cancellation. In the embodiments of Fig. 5, the convolution module 541 includes a time-domain encoder 543, a frequency-domain encoder 545, and a time-domain decoder 547. In some examples, the time-domain encoder 543 is a convolutional time-domain encoder, the frequency-domain encoder 545 is a convolutional frequency-domain spectrum encoder, and the time-domain decoder 547 is a convolutional time-domain decoder. In various examples, the convolution module 541 also receives an input vector with x, y, and z components representing the head rotation direction, which is used to transform the binaural audio to match the head rotation direction. In other embodiments, alternative configurations, different, or additional components may be included in the convolution module 541.Furthermore, functionality associated with a component of the convolution module 541 may be achieved by another component included in the convolution module 541 or another module or system.

[0037] Encoder 545 receives STFT (Short Form Fourier Transform) spectra. In various examples, the input data to encoder 545 is frequency-domain STFT spectra derived from input audio data. The input data includes input tensors, each of which may comprise multiple frames of data.

[0038] In various examples, an STFT is a Fourier-related transform used to determine the sinusoidal frequency and phase content of local portions of a signal as it changes over time. Generally, STFTs are computed by dividing a longer time-domain signal into shorter segments of equal length and then calculating the Fourier transform separately on each shorter segment. This results in the Fourier spectrum on each shorter segment. The changing spectra can be plotted as a function of time, for example, as a spectrogram. In some examples, the STFT is a discrete-time STFT, so the data to be transformed is divided into tensors or frames (which usually overlap to reduce artifacts at the boundary).Each tensor or frame is Fourier transformed, and the complex result is added to a matrix that records the magnitude and phase for each time point and frequency. In some examples, an input tensor has a size HxBxC, where H denotes the height of the input tensor (e.g., the number of rows in the input tensor or the number of data elements in a row), W denotes the width of the input tensor (e.g., the number of columns in the input tensor or the number of data elements in a row), and C denotes the depth of the input tensor (e.g., the number of input channels).

[0039] By inverting the STFT, an inverse STFT can be generated. In various examples, the STFT is processed by the DNN and is then inverted in the decoder 547 or before being input to the decoder 547. By inverting the STFT, the encoded frequency-domain signal from the frequency encoder 545 can be recombined with the encoded time-domain signal from the time encoder 543. One way to invert the STFT is to use the overlap-addition method, which also allows modifications to the complex STFT spectrum. This results in a versatile signal processing method referred to as the overlap-addition method. In various examples, the output of the decoder 547 is a rotated binaural time-domain audio output signal.

[0040] The data store 551 stores data received, generated, used, or otherwise associated with the DNN module 501. For example, the data store 551 stores the datasets used by the training module 521 and the validation module 531. The data store 551 may also store data generated by the training module 521 and the validation module 531, such as the hyperparameters for training the DNN, internal parameters of the trained DNN (e.g., weights, etc.), sparse acceleration data (e.g., sparse bitmap, etc.), and so on. In some embodiments, the data store 551 is a component of the DNN module 501. In other embodiments, the data store 551 may be external to the DNN module 501 and may communicate with the DNN module 501 over a network.

[0041] Fig. 6 is a block diagram of an example neural network architecture 600 that can perform virtual rotation of input sound locations, according to various embodiments. The neural network architecture includes binaural time-domain input signals 612a, 612b and binaural frequency-domain input signals 622a, 622b. In some examples, the frequency-domain input signals 622a, 622b are spectra transformed from the time-domain input signals 612a, 612b. In some examples, a short-time Fourier transform is used to transform the time-domain input signals 612a, 612b into frequency-domain STFT spectra.

[0042] The time-domain input signals 612a, 612b are input to a time-domain encoder 610, which includes multiple time-domain encoder layers 610a, 610b, 610c, 610d, 610e. In some examples, the time-domain encoder 610 is a convolutional encoder and includes convolutional U-Nets for time-domain signals. The time-domain encoder 610 receives the two channels of the time-domain input signals 612a, 612b in a first time-domain encoder convolutional layer 610a. The time-domain encoder 610 also receives an input representing the spherical direction in which the head has rotated. The spherical direction in which the head has turned can be input into the neural network architecture 600 as Cartesian coordinates 602 (components x, y, z of a vector representing the direction in which the head has turned), or using another coordinate system.In various examples, the vector component input, including the Cartesian coordinates 602, is processed by multiple fully connected neural network layers (vector component layers 604). In some examples, the vector component layers 604 expand the coordinates 602 into a vector and / or a tensor using methods similar to those used by neural network encoders and / or neural network decoders. In some examples, expanding the input to the neural network can improve training routines and / or training results of the neural network. The output 606 of the vector component layers 604 is input to each layer 610a, 610b, 610c, 610d, 610e of the time-domain encoder 610.The neural network 600 uses the output 606 of the vector component layers 604 to transform the audio time domain input signals 612a, 612b to generate the binaural audio output 632a, 632b consistent with the new perspective of the head.

[0043] The first time-domain encoder convolutional layer 610a processes the two channels of time-domain input signals 612a, 612b and the output of the vector component layers 604 and outputs 128 channels of time-domain outputs to a second time-domain encoder convolutional layer 610b.The second time-domain encoder convolutional layer 610b receives the 128 channels of time-domain signals and the output of the vector component layers 604 and outputs 256 channels of time-domain outputs to a third time-domain encoder convolutional layer 610c. The third time-domain encoder convolutional layer 610c receives the 256 channels of time-domain signals and the output of the vector component layers 604 and outputs 512 channels of time-domain outputs to a fourth time-domain encoder convolutional layer 610d. The fourth time-domain encoder convolutional layer 610d receives the 512 channels of time-domain signals and the output of the vector component layers 604 and outputs 1024 channels of time-domain outputs to a fifth time-domain encoder convolutional layer 610e. The fifth time-domain encoder convolutional layer 610e receives the 1024 channels of time domain signals and the output of the vector component layers 604 and outputs 2048 channels of time domain outputs.In some examples, the output of the fifth time-domain encoder convolutional layer 610e is the output of the time-domain encoder 610. The output of the time-domain encoder 610 is input to an adder 650.

[0044] The frequency-domain input signals 622a, 622b are input to a frequency-domain encoder 620, which includes a plurality of time-domain encoder layers 620a, 620b, 620c, 620d, 620e. In some examples, the frequency-domain encoder 620 is a convolutional encoder for frequency-domain STFT spectra. The frequency-domain encoder 620 receives the two channels of spectra (frequency-domain input signals 622a, 622b) in a first frequency-domain encoder convolutional layer 620a and outputs 128 channels of frequency-domain outputs to a second frequency-domain encoder convolutional layer 620b.The second frequency-domain encoder convolutional layer 620b receives the 128 channels of time-domain signals and outputs 256 channels of frequency-domain outputs to a third frequency-domain encoder convolutional layer 620c. The third frequency-domain encoder convolutional layer 620c receives the 256 channels of frequency-domain signals and outputs 512 channels of frequency-domain outputs to a fourth frequency-domain encoder convolutional layer 620d. The fourth frequency-domain encoder convolutional layer 620d receives the 512 channels of frequency-domain signals and outputs 1024 channels of frequency-domain outputs to a fifth frequency-domain encoder convolutional layer 620e. The fifth frequency-domain encoder convolutional layer 620e receives the 1024 channels of frequency-domain signals and outputs 2048 channels of frequency-domain outputs. In some examples, the output of the fifth time-domain encoder convolutional layer 620e is the output of the time-domain encoder 620.The output of the frequency domain encoder 620 is input to the adder 650 and combined there with the output of the time domain encoder 610.

[0045] The output of adder 650 is received by a time-domain decoder 630. Time-domain decoder 630 includes multiple time-domain decoder layers 630a, 630b, 630c, 630d, 630e. In some examples, time-domain encoder 630 is a convolutional encoder and includes convolutional U-Nets for time-domain signals. Time-domain encoder 630 also receives an input representing the spherical direction in which the head has rotated. The output 606 of vector component layers 604 is input to each layer 630a, 630b, 630c, 630d, 630e of time-domain encoder 630. The neural network 600 uses the output 606 of the vector component layers 604 to generate the binaural audio output 632a, 632b consistent with the new head perspective. The time-domain decoder also receives output signals directly from corresponding layers of the time-domain encoder.In particular, each layer 630a, 630b, 630c, 630d, 630e of the time domain decoder receives the output signal from the corresponding layer 610a, 610b, 610c, 610d, 610e that has generated the same number of output channels, and the decoder layer 630a, 630b, 630c, 630d, 630e receives as input.

[0046] The first time-domain decoder convolutional layer 630a processes the 2048 channels of time-domain input signals from the adder 650, 2048 channels of time-domain encoder signals from the time-domain encoder layer 610e, and the output 606 of the vector component layers 604, and outputs 1024 channels of time-domain outputs to a second time-domain decoder convolutional layer 630b.The second time-domain decoder convolutional layer 630b receives the 1024 channels of time-domain signals from the first time-domain decoder convolutional layer 630a, 1024 channels of time-domain encoder signals from the time-domain encoder layer 610d, and the output 606 of the vector component layers 604, and outputs 512 channels of time-domain outputs to a third time-domain decoder convolutional layer 630c. The third time-domain encoder convolutional layer 630c receives the 512 channels of time-domain signals from the second time-domain decoder convolutional layer 630b, 512 channels of time-domain encoder signals from the time-domain encoder layer 610c, and the output 606 of the vector component layers 604, and outputs 256 channels of time-domain outputs to a fourth time-domain decoder convolutional layer 630d.The fourth time-domain decoder convolutional layer 630d receives the 256 channels of time-domain signals from the third time-domain decoder convolutional layer 630c, 256 channels of time-domain encoder signals from the second time-domain encoder layer 610b, and the output 606 of the vector component layers 604, and outputs 128 channels of time-domain outputs to a fifth time-domain decoder convolutional layer 630e. The fifth time-domain decoder convolutional layer 630e receives the 128 channels of time-domain signals from the fourth time-domain decoder convolutional layer 630d, 128 channels of time-domain signals from the first time-domain encoder convolutional layer 610a, and the output 606 of the vector component layers 604, and outputs two channels of time-domain outputs 632a, 632b.In some examples, the output of the fifth time-domain decoder convolutional layer 630e is the output of the time-domain decoder 630, and the output of the decoder 630 is the output of the neural network architecture 600.

[0047] In some examples, the time domain encoder 610 and the frequency domain encoder 620 may have a shared cross-domain bottleneck such that both the output of the time domain encoder 610 and the output of the frequency domain encoder 620 are added into additional encoder layers before reaching the decoder 630.

[0048] The neural network architecture 600, which includes the time-domain encoder 610 and the time-domain decoder 630, with multiple blocks and block-to-block skip connections, may be a U-Net. The addition of the frequency-domain encoder 620 results in an additional output of the frequency-domain encoder 620, which is combined with the output of the time-domain encoder 610 at the U-Net bottleneck in the adder 650. Thus, the neural network architecture 600 is a multi-domain architecture.

[0049] According to various implementations, the neural network architecture 600 is in Fig. 6 shows an example of a neural network that can be used for virtual rotation of signal source locations. In various examples, the neural network may comprise an architecture similar to DEMUCS and / or hybrid DEMUCS. In some examples, the architecture may comprise a U-Net encoder and / or a U-Net decoder structure. In some examples, the encoder and decoder may comprise symmetric structures. In some examples, an encoder layer comprises a convolution. In one example, the convolution may comprise a kernel size of eight, a stride of four, a first layer with a fixed number of channels (e.g., 48 or 64), and a doubling of the number of channels in subsequent layers. The neural network architecture may comprise a ReLU (Rectified Linear Unit), and the neural network architecture may comprise a 1x1 convolution with a gated linear unit activation.A decoder layer can sum the contribution from the U-Net skip connection and the previous layer, applying a 1x1 convolution with GLU. In some examples, the input audio data is 44.1 kHz audio. In some examples, the input audio data is upsampled by a factor of two before entering the encoder (to limit aliasing from the outermost layers), and the decoder output is downsampled by a factor of two.

[0050] In some examples, the neural network is an architecture inspired by hybrid DEMUs and includes multi-domain analysis and prediction capabilities. The architecture may include a temporal branch, a spectral branch, and shared layers. The temporal branch receives as input a waveform and processes the waveform. In some examples, the temporal branch includes GELU (Gaussian Error Linear Units) for activations. In some examples, the temporal branch includes multiple layers (e.g., five layers), and the layers reduce the number of time steps by a factor of 1024. The spectral branch receives as input a spectrogram generated using an STFT function. The spectrogram is a frequency representation of the waveform input to the temporal branch. In some examples, the STFT is obtained over 4096 time steps with a hop length of 1024.Thus, in some examples, the number of time steps for the spectral branch matches the output of the temporal branch's encoder. In some examples, the spectral branch performs the same convolutions as the temporal branch, but the spectral branch performs the convolutions in the frequency dimension. Each layer of the spectral branch reduces the number of frequencies by a factor of four. In some examples, a fifth layer of the spectral branch reduces the number of frequencies by a factor of eight.

[0051] In some examples, the spectral branch may perform frequency-wise convolutions. The number of frequency bins at each layer of the neural network may be divided by four. In some examples, the last layer has eight frequency bins, which can be reduced to one using a convolution with a kernel size of eight and no padding. In some examples, the spectrogram input to the neural network may be represented as an amplitude spectrogram or as complex numbers. In some examples, the spectral branch output is transformed into a waveform and summed with the temporal branch output, and the output from the summer is in the waveform domain. Example procedure for virtual rotation

[0052] Fig. 7 is a flowchart of an example method 700 for virtual rotation of binaural audio signals according to various embodiments. In step 710, a binaural audio input signal is received, including a right audio input signal and a left audio input signal. In step 720, the binaural audio signal and a head rotation angle are input to a neural network. In various examples, the neural network is configured to output a virtual rotation of the binaural audio input signal, wherein the virtual rotation includes processing the binaural audio input signal to change the perceived source position of sounds in the binaural audio input signal such that the sound locations are perceived as unchanged despite the head rotation. Accordingly, in various examples, the sound locations are not perceived to rotate with the head, even when a listener is wearing headphones.

[0053] In step 730, a virtual rotation angle is determined based on the head rotation angle. As described above, the virtual rotation angle may be the opposite of the head rotation angle. In step 740, the binaural audio input signal is transformed into a binaural frequency-domain signal. In some examples, a Fourier transform is used to transform the signal, and in some examples, an STFT is used to transform the signal into a binaural frequency-domain signal.

[0054] In step 750, a rotated right audio signal and a rotated left audio signal are generated based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal. As described above, a neural network, such as the DNN module 501 of Fig. 5 or the neural network 600 of Fig. 6 generate virtually right- and left-rotated audio signals based on the virtual rotation angle and the binaural audio input signal. According to various examples, the virtually right- and left-rotated audio signals are adjusted to change the perceived source position of various sounds in the right and left audio input signals such that the perceived source position of various sounds is virtually rotated by the virtual rotation angle. In step 760, a binaural audio output signal rotated by the virtual rotation angle is output from the neural network. The binaural audio output signal includes the rotated right audio signal and the rotated left audio signal. Example data processing device

[0055] Fig. Figure 8 is a block diagram of an exemplary computing device 800 according to various embodiments. In some embodiments, the computing device 800 may be implemented for at least a portion of the DNN module 501 in Fig. 5 and the neural network 600 of Fig. 6. A number of components are in Fig. 8 as included in the computing device 800, but any one or more of these components may be omitted or duplicated as appropriate for the application. In some embodiments, some or all of the components included in the computing device 800 may be attached to one or more motherboards. In some embodiments, some or all of these components may be fabricated on a single SoC (System-on-a-Chip) die. Additionally, in various embodiments, the computing device 800 may include one or more of the Fig. 8, but the computing device 800 may include interface circuitry for coupling to the one or more components. For example, the computing device 800 may not include a display device 806, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 806 may be coupled. In another set of examples, the computing device 800 may not include a video input device 818 or a video output device 808, but may include video input or output device interface circuitry (e.g., connectors and supporting circuitry) to which a video input device 818 or a video output device 808 may be coupled.

[0056] Computing device 800 may include a processing device 802 (e.g., one or more processing devices). Processing device 802 processes electronic data from registers and / or from memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. Computing device 800 may include memory 804, which may itself include one or more storage devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM), high-bandwidth memory (HBM)), flash memory, solid-state memory, and / or a hard disk. In some embodiments, memory 804 may include memory that shares a die with processing device 802.In some embodiments, memory 804 includes one or more non-transitory computer-readable media storing instructions necessary for occupancy allocation, e.g., as described above in connection with. Fig. 7 described method 700 or some operations performed by the DNN system 501 in Fig. 6 are executable. The instructions stored in the one or more non-transitory computer-readable media can be executed by the processing device 802.

[0057] In some embodiments, computing device 800 may include a communications chip 812 (e.g., one or more communications chips). For example, communications chip 812 may be configured to manage wireless communications for transferring data to and from computing device 800. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, etc., that can communicate data using modulated electromagnetic radiation through a non-solid medium. The term does not imply that the associated devices do not include wires, although in some embodiments they may not.

[0058] The communication chip 812 may implement any of a number of wireless standards or protocols, including, but not limited to, Institute for Electrical and Electronic Engineers (IEEE) standards, including Wi-Fi (IEEE 802.10 family), IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment), the Long-Term Evolution (LTE) project, along with all modifications, updates, and / or revisions (e.g., the Advanced LTE project, the Ultra Mobile Broadband (UMB) project (also referred to as "3GPP2"), etc.). IEEE 1202.16-compliant Broadband Wireless Access (BWA) networks are commonly referred to as WiMAX networks, an acronym that stands for Worldwide Interoperability for Microwave Access, a certification mark for products that pass conformance and interoperability tests for the IEEE 1202.16 standards.The communication chip 812 may operate according to a GSM (Global System for Mobile Communication), GPRS (General Packet Radio Service), UMTS (Universal Mobile Telecommunications System), HSPA (High Speed Packet Access), E-HSPA (Evolved HSPA), or LTE network. The communication chip 812 may operate according to EDGE (Enhanced Data for GSM Evolution), GERAN (GSM EDGE Radio Access Network), UTRAN (Universal Terrestrial Radio Access Network), or E-UTRAN (Evolved UTRAN). The communication chip 812 may operate according to CDMA (Code-Division Multiple Access), TDMA (Time Division Multiple Access), DECT (Digital Enhanced Cordless Telecommunications), EV-DO (Evolution-Data Optimized), and derivatives thereof, as well as any other wireless protocols referred to as 3G, 4G, 5G, and beyond. In other embodiments, the communication chip 812 may operate according to other wireless protocols.The computing device 800 may include an antenna 822 to support wireless transmissions and / or to receive other wireless transmissions (such as AM or FM radio transmissions).

[0059] In some embodiments, the communication chip 812 may manage wired communications, such as electrical, optical, or other suitable communication protocols (e.g., Ethernet). As mentioned above, the communication chip 812 may include multiple communication chips. For example, a first communication chip 812 may be dedicated to shorter-range wireless communications, such as Wi-Fi or Bluetooth, and a second communication chip 812 may be dedicated to longer-range wireless communications, such as GPS (Global Positioning System), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication chip 812 may be dedicated to wireless communications, and a second communication chip 812 may be dedicated to wired communications.

[0060] Computing device 800 may include battery / power supply circuitry 814. Battery / power supply circuitry 814 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of computing device 800 to a power source separate from computing device 800 (e.g., AC power from the mains).

[0061] Computing device 800 may include a display device 806 (or corresponding interface circuitry, as discussed above). Display device 806 may include any visual indicator, such as a point-of-view display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD display), a light-emitting diode display, or a flat panel display.

[0062] Computing device 800 may include a video output device 808 (or corresponding interface circuitry, as discussed above). Video output device 808 may include any device that generates an audible indicator, such as speakers, headphones, or earbuds.

[0063] Data processing device 800 may include a video input device 818 (or corresponding interface circuitry, as discussed above). Video input device 818 may include any device that generates a signal representing sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments with a musical instrument digital interface (MIDI) output).

[0064] Computing device 800 may include a GPS device 816 (or corresponding interface circuitry, as discussed above). GPS device 816 may be in communication with a satellite-based system and may receive a position of computing device 800, as is known in the art.

[0065] Data processing device 800 may include another output device 810 (or corresponding interface circuitry, as discussed above). Examples of another output device 810 would be a video codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.

[0066] The computing device 800 may include another input device 820 (or corresponding interface circuitry, as discussed above). Examples of the other input device 820 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a barcode reader, a quick response code (QR) reader, any sensor, or a radio frequency identification (RFID) reader.

[0067] Computing device 800 may comprise any desired form factor, such as a handheld or mobile computing system (e.g., a cellular phone, a smartphone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a personal digital assistant (PDA), an ultramobile personal computer, etc.), a desktop computing system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computing system. In some embodiments, computing device 800 may be any other electronic device that processes data. Selected examples

[0068] The following paragraphs provide various examples of the embodiments disclosed herein.

[0069] Example 1 provides a computer-implemented method, comprising: receiving a binaural audio input signal comprising a right audio input signal and a left audio input signal; inputting the binaural audio input signal and a head rotation angle to a neural network; determining a virtual rotation angle based on the head rotation angle in the neural network; transforming the binaural audio input signal into a binaural frequency-domain signal in the neural network; generating a rotated right audio signal and a rotated left audio signal by the neural network based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency-domain signal; and Outputting a binaural audio output signal rotated by the virtual rotation angle, including the rotated right audio signal and the rotated left audio signal, through the neural network.

[0070] Example 2 provides the computer-implemented method of Example 1, wherein the neural network comprises a time-domain encoder comprising a plurality of time-domain encoder layers, and further comprising: Inputting the right audio input signal, the left audio input signal, and the virtual rotation angle into a first time-domain encoder layer of the plurality of time-domain encoder layers; and outputting a plurality of encoded time-domain signals from the time-domain encoder.

[0071] Example 3 provides the computer-implemented method of Example 2, wherein a second time-domain encoder layer of the plurality of time-domain encoder layers receives an output from the first time-domain encoder layer and the virtual rotation angle.

[0072] Example 4 provides the computer-implemented method of Example 3, wherein a number of channels output from each of the plurality of time-domain encoder layers is greater than a number of channels input to each of the plurality of time-domain encoder layers.

[0073] Example 5 provides the computer-implemented method of Example 2, wherein the neural network comprises a frequency domain encoder comprising a plurality of frequency domain encoder layers, and further comprising inputting the binaural frequency domain signal to a first frequency domain encoder layer of the plurality of frequency domain encoder layers and outputting a plurality of frequency domain encoding signals from the frequency domain encoder.

[0074] Example 6 provides the computer-implemented method of Example 5, wherein the neural network comprises an adder and further comprises adding the plurality of encoded time-domain signals and the plurality of encoded frequency-domain signals to generate a plurality of added encoded signals, and inputting the plurality of added encoded signals to a time-domain decoder.

[0075] Example 7 provides the computer-implemented method of Example 1, wherein the head rotation angle comprises Cartesian coordinates representing a direction in which a head rotated.

[0076] Example 8 provides the computer-implemented method of Example 1, further comprising: training the neural network using synthetic audio samples generated using head rotation transfer functions.

[0077] Example 9 provides one or more non-transitory computer-readable media storing instructions executable to perform operations, the operations comprising: receiving a binaural audio input signal comprising a right audio input signal and a left audio input signal; inputting the binaural audio input signal and a head rotation angle to a neural network; determining a virtual rotation angle based on the head rotation angle in the neural network; transforming the binaural audio input signal into a binaural frequency domain signal in the neural network; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal;Outputting a binaural audio output signal rotated by the virtual rotation angle, including the rotated right audio signal and the rotated left audio signal, through the neural network.;

[0078] Example 10 provides one or more non-transitory computer-readable media of Example 9, wherein the neural network comprises a time-domain encoder comprising a plurality of time-domain encoder layers, and further comprising: Inputting the right audio input signal, the left audio input signal, and the virtual rotation angle into a first time-domain encoder layer of the plurality of time-domain encoder layers; and outputting a plurality of encoded time-domain signals from the time-domain encoder.

[0079] Example 11 provides one or more non-transitory computer-readable media of Example 10, wherein the neural network comprises a frequency domain encoder comprising a plurality of frequency domain encoder layers, and the operations further comprise inputting the binaural frequency domain signal to a first frequency domain encoder layer of the plurality of frequency domain encoder layers and outputting a plurality of frequency domain encoded signals from the frequency domain encoder.

[0080] Example 12 provides one or more non-transitory computer-readable media of Example 11, wherein the neural network comprises an adder, and the operations further comprise adding the plurality of encoded time-domain signals and the plurality of encoded frequency-domain signals to generate a plurality of added encoded signals, and inputting the plurality of added encoded signals to a time-domain decoder.

[0081] Example 13 provides one or more non-transitory computer-readable media of Example 9, wherein the head rotation angle comprises Cartesian coordinates representing a direction in which a head rotated.

[0082] Example 14 provides one or more non-transitory computer-readable media of Example 9, the operations further comprising: training the neural network using synthetic audio samples generated using head rotation transfer functions.

[0083] Example 15 provides an apparatus comprising a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations, the operations comprising: receiving a binaural audio input signal comprising a right audio input signal and a left audio input signal; inputting the binaural audio input signal and a head rotation angle to a neural network; determining a virtual rotation angle based on the head rotation angle in the neural network; transforming the binaural audio input signal into a binaural frequency domain signal in the neural network;Generating a rotated right audio signal and a rotated left audio signal by the neural network based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; Outputting a binaural audio output signal rotated by the virtual rotation angle, including the rotated right audio signal and the rotated left audio signal, by the neural network.

[0084] Example 16 provides the apparatus of Example 15, wherein the neural network comprises a time-domain encoder comprising a plurality of time-domain encoder layers, and further comprising: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle to a first time-domain encoder layer of the plurality of time-domain encoder layers; and outputting a plurality of encoded time-domain signals from the time-domain encoder.

[0085] Example 17 provides the apparatus of Example 16, wherein the neural network comprises a frequency domain encoder comprising a plurality of frequency domain encoder layers, and further comprising inputting the binaural frequency domain signal to a first frequency domain encoder layer of the plurality of frequency domain encoder layers and outputting a plurality of frequency domain encoding signals from the frequency domain encoder.

[0086] Example 18 provides the apparatus of Example 17, wherein the neural network comprises an adder, and the operations further comprise: adding the plurality of encoded time domain signals and the plurality of encoded frequency domain signals to generate a plurality of added encoded signals; and Inputting the multiple added coded signals into a time domain decoder.

[0087] Example 19 provides the apparatus of Example 15, wherein the head rotation angle comprises Cartesian coordinates representing a direction in which a head rotated.

[0088] Example 20 provides the apparatus of Example 15, wherein the operations further comprise: training the neural network using synthetic audio samples generated using head rotation transfer functions.

[0089] The above description of illustrated implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of and examples of the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as will be appreciated by those skilled in the art. These modifications may be made to the disclosure in light of the above detailed description.

Claims

[1] Computer-implemented method comprising: Receiving a binaural audio input signal comprising a right audio input signal and a left audio input signal; Inputting the binaural audio input signal and a head rotation angle into a neural network; Determining a virtual rotation angle based on the head rotation angle in the neural network; Transforming the binaural audio input signal into a binaural frequency domain signal in the neural network; Generating a rotated right audio signal and a rotated left audio signal by the neural network based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; and Outputting a binaural audio output signal rotated by the virtual rotation angle, including the rotated right audio signal and the rotated left audio signal, through the neural network. [2] The computer-implemented method of claim 1, wherein the neural network comprises a time-domain encoder comprising a plurality of time-domain encoder layers, and further comprising: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle into a first time-domain encoder layer of the plurality of time-domain encoder layers; and Outputting a plurality of encoded time domain signals from the time domain encoder. [3] The computer-implemented method of claim 2, wherein a second time-domain encoder layer of the plurality of time-domain encoder layers receives an output from the first time-domain encoder layer and the virtual rotation angle. [4] The computer-implemented method of claims 2-3, wherein a number of channels output from each of the plurality of time-domain encoder layers is greater than a number of channels input to each of the plurality of time-domain encoder layers. [5] A computer-implemented method according to claims 2-4, wherein the time domain encoder is a convolutional encoder. [6] The computer-implemented method of claim 5, wherein the neural network comprises a frequency domain encoder comprising a plurality of frequency domain encoder layers, and further comprising: inputting the binaural frequency domain signal into a first frequency domain encoder layer of the plurality of frequency domain encoder layers; and Outputting a plurality of coded frequency domain signals from the frequency domain encoder. [7] The computer-implemented method of claim 6, wherein the neural network comprises an adder and further comprises adding the plurality of encoded time-domain signals and the plurality of encoded frequency-domain signals to generate a plurality of added encoded signals, and inputting the plurality of added encoded signals to a time-domain decoder. [8] The computer-implemented method of claims 1-7, wherein the head rotation angle comprises Cartesian coordinates representing a direction in which a head rotated. [9] A computer-implemented method according to claims 1-8, further comprising: Training the neural network using synthetic audio samples generated using head rotation transfer functions. [10] One or more non-transitory computer-readable media storing instructions executable to perform operations, the operations comprising: Receiving a binaural audio input signal comprising a right audio input signal and a left audio input signal; Inputting the binaural audio input signal and a head rotation angle into a neural network; Determining a virtual rotation angle based on the head rotation angle in the neural network; Transforming the binaural audio input signal into a binaural frequency domain signal in the neural network; Generating a rotated right audio signal and a rotated left audio signal by the neural network based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; and Outputting a binaural audio output signal rotated by the virtual rotation angle, including the rotated right audio signal and the rotated left audio signal, through the neural network. [11] One or more non-transitory computer-readable media according to claim 10, wherein the neural network comprises a time-domain encoder comprising a plurality of time-domain encoder layers, and further comprising: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle into a first time-domain encoder layer of the plurality of time-domain encoder layers; and Outputting a plurality of encoded time domain signals from the time domain encoder. [12] One or more non-transitory computer-readable media according to claim 11, wherein a second time-domain encoder layer of the plurality of time-domain encoder layers receives an output from the first time-domain encoder layer and the virtual rotation angle. [13] The computer-implemented method of claims 10-12, wherein a number of channels output from each of the plurality of time-domain encoder layers is greater than a number of channels input to each of the plurality of time-domain encoder layers. [14] One or more non-transitory computer-readable media according to claim 10, wherein the neural network comprises a time-domain encoder comprising a plurality of time-domain encoder layers, and the operations further comprise: inputting the binaural frequency domain signal into a first frequency domain encoder layer of the plurality of frequency domain encoder layers; and Outputting a plurality of coded frequency domain signals from the frequency domain encoder. [15] One or more non-transitory computer-readable media according to claim 14, wherein the neural network comprises an adder and the operations further comprise: adding the plurality of encoded time domain signals and the plurality of encoded frequency domain signals to generate a plurality of added encoded signals; and Inputting the multiple added coded signals into a time domain decoder. [16] One or more non-transitory computer-readable media according to claims 10-15, wherein the head rotation angle comprises Cartesian coordinates representing a direction in which a head rotated. [17] One or more non-transitory computer-readable media according to claims 10-16, wherein the operations further comprise: training the neural network using synthetic audio samples generated using head rotation transfer functions. [18] Device comprising: a computer processor for executing computer program instructions and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including: Receiving a binaural audio input signal comprising a right audio input signal and a left audio input signal; Inputting the binaural audio input signal and a head rotation angle into a neural network; Determining a virtual rotation angle based on the head rotation angle in the neural network; Transforming the binaural audio input signal into a binaural frequency domain signal in the neural network; Generating a rotated right audio signal and a rotated left audio signal by the neural network based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; and Outputting a binaural audio output signal rotated by the virtual rotation angle, including the rotated right audio signal and the rotated left audio signal, through the neural network. [19] The apparatus of claim 18, wherein the neural network comprises a time domain encoder comprising a plurality of time domain encoder layers, and the operations further comprise: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle into a first time-domain encoder layer of the plurality of time-domain encoder layers; and Outputting a plurality of encoded time domain signals from the time domain encoder. [20] The apparatus of claim 19, wherein a second time-domain encoder layer of the plurality of time-domain encoder layers receives an output from the first time-domain encoder layer and the virtual rotation angle. [21] The apparatus of claims 19-20, wherein a number of channels output from each of the plurality of time domain encoder layers is greater than a number of channels input to each of the plurality of time domain encoder layers. [22] The apparatus of claims 19-21, wherein the neural network comprises a frequency domain encoder comprising a plurality of frequency domain encoder layers, and the operations further comprise: inputting the binaural frequency domain signal into a first frequency domain encoder layer of the plurality of frequency domain encoder layers; and Outputting a plurality of coded frequency domain signals from the frequency domain encoder. [23] The apparatus of claim 22, wherein the neural network comprises an adder and the operations further comprise: adding the plurality of encoded time domain signals and the plurality of encoded frequency domain signals to generate a plurality of added encoded signals; and Inputting the multiple added coded signals into a time domain decoder. [24] The apparatus of claims 18-23, wherein the head rotation angle comprises Cartesian coordinates representing a direction in which a head rotated. [25] The apparatus of claims 18-24, further comprising: training the neural network using synthetic audio samples generated using head rotation transfer functions.