Deep learning solution for virtual rotation of two-channel audio signals

The audio regression method of deep learning process two-channel audio signals is processed, and the rotated audio output is generated, solving the problem that the two-channel audio signal does not change when the head rotates, and improving the naturalness and immersion of the listening experience.

CN120455920APending Publication Date: 2025-08-08INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510014215.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2025-01-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to realize the automatic change of the two-channel audio signal when the head rotates, resulting in an unnatural listening experience.

Method used

Using deep learning-based audio regression method, a new two-channel audio output signal with corresponding rotation environment is generated by a deep neural network to process the two-channel audio signal and rotation angle.

Benefits of technology

The two-channel audio signal is automatically changed with the head rotation, which enhances the immersion and naturalness of listening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455920A_ABST
    Figure CN120455920A_ABST
Patent Text Reader

Abstract

Techniques are provided herein for providing a binaural sound signal that is virtually rotated to match the head rotation such that when a user rotates their head, an audio output for a headphone is perceived to maintain its position relative to the user. In particular, techniques are presented for extracting spherical position information that has been embedded in a binaural signal to generate a binaural sound signal that changes to match head rotation. The deep learning-based audio regression method may use a 2-channel two-channel audio signal and a rotation angle as inputs, and generate a new two-channel audio output signal having a rotated environment corresponding to the rotation angle. The deep learning-based audio regression method may be implemented as a neural network, and may include deep learning operations, such as convolution, pooling, element-by-element operations, linear operations, and non-linear operations. Deep learning operations may be performed on internal parameters and one or more activations of the DNN.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to two-channel audio signals, and in particular to virtual rotation of two-channel audio signals. Background Art

[0002] Binaural sound is audio that can be heard by both ears. In a natural environment, sound waves originating from directions other than the center of the listener's head arrive at each ear at slightly different times and at different volumes. Using these differences in the arrival times and volume of the sound waves, the listener's brain can calculate the source of the sound. By using binaural recordings and creating binaural sound, the perception of sound originating from a selected location relative to the listener's head can be replicated, with each speaker of a pair of stereo headphones (or in-ear headphones) producing slightly different sounds. When binaural sound is played through headphones, the listener hears three-dimensional (3D) binaural sound. However, binaural recordings used to generate binaural sound are recorded and / or generated with a fixed head orientation. This fixed head orientation for binaural sound results in a phenomenon where, when the listener rotates their head, the binaural sound rotates with the head, rather than automatically changing with the rotation of the head. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate description, similar reference numerals designate similar structural elements. In the accompanying drawings, the embodiments are illustrated by way of example and not limitation.

[0004] Figure 1 is a block diagram of an example of a sound localization process using binaural sound perception.

[0005] Figure 2 is a block diagram illustrating a system for generating binaural sound with selected perceived positional source angles using HRTFs.

[0006] Figure 3A Four stationary sound sources surrounding a person's head are shown, according to various embodiments.

[0007] Figure 3B The rotation of four stationary sound sources around a person's head is shown in accordance with various embodiments.

[0008] Figure 4 is a diagram of an example virtual rotation system in accordance with various embodiments.

[0009] Figure 5 is a block diagram of a deep neural network (DNN) module that can be used for a virtual rotation system according to various embodiments.

[0010] Figure 6 is a block diagram illustrating an example of a neural network architecture 600 that can perform a virtual rotation of an input sound position, according to various embodiments.

[0011] Figure 7 is a flow chart illustrating an example method for virtual rotation according to various embodiments.

[0012] Figure 8 is a block diagram of an example computing device in accordance with various embodiments. DETAILED DESCRIPTION

[0013] Overview

[0014] Listeners can generally determine where sounds in the environment are coming from because sounds arrive at their ears differently, depending on their position relative to the head. For example, a sound coming directly from the right side of a person's head will arrive at the right ear clearly and crisply, while the left ear will receive a version of the sound that is slightly attenuated by the presence of the head and body in the sound's acoustic path. The human brain uses the differences in the sounds heard at each ear to estimate the sound's location. However, note that for even very brief and impactful sounds, such as a clap or a gunshot, people don't notice any delay or difference between the sounds heard at each ear; the process is fast and organic.

[0015] Binaural audio can be recorded using a specific capture method that attempts to simulate the way a real human head hears audio. The capture method may include using a special mannequin head and using a binaural headset microphone placed in the person's ears during recording. In some examples, instead of recording the binaural audio directly, a special digital filter based on the head-related transfer function (HRTF) can be used to generate the binaural audio. The HRTF filter is based on a selected direction from which the sound source is to be perceived, with different filter parameters for each selected direction.

[0016] By using a binaural recording and creating binaural sound, the perception of sound originating from a selected position relative to the listener's head can be replicated, with the output for each speaker of a pair of headphones being slightly different. Alternatively, by continuously filtering the audio recording using HRTF filters, the perception of sound originating from a selected position relative to the listener's head can be replicated to generate binaural audio that creates a selected perception of sound direction. When the binaural sound is played through headphones, the listener can hear the binaural sound in three dimensions (3D).

[0017] Typically, binaural recordings used to generate binaural sound are generated with a fixed head orientation, so that when a listener rotates their head, the binaural sound rotates with the head, rather than automatically changing with the head rotation. However, head rotation and the associated changes in the sound field are important listening and immersion characteristics.

[0018] In order to achieve the changes corresponding to the head rotation in binaural sound, it is common to recapture the binaural audio recording using different head positions, or to make the recording using a complex and expensive directional microphone array. However, these expensive and time-consuming strategies may not be suitable for sounds that cannot be re-recorded. Another method that can be used to achieve the changes corresponding to the head rotation in binaural sound is to filter the audio source using a head-related transfer function (HRTF) designed to simulate a binaural experience. However, using HRTF means filtering the audio signal and adding multiple audio channels to generate a corresponding HRTF for each source position. Additionally, HRTF requires multiple microphone capture arrays and a known source position for each signal. Therefore, a simple binaural signal cannot be used to generate a binaural signal rotation using HRTF.

[0019] This paper presents systems and methods for providing a binaural sound signal that can be altered to match head rotation. In particular, techniques are presented for extracting spherical position information already embedded in a binaural signal to generate a binaural sound signal that can be altered to match head rotation. This paper discusses a deep learning-based audio regression method that can use a 2-channel binaural audio signal and a selected rotation angle as input and generate a new binaural audio output signal with a rotated environment corresponding to the selected rotation angle.

[0020] According to various implementations, the virtual audio rotation module may include a deep learning-based audio regression method, which may be implemented as a neural network, such as a deep neural network (DNN). As described herein, a DNN layer may include one or more deep learning operations, such as convolution, pooling, element-by-element operations, linear operations, nonlinear operations, and the like. Deep learning operations in a DNN may be performed on one or more internal parameters (e.g., weights) and one or more activations of the DNN determined during the training phase. Activations may be data points (also referred to as "data elements" or "elements"). The activations or weights of a DNN layer may be elements of a tensor of the DNN layer. A tensor is a data structure with multiple elements across one or more dimensions. Example tensors include vectors (which are one-dimensional tensors) and matrices (which are two-dimensional tensors). Three-dimensional tensors and even higher-dimensional tensors may also exist. A DNN layer may have an input tensor (generated during feature extraction), comprising one or more input activations (also referred to as "input elements") and a weight tensor comprising one or more weights. Weights are elements in the weight tensor. The weight tensor of a convolution may be a kernel, a filter, or a set of filters. The output data of a DNN layer can be an output tensor, which includes one or more output activations (also called "output elements").

[0021] For purposes of explanation, specific numbers, materials, and configurations are set forth to provide a thorough understanding of the illustrative implementations. However, those skilled in the art will appreciate that the present disclosure may be practiced without these specific details, or / and the present disclosure may be practiced using only some of the described aspects. In other cases, well-known features have been omitted or simplified to avoid obscuring the illustrative implementations.

[0022] In addition, reference is made to the accompanying drawings which form a part of this document and in which are shown by way of example embodiments that may be practiced. It should be understood that other embodiments may be utilized and that structural or logical changes may be made without departing from the scope of this disclosure. Therefore, the following detailed description should not be considered restrictive.

[0023] The various operations may be described as a plurality of discrete actions or operations performed sequentially, described in a manner that is most helpful for understanding the claimed subject matter. However, the order of description should not be interpreted as implying that the operations are necessarily order-dependent. In particular, the operations may not be performed in the order presented. The operations may be performed in an order different from that of the described embodiment. In additional embodiments, various additional operations may be performed or the operations described may be omitted.

[0024] For the purposes of this disclosure, the phrase "A and / or B" or the phrase "A or B" means (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B, and / or C" or the phrase "A, B, or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). When the term "between" is used in reference to a measurement range, it includes both ends of the measurement range.

[0025] The phrases "in an embodiment" or "in embodiments" are used in the description, each of which may refer to one or more of the same or different embodiments. The terms "comprising," "including," "having," and the like, as used with respect to the embodiments of the present disclosure are synonymous. The present disclosure may use perspective-based descriptions, such as "above," "below," "top," "bottom," and "side," to explain various features of the accompanying drawings, but these terms are used only to facilitate discussion and do not imply a desired or required orientation. The drawings are not necessarily drawn to scale. Unless otherwise stated, the use of ordinal adjectives "first," "second," and "third," etc., to describe common objects merely indicates different instances of the similar objects referred to, and is not intended to imply that the objects described must be arranged in a given sequence, whether temporal, spatial, ranked, or any other sequence.

[0026] In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.

[0027] The terms "substantially," "close," "approximately," "near," and "about" generally refer to input operands that are within + / - 20% of target values, based on specific values described herein or known in the art. Similarly, terms indicating the orientation of various elements, such as "coplanar," "perpendicular," "orthogonal," "parallel," or any other angle between elements, generally refer to input operands that are within + / - 5-20% of target values, based on specific values described herein or known in the art.

[0028] Additionally, the terms "comprise," "comprising," "include," "including," "have," "having," or any other variations thereof, are intended to cover a non-exclusive inclusion. For example, a method, process, apparatus, or system that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, apparatus, or system. Furthermore, the term "or" refers to an inclusive or rather than an exclusive or.

[0029] The systems, methods, and devices of the present disclosure each have several innovative aspects, no single one of which is solely responsible for all of the desirable attributes disclosed herein. The details of one or more implementations of the subject matter described in this specification are set forth in the following description and drawings.

[0030] Example binaural sound position

[0031] Figure 1 is a block diagram of an example of a sound location process using binaural sound perception. In particular, as shown in the leftmost panel, sounds originate from selected locations around the head 102. 104, of which is the angle from the central axis in the horizontal plane, and θ is the angle from the central axis in the vertical plane, where the central axis is located directly in front of the face, i.e., directly in front of the head. As shown in the middle panel, sound from location 104 reaches each ear using a different acoustic path. In particular, sound reaches the left ear using a direct path 106, and sound reaches the right ear using one or more indirect paths 108a, 108b. The acoustic paths vary depending on head size, head shape, pinna, and environment. The right panel shows that the brain uses the differences between the sounds received at each ear to estimate the location of the sound. 104.

[0032] use Figure 1 Figure 3 shows a graph of a monophonic sound image, which can be generated by selecting a digital filter called an HRTF to filter the monophonic sound using a selected sound source position angle and sending the output signal to a pair of stereo headphones to generate a binaural sound with a selected perceived position source angle. Figure 2is a block diagram illustrating a system for generating binaural sound with selected perceived positional source angles using HRTFs. A virtual sound source position angle 202 indicates the desired perceived position of a sound source. At block 204, an HRTF for each ear is selected based on the virtual sound source position angle 202. An original mono signal 206 is input to a processing block along with the selected HRTF from block 204. The original mono signal 206 is processed using a right HRTF 208a to generate a right ear signal to be played at a right headphone speaker 210a, and the original mono signal 206 is processed using a left HRTF 208b to generate a left ear signal to be played at a left headphone speaker 210b. When a listener listens to the signal in both the right and left headphone speakers 210a and 210b simultaneously, the listener perceives the signal as originating from the sound source specified by the virtual sound source position angle 202.

[0033] Virtual rotation of an example two-channel audio signal

[0034] Figure 3A Four fixed sound sources around a person's head are shown according to various embodiments. In particular, Figure 3A As shown, the sound sources (phone, airplane, car, and cat) remain fixed on an illustrative sphere, and as a person rotates their head, the positions of the sound sources relative to their ears change. Thus, a person can perceive the approximate location of each fixed source in their environment. When a person rotates their head by a selected angle, the angular spherical position of each sound source changes by an equivalent angle in the opposite direction of the head rotation.

[0035] Figure 3B FIG2 shows the rotation of four fixed sound sources around a person's head according to various embodiments. In particular, Figure 3B A virtual environment is shown in which a person's head remains fixed and, conversely, the source positions move relative to the head, which causes the approximate positions of the source positions relative to the head to change in the same manner as if the head were rotated. Thus, a counterclockwise rotation of the person's head by a selected angle is equivalent to a clockwise rotation of each sound source around the person's head by the same selected angle. When a listener is wearing headphones (or in-ear headphones), and the position of the user's head relative to the headphone speakers is fixed, the equivalence can be used to change the perceived sound source positions. In particular, using the equivalence (of counterclockwise rotation of the person's head to clockwise rotation of each sound source), when the user is wearing headphones that play a binaural signal to generate a binaural sound environment, by rotating the perceived audio source positions in the binaural signal by an angle equal to but opposite to the angle of head movement, the perceived sound source positions can move as the person's head moves, as shown in FIG. Figure 3B shown.

[0036] Example Virtual Rotation System Overview

[0037] Figure 4 is a diagram of an example virtual rotation system 400 according to various embodiments. In particular, Figure 4 As shown, system 400 includes a neural network 406 that receives a binaural audio signal 402 and a rotation angle 404. Neural network 406 may be a recurrent neural network that uses binaural audio signal 402 and rotation angle 404 as inputs and generates a new binaural audio output signal 408 having a rotated environment indicated by rotation angle 404. According to various implementations, neural network 406 may be a deep neural network, as described in more detail below.

[0038] Example DNN system

[0039] Figure 5 is a block diagram of a deep neural network (DNN) module 501 that can be used in a virtual rotation system according to various embodiments. Figure 5 In the embodiment of the present invention, DNN module 501 includes: interface module 511, training module 521, verification module 531, convolution module 541, and data storage 551. In other embodiments, DNN module 501 may include: alternative configurations, different or additional components. In addition, the functionality attributed to the components of DNN module 501 may be implemented by different components or different modules or systems included in DNN module 501 (e.g., any neural network and / or deep learning system described herein).

[0040] Interface module 511 facilitates communication between DNN module 501 and other modules or systems. For example, interface module 511 establishes communication between DNN module 501 and an external database to receive data that can be used to train the DNN or input into the DNN to perform a task. As another example, interface module 511 supports DNN module 501 in distributing the DNN to other systems, such as computing devices configured to apply the DNN to perform tasks.

[0041] The training module 521 uses a training data set to train the DNN. In some examples, a training data set can be generated using synthetic audio samples using HRTF. Multiple data sets can be used to provide various audio source types (e.g., human voice, musical instruments, animals, environmental sounds, etc.). For each sample in the data set, multiple random audio sources can be selected, and a random angular direction can be assigned to each selected audio source. Each selected audio source in the selected audio source can be rendered from the corresponding direction based on the assigned random angle using the corresponding HRTF, and the audio sources can be mixed to generate training input data. In order to generate a training output, a second random direction can be determined to represent a new head orientation, and based on the new head orientation and the input direction of each audio source in the audio source, a new relative audio source direction relative to the head can be determined for each audio source in the audio source.

[0042] In one embodiment, where the training module 521 trains a DNN to generate a binaural signal with selected perceived source positions, the training dataset includes: a training binaural signal and training labels. The training labels describe the true position of the sound source in the training signal. In some embodiments, each label in the training dataset corresponds to an angle in the training stereo space (relative to a centerline and a plane). In some embodiments, each label in the training dataset corresponds to a position in the training stereo space. In some embodiments, a portion of the training dataset can be used to initially train the DNN, while the remainder of the training dataset can be retained as a validation subset, which is used by the validation module 531 to validate the performance of the trained DNN. The portion of the training dataset that does not include the adjustment subset and the validation subset can be used to train the DNN.

[0043] The training module 521 also determines hyperparameters for training the DNN. Hyperparameters are variables that specify the DNN training process. Hyperparameters are distinct from parameters within the DNN (e.g., filter weights). In some embodiments, hyperparameters include variables that determine the DNN architecture, such as the number of hidden layers. Hyperparameters also include variables that determine how the DNN is trained, such as the batch size and the number of epochs. The batch size defines the number of training samples to be processed before updating the DNN's parameters. The batch size is equal to or less than the number of samples in the training dataset. The training dataset can be divided into one or more batches. The number of epochs defines the number of times the entire training dataset is passed forward and backward through the entire network. The number of epochs defines the number of times the deep learning algorithm processes the entire training dataset. An epoch means that each training sample in the training dataset has an opportunity to update the DNN's internal parameters. An epoch may include one or more batches. The number of epochs may be 3, 30, 300, 500, or even more.

[0044] Training module 521 defines the architecture of the DNN, for example, based on some of the hyperparameters. The DNN architecture includes an input layer, an output layer, and multiple hidden layers. The DNN input layer may include tensors (e.g., multidimensional arrays) that specify attributes of the input signal (e.g., frequency, volume, and other spectral characteristics). The output layer includes labels for the angle and / or position of the sound source in the input layer. Hidden layers are layers between the input and output layers. Hidden layers include one or more convolutional layers and one or more other types of layers, such as pooling layers, fully connected layers, normalization layers, softmax layers, or logistic layers. The convolutional layers of the DNN abstract the input signal to perform feature extraction. In some examples, feature extraction is based on the spectrogram of the input sound signal. The pooling layer is used to reduce the volume of the input signal after convolution. It is used between two convolutional layers. The fully connected layer involves weights, biases, and neurons. It connects neurons in one layer to neurons in another layer. It is used to classify signals between different categories through training. Note that training a DNN is different from using it in real time, and that latency can be an issue when using a DNN to process data received in real time, an issue that does not exist during training when the dataset can be preloaded.

[0045] In the process of defining the architecture of the DNN, the training module 521 also adds an activation function to the hidden layer or output layer. The activation function of a layer transforms the weighted sum of the input of the layer into the output of the layer. For example, the activation function can be a rectified linear unit activation function, a tangent activation function, or other types of activation functions.

[0046] After the training module 521 defines the architecture of the DNN, the training module 521 inputs a training dataset into the DNN. The training dataset includes a plurality of training samples. Examples of training samples include: the source location of a feature in an audio sample and the ground-truth location of the feature. The training module 521 modifies parameters inside the DNN ("internal parameters of the DNN") to minimize the error between the labels of the training features generated by the DNN and the ground-truth labels of the features. The internal parameters include the weights of the filters in the convolutional layers of the DNN. In some embodiments, the training module 521 uses a cost function to minimize the error.

[0047] Training module 521 can train the DNN for a predetermined number of epochs. The number of epochs is a hyperparameter that defines how many times the deep learning algorithm will work through the entire training dataset. One hyperparameter indicates that each sample in the training dataset has had an opportunity to update the internal parameters of the DNN. After training module 521 completes the predetermined number of epochs, training module 521 can stop updating the parameters in the DNN. A DNN with updated parameters is called a trained DNN.

[0048] The verification module 531 verifies the accuracy of the trained or compressed DNN. In some embodiments, the verification module 531 inputs samples from the validation dataset into the trained DNN and uses the output of the DNN to determine model accuracy. In some embodiments, the validation dataset may consist of some or all samples from the training dataset. Additionally or alternatively, the validation dataset includes additional samples beyond those in the training set. In some embodiments, the verification module 531 may determine an accuracy score that measures the precision, recall, or a combination of precision and recall of the DNN. The verification module 531 may determine the accuracy score using the following metrics: precision = TP / (TP+FP) and recall = TP / (TP+FN), where precision can be the number of correct predictions (TP or true positives) made by the reference classification model out of the total number of its predictions (TP+FP or false positives), and recall can be the number of correct predictions (TP) made by the reference classification model out of the total number of objects that actually possess the attribute in question (TP+FN or false negatives). The F-score (F-score = 2*PR / (P+R)) unifies precision and recall into a single measure.

[0049] The verification module 531 can compare the accuracy score to a threshold score. In the example where the verification module 531 determines that the accuracy score of the enhanced model is less than the threshold score, the verification module 531 instructs the training module 521 to retrain the DNN. In one embodiment, the training module 521 can iteratively retrain the DNN until a stopping condition occurs, for example, the accuracy measurement indicates that the DNN may be sufficiently accurate, or a number of rounds of training have been performed.

[0050] The convolution module 541 performs real-time data processing, for example, for speech enhancement, dynamic noise suppression, blind source separation and / or self-noise cancellation. Figure 5In an embodiment of the invention, the convolution module 541 includes: a time domain encoder 543, a frequency domain encoder 545 and a time domain decoder 547. In some examples, the time domain encoder 543 is a convolutional time domain encoder, the frequency domain encoder 545 is a convolutional frequency domain spectrum encoder, and the time domain decoder 547 is a convolutional time domain decoder. In various examples, the convolution module 541 also receives an input vector having x, y and z components representing the direction of head rotation, which is used to transform the two-channel audio to match the direction of head rotation. In other embodiments, the convolution module 541 may include: alternative configurations, different or additional components. In addition, the functionality of the components attributed to the convolution module 541 may be implemented by different components included in the convolution module 541, the DNN module 501, or different modules or systems.

[0051] The encoder 545 receives a short-form Fourier transform (STFT) spectrum. In various examples, the input data for the encoder 545 is a frequency-domain STFT spectrum derived from the input audio data. The input data includes input tensors, each of which may include multiple frames of data.

[0052] In various examples, the STFT is a Fourier-related transform used to determine the sinusoidal frequency and phase content of a local portion of a signal as it changes over time. Typically, the STFT is calculated by dividing a longer time signal into shorter segments of equal length and then calculating the Fourier transform on each shorter segment separately. This produces a Fourier spectrum on each shorter segment. The changed spectrum can be plotted as a function of time, for example, as a spectrogram. In some examples, the STFT is a discrete-time STFT, so that the data to be transformed is decomposed into tensors or frames (which are often overlapped to reduce artifacts at the boundaries). Each tensor or frame is Fourier transformed, and the complex results are added to a matrix that records the amplitude and phase of each point in time and frequency. In some examples, the input tensor has a size of H×W×C, where H represents the height of the input tensor (e.g., the number of rows in the input tensor or the number of data elements in a row), W represents the width of the input tensor (e.g., the number of columns in the input tensor or the number of data elements in a row), and C represents the depth of the input tensor (e.g., the number of input channels).

[0053] The inverse STFT can be generated by reversing the STFT. In various examples, the STFT is processed by the DNN and then reversed at the decoder 547, or reversed before being input to the decoder 547. By reversing the STFT, the encoded frequency domain signal from the frequency encoder 545 can be recombined with the encoded time domain signal from the time domain encoder 543. One method of reversing the STFT is by using an overlap-add method, which also allows the STFT complex spectrum to be modified. This forms a general signal processing method called an overlap and add method with modification. In various examples, the output from the decoder 547 is a rotated two-channel time domain audio output signal.

[0054] The data storage 551 stores data received, generated, used, or otherwise associated with the DNN module 501. For example, the data storage 551 stores data sets used by the training module 521 and the verification module 531. The data storage 551 may also store data generated by the training module 521 and the verification module 531, such as hyperparameters used to train the DNN, internal parameters of the trained DNN (e.g., weights, etc.), data used for sparse acceleration (e.g., sparse bitmaps, etc.). In some embodiments, the data storage 551 is a component of the DNN module 501. In other embodiments, the data storage 551 may be located outside the DNN module 501 and communicate with the DNN module 501 via a network.

[0055] Figure 6 is a block diagram illustrating an example of a neural network architecture 600 that can perform a virtual rotation of an input sound position according to various embodiments. The neural network architecture includes: a time-domain binaural input signal 612a, 612b and a frequency-domain binaural input signal 622a, 622b. In some examples, the frequency-domain input signals 622a, 622b are spectra transformed from the time-domain input signals 612a, 612b. In some examples, the time-domain input signals 612a, 612b are transformed into frequency-domain STFT spectra using a short-time Fourier transform.

[0056] Time-domain input signals 612a, 612b are input to a time-domain encoder 610, which includes a plurality of time-domain encoder layers 610a, 610b, 610c, 610d, and 610e. In some examples, the time-domain encoder 610 is a convolutional encoder and includes a convolutional U-Net for time-domain signals. The time-domain encoder 610 receives two channels of time-domain input signals 612a, 612b at a first time-domain encoder convolutional layer 610a. The time-domain encoder 610 also receives an input representing the spherical direction in which the head has rotated. The spherical direction in which the head has rotated can be input to the neural network architecture 600 as Cartesian coordinates 602 (the x, y, and z components of a vector representing the direction in which the head has rotated) or using another coordinate system. In various examples, the vector component inputs (including the Cartesian coordinates 602) are processed by a plurality of fully connected neural network layers (vector component layers 604). In some examples, the vector component layer 604 expands the coordinates 602 into vectors and / or tensors using methods similar to those used by the neural network encoder and / or neural network decoder. In some examples, expanding the input to the neural network can improve the neural network training routine and / or training outcomes. The output 606 of the vector component layer 604 is input to each layer 610a, 610b, 610c, 610d, 610e of the time domain encoder 610. The neural network 600 uses the output 606 from the vector component layer 604 to transform the audio time domain input signals 612a, 612b to produce binaural audio outputs 632a, 632b that match the new perspective of the head.

[0057] The first time-domain encoder convolutional layer 610a processes the two-channel time-domain input signals 612a and 612b and the output of the vector component layer 604, and outputs the 128-channel time-domain output to the second time-domain encoder convolutional layer 610b. The second time-domain encoder convolutional layer 610b receives the 128-channel time-domain signals and the output of the vector component layer 604, and outputs the 256-channel time-domain output to the third time-domain encoder convolutional layer 610c. The third time-domain encoder convolutional layer 610c receives the 256-channel time-domain signals and the output of the vector component layer 604, and outputs the 512-channel time-domain output to the fourth time-domain encoder convolutional layer 610d. The fourth time-domain encoder convolutional layer 610d receives the 512-channel time-domain signals and the output of the vector component layer 604, and outputs the 1024-channel time-domain output to the fifth time-domain encoder convolutional layer 610e. The fifth time-domain encoder convolution layer 610e receives the 1024-channel time-domain signal and the output of the vector component layer 604, and outputs a 2048-channel time-domain output. In some examples, the output from the fifth time-domain encoder convolution layer 610e is the output from the time-domain encoder 610. The output from the time-domain encoder 610 is input to the adder 650.

[0058] Frequency domain input signals 622a, 622b are input to a frequency domain encoder 620, which includes multiple frequency domain encoder layers 620a, 620b, 620c, 620d, and 620e. In some examples, the frequency domain encoder 620 is a convolutional encoder for frequency domain STFT spectra. The frequency domain encoder 620 receives two channels of spectrum (frequency domain input signals 622a, 622b) at a first frequency domain encoder convolutional layer 620a and outputs 128 channels of frequency domain output to a second frequency domain encoder convolutional layer 620b. The second frequency domain encoder convolutional layer 620b receives 128 channels of time domain signals and outputs 256 channels of frequency domain output to a third frequency domain encoder convolutional layer 620c. The third frequency domain encoder convolutional layer 620c receives 256 channels of frequency domain signals and outputs 512 channels of frequency domain output to a fourth frequency domain encoder convolutional layer 620d. The fourth frequency domain encoder convolution layer 620d receives a frequency domain signal of 512 channels and outputs a frequency domain output of 1024 channels to the fifth frequency domain encoder convolution layer 620e. The fifth frequency domain encoder convolution layer 620e receives a frequency domain signal of 1024 channels and outputs a frequency domain output of 2048 channels. In some examples, the output from the fifth frequency domain encoder convolution layer 620e is the output from the frequency domain encoder 620. The output from the frequency domain encoder 620 is input to the adder 650, where the output from the frequency domain encoder 620 is combined with the output from the time domain encoder 610.

[0059] The output from the adder 650 is received by the time-domain decoder 630. The time-domain decoder 630 includes a plurality of time-domain decoder layers 630a, 630b, 630c, 630d, and 630e. In some examples, the time-domain decoder 630 is a convolutional decoder and includes a convolutional U-Net for the time-domain signal. The time-domain decoder 630 also receives an input representing the spherical direction of the head rotation. Specifically, the output 606 of the vector component layer 604 is input to each layer 630a, 630b, 630c, 630d, and 630e of the time-domain decoder 630. The neural network 600 uses the output 606 from the vector component layer 604 to generate binaural audio outputs 632a and 632b that match the new perspective of the head. The time-domain decoder also receives output signals directly from the corresponding layers of the time-domain encoder. In particular, each layer 630a, 630b, 630c, 630d, 630e of the time domain decoder receives the output signal from the output signal of the corresponding layer 610a, 610b, 610c, 610d, 610e that produces the same number of output channels, and the decoder layer 630a, 630b, 630c, 630d, 630e receives the output signal as input.

[0060] The first time-domain decoder convolutional layer 630a processes the 2048-channel time-domain input signal from the adder 650, the 2048-channel time-domain encoder signal from the time-domain encoder layer 610e, and the output 606 of the vector component layer 604, and outputs the 1024-channel time-domain output to the second time-domain decoder convolutional layer 630b. The second time-domain decoder convolutional layer 630b receives the 1024-channel time-domain signal from the first time-domain decoder convolutional layer 630a, the 1024-channel time-domain encoder signal from the time-domain encoder layer 610d, and the output 606 of the vector component layer 604, and outputs the 512-channel time-domain output to the third time-domain decoder convolutional layer 630c. The third time-domain encoder convolution layer 630c receives the 512-channel time-domain signals from the second time-domain decoder convolution layer 630b, the 512-channel time-domain encoder signals from the time-domain encoder layer 610c, and the output 606 of the vector component layer 604, and outputs the 256-channel time-domain output to the fourth time-domain decoder convolution layer 630d. The fourth time-domain decoder convolution layer 630d receives the 256-channel time-domain signals from the third time-domain decoder convolution layer 630c, the 256-channel time-domain encoder signals from the second time-domain encoder layer 610b, and the output 606 of the vector component layer 604, and outputs the 128-channel time-domain output to the fifth time-domain decoder convolution layer 630e. The fifth time domain decoder convolutional layer 630e receives the 128-channel time domain signal from the fourth time domain decoder convolutional layer 630d, the 128-channel time domain signal from the first time domain encoder convolutional layer 610a, and the output 606 of the vector component layer 604, and outputs two channels of audio output 632a, 632b. In some examples, the output from the fifth time domain decoder convolutional layer 630e is the output from the time domain decoder 630, and the output from the decoder 630 is the output from the neural network architecture 600.

[0061] In some examples, the time domain encoder 610 and the frequency domain encoder 620 may have a shared cross-domain bottleneck such that both the time domain encoder 610 output and the frequency domain encoder 620 output are added to an additional encoder layer before reaching the decoder 630 .

[0062] The neural network architecture 600 (with multiple blocks and per-block skip connections) including the time domain encoder 610 and the time domain decoder 630 can be a U-Net. The addition of the frequency domain encoder 620 results in an additional frequency domain encoder 620 output, which is combined with the time domain encoder 610 output at the U-Net bottleneck at the adder 650. Therefore, the neural network architecture 600 is a multi-domain architecture.

[0063] According to various implementations, Figure 6 The neural network architecture 600 shown in FIG is an example of a neural network that can be used for virtual rotation of a signal source position. In various examples, the neural network can have an architecture similar to demucs and / or hybrid demucs. In some examples, the architecture can include a U-Net encoder and / or decoder structure. In some examples, the encoder and decoder can have a symmetrical structure. In some examples, the encoder layer includes convolutions. In one example, the convolution can have a kernel size of eight, a stride of four, a first layer with a fixed number of channels (e.g., 48 or 64), and subsequent layers with double the number of channels. The neural network architecture can include rectified linear units (ReLUs), and the neural network architecture can include 1x1 convolutions with gated linear unit activations. The decoder layer can add contributions from the U-Net skip connection and the previous layer, applying a 1x1 convolution with a GLU. In some examples, the audio input data is 44.1kHz audio. In some examples, the input audio data is upsampled by two before being input to the encoder (to limit aliasing from the outermost layer), while the output from the decoder is downsampled by two.

[0064] In some examples, the neural network is an architecture inspired by hybrid demucics and includes multi-domain analysis and prediction capabilities. The architecture may include a time branch, a spectral branch, and shared layers. The time branch receives a waveform as input and processes the waveform. In some examples, the time branch includes a Gaussian Error Linear Unit (GELU) for activation. In some examples, the time branch includes multiple layers (e.g., five layers), and these layers reduce the number of time steps by a factor of 1024. The spectral branch receives as input a spectrogram generated using an STFT function. The spectrogram is a frequency representation of the waveform input to the time branch. In some examples, the STFT is obtained using 4096 time steps with a hop length of 1024. Therefore, in some examples, the number of time steps in the spectral branch matches the number of time steps in the output of the time branch encoder. In some examples, the spectral branch performs the same convolution as the time branch, but in the frequency dimension. Each layer in the spectral branch reduces the number of frequencies by a factor of four. In some examples, the fifth layer in the spectral branch reduces the number of frequencies by a factor of eight.

[0065] In some examples, the spectrum branch can perform frequency-by-frequency convolution. At each layer of the neural network, the number of frequency bins can be divided by four. In some examples, the last layer has eight frequency bins, which can be reduced to one frequency bin using a convolution with a kernel size of eight and no padding. In some examples, the spectrogram input to the neural network can be represented as an amplitude spectrogram, or can be represented as a complex number. In some examples, the spectrum branch output is converted to a waveform and added to the time branch output, and the output from the adder is in the waveform domain.

[0066] Example method for virtual rotation

[0067] Figure 7 7 is a flow chart of an example method 700 for virtual rotation of a binaural audio signal according to various embodiments. At step 710, a binaural audio input signal is received, the binaural audio input signal comprising a right audio input signal and a left audio input signal. At step 720, the binaural audio signal and a head rotation angle are input to a neural network. In various examples, the neural network is configured to output a virtual rotation of the binaural audio input signal, wherein the virtual rotation comprises: processing the binaural audio input signal to change the perceived source position of a sound in the binaural audio input signal such that the sound position is perceived as unchanged despite a rotation of the head. Thus, in various examples, the sound position is not perceived as rotating with the head even when the listener is wearing headphones.

[0068] At step 730, a virtual rotation angle is determined based on the head rotation angle. As described above, the virtual rotation angle can be opposite to the head rotation angle. At step 740, the two-channel audio input signal is transformed into a two-channel frequency domain signal. In some examples, the signal is transformed using a Fourier transform, and in some examples, the signal is transformed into a two-channel frequency domain signal using an STFT.

[0069] At step 750, a rotated right audio signal and a rotated left audio signal are generated based on the virtual rotation angle, the two-channel audio input signal, and the two-channel frequency domain signal. Figure 5 DNN module 501 or Figure 6A neural network, such as neural network 600, can generate virtually rotated right and left audio signals based on a virtual rotation angle and a two-channel audio input signal. According to various examples, the virtually rotated right and left audio signals are adjusted to change the perceived source positions of various sounds in the right and left audio input signals, such that the perceived source positions of the various sounds are virtually rotated by the virtual rotation angle. At step 760, a two-channel audio output signal rotated by the virtual rotation angle is output from the neural network. The two-channel audio output signal includes the rotated right and left audio signals.

[0070] Example computing device

[0071] Figure 8 is a block diagram of an example computing device 800 according to various embodiments. In some embodiments, the computing device 800 may be used to Figure 5 DNN module 501 in Figure 6 At least part of the neural network 600 in . Several components in Figure 8 800, but any one or more of these components may be omitted or repeated as appropriate for the application. In some embodiments, some or all of the components included in the computing device 800 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system-on-chip (SoC) die. Additionally, in various embodiments, the computing device 800 may not include a processor. Figure 8 800 may include one or more of the components shown in , but the computing device 800 may include interface circuitry for coupling to one or more of the components. For example, the computing device 800 may not include a display device 806, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which the display device 806 may be coupled. In another set of examples, the computing device 800 may not include an audio input device 818 or an audio output device 808, but may include audio input or output device interface circuitry (e.g., a connector and supporting circuitry) to which the audio input device 818 or the audio output device 808 may be coupled.

[0072] The computing device 800 may include a processing device 802 (e.g., one or more processing devices). The processing device 802 processes electronic data from registers and / or memory to transform the electronic data into other electronic data that can be stored in registers and / or memory. The computing device 800 may include a memory 804, which itself may include one or more memory devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM)), high bandwidth memory (HBM), flash memory, solid-state memory, and / or a hard drive. In some embodiments, the memory 804 may include a memory that shares a die with the processing device 802. In some embodiments, the memory 804 includes one or more non-transitory computer-readable media that stores instructions that are executable to perform occupancy mapping, such as described above in conjunction with Figure 7 The method 700 described or by Figure 5 Some operations performed by the DNN system 501 in Figure 6 The instructions stored in one or more non-transitory computer-readable media may be executed by the processing device 802.

[0073] In some embodiments, computing device 800 may include a communication chip 812 (e.g., one or more communication chips). For example, communication chip 812 may be configured to manage wireless communications for transferring data to and from computing device 800. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, and the like that may use modulated electromagnetic radiation to transfer data through a non-solid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they may not.

[0074] The communication chip 812 can implement any of several wireless standards or protocols, including but not limited to Institute of Electrical and Electronics Engineers (IEEE) standards, including Wi-Fi (IEEE 802.10 series), IEEE 802.16 standards (e.g., IEEE 802.16-2005 amendment), Long Term Evolution (LTE) project and any amendments, updates and / or revisions (e.g., LTE-Advanced project, Ultra Mobile Broadband (UMB) project (also known as "3GPP2"), etc.). IEEE 802.16-compliant broadband wireless access (BWA) networks are commonly referred to as WiMAX networks. This abbreviation stands for Worldwide Interoperability for Microwave Access, which is a certification mark for products that have passed IEEE 802.16 standard conformance and interoperability testing. The communication chip 812 can operate according to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. The communication chip 812 can operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). The communication chip 812 can operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and its derivatives, as well as any other wireless protocol designated as 3G, 4G, 5G, and higher. In other embodiments, the communication chip 812 can operate in accordance with other wireless protocols. The computing device 800 may include an antenna 822 for facilitating wireless communications and / or receiving other wireless communications (e.g., AM or FM radio transmissions).

[0075] In some embodiments, the communication chip 812 can manage wired communications, such as electrical, optical, or any other suitable communication protocol (e.g., Ethernet). As described above, the communication chip 812 can include multiple communication chips. For example, the first communication chip 812 can be dedicated to shorter-range wireless communications, such as Wi-Fi or Bluetooth, and the second communication chip 812 can be dedicated to longer-range wireless communications, such as Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or other. In some embodiments, the first communication chip 812 can be dedicated to wireless communications, and the second communication chip 812 can be dedicated to wired communications.

[0076] Computing device 800 may include battery / power circuitry 814. Battery / power circuitry 814 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of computing device 800 to an energy source separate from computing device 800 (e.g., AC line power).

[0077] Computing device 800 may include a display device 806 (or corresponding interface circuitry, as described above). Display device 806 may include any visual indicator, such as a heads-up display, a computer monitor, a projector, a touch screen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat-panel display.

[0078] Computing device 800 may include an audio output device 808 (or corresponding interface circuitry, as described above). Audio output device 808 may include any device that generates an audible indicator, such as a speaker, a headset, or earphones.

[0079] Computing device 800 may include an audio input device 818 (or corresponding interface circuitry, as described above). Audio input device 818 may include any device that generates a signal representing sound, such as a microphone, a microphone array, or a digital musical instrument (e.g., an instrument with a Musical Instrument Digital Interface (MIDI) output).

[0080] Computing device 800 may include a GPS device 816 (or corresponding interface circuitry, as described above). GPS device 816 may communicate with a satellite-based system and may receive the location of computing device 800, as is known in the art.

[0081] Computing device 800 may include another output device 810 (or corresponding interface circuitry, as described above). Examples of other output devices 810 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.

[0082] The computing device 800 may include another input device 820 (or corresponding interface circuitry, as described above). Examples of other input devices 820 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device (e.g., a mouse), a stylus, a touchpad, a barcode reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.

[0083] The computing device 800 can have any desired form factor, such as a handheld or mobile computer system (e.g., a mobile phone, a smartphone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a personal digital assistant (PDA), an ultra-mobile personal computer, etc.), a desktop computer system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computer system. In some embodiments, the computing device 800 can be any other electronic device that processes data.

[0084] Selected Examples

[0085] The following paragraphs provide various examples of the embodiments disclosed herein.

[0086] Example 1 provides a computer-implemented method, comprising: receiving a two-channel audio input signal, the two-channel audio input signal including a right audio input signal and a left audio input signal; inputting the two-channel audio signal and a head rotation angle into a neural network; determining a virtual rotation angle at the neural network based on the head rotation angle; transforming the two-channel audio input signal into a two-channel frequency domain signal at the neural network; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the two-channel audio input signal, and the two-channel frequency domain signal; and outputting, by the neural network, a two-channel audio output signal rotated by the virtual angle, the two-channel audio output signal rotated by the virtual angle including a rotated right audio signal and a rotated left audio signal.

[0087] Example 2 provides the computer-implemented method of Example 1, wherein the neural network includes a time-domain encoder, the time-domain encoder includes a plurality of time-domain encoder layers, and the computer-implemented method further comprises: inputting a right audio input signal, a left audio input signal, and a virtual rotation angle to a first time-domain encoder layer of the plurality of time-domain encoder layers, and outputting a plurality of time-domain encoded signals from the time-domain encoder.

[0088] Example 3 provides the computer-implemented method of Example 2, wherein a second time-domain encoder layer of the plurality of time-domain encoder layers receives the output from the first time-domain encoder layer and the virtual rotation angle.

[0089] Example 4 provides a computer-implemented method as described in Example 3, wherein the number of channels output from each of the multiple time-domain encoder layers is greater than the number of channels input to each of the multiple time-domain encoder layers.

[0090] Example 5 provides the computer-implemented method of Example 2, wherein the neural network includes a frequency domain encoder, the frequency domain encoder includes a plurality of frequency domain encoder layers, and the computer-implemented method further comprises: inputting a binaural frequency domain signal to a first frequency domain encoder layer of the plurality of frequency domain encoder layers, and outputting a plurality of frequency domain encoded signals from the frequency domain encoder.

[0091] Example 6 provides the computer-implemented method of Example 5, wherein the neural network includes an adder, and the computer-implemented method further comprises: adding the plurality of time-domain encoded signals and the plurality of frequency-domain encoded signals to generate a plurality of added encoded signals, and inputting the plurality of added encoded signals to a time-domain decoder.

[0092] Example 7 provides the computer-implemented method of Example 1, wherein the head rotation angle comprises a Cartesian coordinate representing a direction in which the head is rotated.

[0093] Example 8 provides the computer-implemented method of Example 1, further comprising: training the neural network using synthetic audio samples, the synthetic audio samples generated using a head rotation transfer function.

[0094] Example 9 provides a non-transitory computer-readable medium storing one or more instructions, which are executable to perform operations, the operations including: receiving a two-channel audio input signal, the two-channel audio input signal including a right audio input signal and a left audio input signal; inputting the two-channel audio signal and a head rotation angle into a neural network; determining a virtual rotation angle at the neural network based on the head rotation angle; transforming the two-channel audio input signal into a two-channel frequency domain signal at the neural network; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the two-channel audio input signal, and the two-channel frequency domain signal; and outputting, by the neural network, a two-channel audio output signal rotated by the virtual angle, the two-channel audio output signal rotated by the virtual angle including a rotated right audio signal and a rotated left audio signal.

[0095] Example 10 provides the one or more non-transitory computer-readable media of Example 9, wherein the neural network includes a time-domain encoder, the time-domain encoder including a plurality of time-domain encoder layers, and the operation further comprises: inputting a right audio input signal, a left audio input signal, and a virtual rotation angle to a first time-domain encoder layer of the plurality of time-domain encoder layers, and outputting a plurality of time-domain encoded signals from the time-domain encoder.

[0096] Example 11 provides one or more non-transitory computer-readable media of Example 10, wherein the neural network includes a frequency domain encoder, the frequency domain encoder including a plurality of frequency domain encoder layers, and the operation further includes: inputting a binaural frequency domain signal to a first frequency domain encoder layer of the plurality of frequency domain encoder layers; and outputting a plurality of frequency domain encoded signals from the frequency domain encoder.

[0097] Example 12 provides the one or more non-transitory computer-readable media of Example 11, wherein the neural network includes an adder, and the operations further include: adding the plurality of time-domain encoded signals and the plurality of frequency-domain encoded signals to generate a plurality of added encoded signals; and inputting the plurality of added encoded signals to a time-domain decoder.

[0098] Example 13 provides the one or more non-transitory computer-readable media of Example 9, wherein the head rotation angle comprises a Cartesian coordinate representing a direction in which the head is rotated.

[0099] Example 14 provides the one or more non-transitory computer-readable media of Example 9, wherein the operations further comprise training the neural network using synthetic audio samples, the synthetic audio samples being generated using a head rotation transfer function.

[0100] Example 15 provides an apparatus comprising: a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions, the computer program instructions being executable by the computer processor to perform operations comprising: receiving a two-channel audio input signal, the two-channel audio input signal comprising a right audio input signal and a left audio input signal; inputting the two-channel audio signal and a head rotation angle into a neural network; determining, at the neural network, a virtual rotation angle based on the head rotation angle; transforming, at the neural network, the two-channel audio input signal into a two-channel frequency domain signal; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the two-channel audio input signal, and the two-channel frequency domain signal; and outputting, by the neural network, a two-channel audio output signal rotated by the virtual angle, the two-channel audio output signal rotated by the virtual angle comprising a rotated right audio signal and a rotated left audio signal.

[0101] Example 16 provides the apparatus of Example 15, wherein the neural network includes a time domain encoder, the time domain encoder including a plurality of time domain encoder layers, and the operation further comprises: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle to a first time domain encoder layer of the plurality of time domain encoder layers, and outputting a plurality of time domain encoded signals from the time domain encoder.

[0102] Example 17 provides the apparatus of Example 16, wherein the neural network includes a frequency domain encoder, the frequency domain encoder including a plurality of frequency domain encoder layers, and the operation further includes: inputting a binaural frequency domain signal to a first frequency domain encoder layer of the plurality of frequency domain encoder layers; and outputting a plurality of frequency domain encoded signals from the frequency domain encoder.

[0103] Example 18 provides the apparatus of Example 17, wherein the neural network comprises: an adder, and the operation further comprises: adding the plurality of time-domain encoded signals and the plurality of frequency-domain encoded signals to generate a plurality of added encoded signals; and inputting the plurality of added encoded signals to the time-domain decoder.

[0104] Example 19 provides the apparatus of Example 15, wherein the head rotation angle comprises a Cartesian coordinate representing the direction in which the head is rotated.

[0105] Example 20 provides the apparatus of Example 15, wherein the operation further comprises: training the neural network using the synthesized audio samples, wherein the synthesized audio samples are generated using a head rotation transfer function.

[0106] The above description of the illustrative implementation of the present disclosure (including the content described in the abstract) is not intended to be exhaustive or to limit the present disclosure to the disclosed precise form. Although the specific implementation and examples of the present disclosure are described herein for illustrative purposes, those skilled in the relevant art will recognize that various equivalent modifications can be made within the scope of the present disclosure. These modifications can be made to the present disclosure based on the above detailed description.

Claims

1. A computer-implemented method comprising: receiving a two-channel audio input signal, wherein the two-channel audio input signal includes a right audio input signal and a left audio input signal; Inputting the binaural audio input signal and the head rotation angle into a neural network; determining, at the neural network, a virtual rotation angle based on the head rotation angle; At the neural network, transforming the two-channel audio input signal into a two-channel frequency domain signal; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; as well as The neural network outputs a two-channel audio output signal rotated by the virtual rotation angle, and the two-channel audio output signal rotated by the virtual rotation angle includes the rotated right audio signal and the rotated left audio signal.

2. The computer-implemented method of claim 1 , wherein: The neural network includes a time-domain encoder, the time-domain encoder including a plurality of time-domain encoder layers, and the computer-implemented method further includes: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle to a first time domain encoder layer among the plurality of time domain encoder layers; and A plurality of time-domain encoded signals are output from the time-domain encoder.

3. The computer-implemented method of claim 2, wherein: A second time-domain encoder layer among the plurality of time-domain encoder layers receives the output from the first time-domain encoder layer and the virtual rotation angle.

4. The computer-implemented method according to any one of claims 2 to 3, wherein: The number of channels output from each of the plurality of time-domain encoder layers is greater than the number of channels input to each of the plurality of time-domain encoder layers.

5. The computer-implemented method according to any one of claims 2-3, wherein: The time domain encoder is a convolutional encoder.

6. The computer-implemented method of claim 2, wherein: The neural network includes a frequency domain encoder, the frequency domain encoder including a plurality of frequency domain encoder layers, and the computer-implemented method further includes: Inputting the two-channel frequency domain signal to a first frequency domain encoder layer among the plurality of frequency domain encoder layers; and A plurality of frequency-domain encoded signals are output from the frequency-domain encoder.

7. The computer-implemented method of claim 6, wherein: The neural network includes an adder, and the computer-implemented method further includes adding the plurality of time-domain encoded signals and the plurality of frequency-domain encoded signals to generate a plurality of added encoded signals, and inputting the plurality of added encoded signals to a time-domain decoder.

8. The computer-implemented method of any one of claims 1-3, 6-7, wherein: The head rotation angle includes Cartesian coordinates, and the Cartesian coordinates represent the direction in which the head is rotated.

9. The computer-implemented method of any one of claims 1-3, 6-7, further comprising training the neural network using synthetic audio samples, the synthetic audio samples generated using a head rotation transfer function.

10. One or more non-transitory computer-readable media storing instructions executable to perform operations comprising: receiving a two-channel audio input signal, wherein the two-channel audio input signal includes a right audio input signal and a left audio input signal; Inputting the binaural audio input signal and the head rotation angle into a neural network; determining, at the neural network, a virtual rotation angle based on the head rotation angle; At the neural network, transforming the two-channel audio input signal into a two-channel frequency domain signal; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; as well as The neural network outputs a two-channel audio output signal rotated by the virtual rotation angle, and the two-channel audio output signal rotated by the virtual rotation angle includes the rotated right audio signal and the rotated left audio signal.

11. The one or more non-transitory computer-readable media of claim 10, wherein: The neural network includes a time-domain encoder, the time-domain encoder includes a plurality of time-domain encoder layers, and the operations further include: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle to a first time domain encoder layer among the plurality of time domain encoder layers; and A plurality of time-domain encoded signals are output from the time-domain encoder.

12. The one or more non-transitory computer-readable media of claim 11, wherein: A second time-domain encoder layer among the plurality of time-domain encoder layers receives the output from the first time-domain encoder layer and the virtual rotation angle.

13. One or more non-transitory computer-readable media according to any one of claims 10-12, wherein: The number of channels output from each of the plurality of time-domain encoder layers is greater than the number of channels input to each of the plurality of time-domain encoder layers.

14. The one or more non-transitory computer-readable media of claim 11, wherein: The neural network includes a frequency domain encoder, the frequency domain encoder includes a plurality of frequency domain encoder layers, and the operations further include: Inputting the binaural frequency domain signal to a first frequency domain encoder layer among the plurality of frequency domain encoder layers; and A plurality of frequency-domain encoded signals are output from the frequency-domain encoder.

15. The one or more non-transitory computer-readable media of claim 14, wherein: The neural network includes an adder, and the operations further include: adding the plurality of time-domain coded signals and the plurality of frequency-domain coded signals to generate a plurality of added coded signals; and The plurality of added encoded signals are input to a time domain decoder.

16. One or more non-transitory computer-readable media according to any one of claims 10-12, 14-15, wherein: The head rotation angle includes Cartesian coordinates, and the Cartesian coordinates represent the direction in which the head is rotated.

17. The one or more non-transitory computer-readable media of any one of claims 10-12, 14-15, the operations further comprising training the neural network using synthetic audio samples, the synthetic audio samples generated using a head rotation transfer function.

18. An apparatus comprising: a computer processor for executing computer program instructions; as well as a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including: receiving a two-channel audio input signal, wherein the two-channel audio input signal includes a right audio input signal and a left audio input signal; Inputting the binaural audio input signal and the head rotation angle into a neural network; determining, at the neural network, a virtual rotation angle based on the head rotation angle; At the neural network, transforming the two-channel audio input signal into a two-channel frequency domain signal; generating, by the neural network, a rotated right audio signal and a rotated left audio signal based on the virtual rotation angle, the binaural audio input signal, and the binaural frequency domain signal; as well as The neural network outputs a two-channel audio output signal rotated by the virtual rotation angle, and the two-channel audio output signal rotated by the virtual rotation angle includes the rotated right audio signal and the rotated left audio signal.

19. The device according to claim 18, wherein The neural network includes a time-domain encoder, the time-domain encoder includes a plurality of time-domain encoder layers, and the operations further include: inputting the right audio input signal, the left audio input signal, and the virtual rotation angle to a first time domain encoder layer among the plurality of time domain encoder layers, and A plurality of time-domain encoded signals are output from the time-domain encoder.

20. The device according to claim 19, wherein A second time-domain encoder layer among the plurality of time-domain encoder layers receives the output from the first time-domain encoder layer and the virtual rotation angle.

21. The device according to any one of claims 19-20, wherein The number of channels output from each of the plurality of time-domain encoder layers is greater than the number of channels input to each of the plurality of time-domain encoder layers.

22. The device according to any one of claims 19-20, wherein The neural network includes a frequency domain encoder, the frequency domain encoder includes a plurality of frequency domain encoder layers, and the operations further include: Inputting the binaural frequency domain signal to a first frequency domain encoder layer among the plurality of frequency domain encoder layers; and A plurality of frequency-domain encoded signals are output from the frequency-domain encoder.

23. The device according to claim 22, wherein The neural network includes an adder, and the operations further include: adding the plurality of time-domain coded signals and the plurality of frequency-domain coded signals to generate a plurality of added coded signals; and The plurality of added encoded signals are input to a time domain decoder.

24. The device according to any one of claims 18 to 20 or 23, wherein: The head rotation angle includes Cartesian coordinates, and the Cartesian coordinates represent the direction in which the head is rotated.

25. The apparatus of any one of claims 18-20, 23, the operations further comprising training the neural network using synthetic audio samples, the synthetic audio samples generated using a head rotation transfer function.