Multi-dimensional acoustic crosstalk cancellation filter interpolation
By determining the user location and orientation in the audio processing system, and generating crosstalk cancellation filters using dimension diagrams and transfer functions, the crosstalk problem caused by user movement in three-dimensional space is solved, real-time compensation and calculation efficiency are improved.
Patent Information
- Application Number
- CN202510008156.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional audio processing systems are difficult to effectively reduce crosstalk in three-dimensional space, especially when the user's head moves over six degrees of freedom, computing resource demand is high and the effect is not good.
By determining the user's position and orientation in the environment, using the weights of a subset of nearest points and transfer functions in the dimension graph, a multi-dimensional acoustic crosstalk cancellation filter is generated to compensate the user's movement in real time and reduce spectral distortion.
Real-time compensation for user movements over six degrees of freedom is achieved, reducing crosstalk and improving the accuracy of audio signals in the left and right ears, and reducing computational costs.
Smart Images

Figure CN120264194A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to audio reproduction, and more particularly, to interpolation of multi-dimensional acoustic crosstalk cancellation filters. Background Art
[0002] An audio processing system uses one or more speakers to generate sound in a given space. The one or more speakers generate a sound field, where a user located in the environment receives the sound included in the sound field. The one or more speakers reproduce sound based on an input signal that typically includes at least two channels, such as a left channel and a right channel. The left channel is intended to be received by the user's left ear, and the right channel is intended to be received by the user's right ear. A binaural rendering algorithm that uses one or more speakers to generate sound relies on a crosstalk cancellation algorithm to ensure that a signal intended for the left ear is received by the left ear without being interfered with by other signals intended for the right ear, and vice versa. To this end, traditional crosstalk cancellation algorithms attempt to filter out interfering signals by characterizing the audio transmission path from the speaker to the entrance of the user's ear canal based on measurements made on a user located at a specific azimuth.
[0003] At least one drawback of traditional crosstalk cancellation techniques is that the traditional techniques are highly focused on working for a specific point in three-dimensional space and will fail if the user moves or rotates their head. Other traditional techniques attempt to compensate for potential lateral displacement of the head in one or two directions. However, traditional crosstalk cancellation techniques have difficulty addressing the actual movement of the head in three-dimensional space, which produces six degrees of freedom ( For example , moving along the x-axis, y-axis, and z-axis (also referred to as forward / backward, left / right, up / down) and rotating along the x-axis, y-axis, and axis (also referred to as pitch, yaw, and roll)). For example, when a user moves in multiple directions, traditional crosstalk cancellation techniques may degrade in effectiveness or, in some cases, result in increased interference. Additionally, for each additional degree of freedom, the computational resources required to cover each degree of freedom increase exponentially. Traditional crosstalk cancellation techniques do not have the computational resources required to cover all six degrees of freedom. Therefore, traditional techniques for reducing crosstalk when playing back audio in three-dimensional space are not sufficient to handle the full range of user movement.
[0004] As shown by the foregoing, the present technology requires more effective techniques for reducing crosstalk when generating sound received by a user in three-dimensional space in an environment. Summary of the Invention
[0005] Various embodiments disclose a computer-implemented method that includes: determining a user's location and orientation in an environment; determining a subgroup of nearest points in a dimensionality map based on the user's location and orientation, where the dimensionality map includes a set of points in a multi-dimensional space, each point associated with a corresponding transfer function, and in the dimensionality map, the points in the subgroup of nearest points are closer to the user's location and orientation in the dimensionality map than other points; determining a respective weight for each transfer function associated with the subgroup of nearest points based on the distance between each point in the subgroup of nearest points and the user's location and orientation in the dimensionality map; determining at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; generating a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; and transmitting the plurality of audio signals to the plurality of loudspeakers for output.
[0006] In addition, other embodiments provide one or more non-transitory computer-readable media and systems configured to implement the methods recited above.
[0007] At least one technical advantage of the disclosed technology over the prior art is that through the disclosed technology, an audio processing system can form an improved crosstalk cancellation filter by compensating for a user's movement in six degrees of freedom in real time. Additionally, spectral distortion caused by user movement is reduced with a reduced computational cost. Additionally, audio intended to be received by a user's left and right ears respectively is more accurately represented in the audio output of the audio processing and playback system. These technical advantages provide one or more technological advancements over prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to enable a more particular understanding of the manner in which the above-recited features of various embodiments can be obtained, the inventive concepts briefly summarized above may be described in more detail by reference to various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only illustrate typical embodiments of the inventive concepts and should therefore in no way be considered limiting in scope, and there are other equally effective embodiments.
[0009] Figure 1 is a schematic diagram showing an audio processing system according to one or more embodiments.
[0010] Figure 2 illustrates an example of how a listener observes crosstalk from an input signal generated by one or more speakers according to one or more embodiments.
[0011] Figure 3Shows an example of triangulation used during crosstalk cancellation based on the observed position and orientation of a listener in three-dimensional space according to one or more embodiments.
[0012] Figure 4 Shows an example of a filter that performs crosstalk cancellation based on the observed position and orientation of a listener in three-dimensional space according to one or more embodiments.
[0013] Figure 5 Shows a flowchart of method steps for combining transfer functions for configuring a filter that performs crosstalk cancellation according to one or more embodiments. Detailed Description
[0014] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, those skilled in the art will appreciate that the inventive concepts may be practiced without one or more of these specific details.
[0015] Figure 1 Is a schematic diagram showing an audio processing system 100 according to various embodiments. As shown, the audio processing system 100 includes, but is not limited to, a computing device 110, an audio source 140, one or more sensors 150, and one or more speakers 160. The computing device 110 includes, but is not limited to, a processing unit 112 and a memory 114. The memory 114 stores, but is not limited to, a crosstalk cancellation application 120, a transfer function 132, a dimensional map 134, and one or more filters 138.
[0016] In operation, the audio processing system 100 processes sensor data from one or more sensors 150 to track the orientation of one or more listeners within a listening environment. One or more sensors 150 track the position of a listener's head in three-dimensional space as well as the pitch, yaw, and roll of the head, which are used to respectively locate the relative orientation of the listener's left and right ears. Based on the position and / or orientation of the head within the three-dimensional environment, the crosstalk cancellation application 120 selects one or more transfer functions 132 for one or more filters 138, which are used to process the audio source 140 for playback by one or more speakers 160 associated with the audio processing system 100. Additionally, if the position of the listener's head in three-dimensional space changes during playback of the audio source 140, the crosstalk cancellation application 120 selects different transfer functions 132 and potentially different filters 138 for processing the audio source 140 for playback via one or more speakers 160.
[0017] The computing device 110 is a device that drives the speaker 160 to partially generate a sound field for a listener by playing back an audio source 140. In various embodiments, the computing device 110 is an audio processing unit in a home theater system, a soundbar, a vehicle system, etc. In some embodiments, the computing device 110 is included in one or more devices, such as consumer products ( For example , portable speakers, gaming products, etc.), vehicles ( For example , head units of sedans, trucks, minivans, etc.), smart home devices ( For example , smart lighting systems, security systems, digital assistants, etc.), communication systems ( For example , conference call systems, video conferencing systems, speaker amplification systems, etc.). In various embodiments, the computing device 110 is located in various environments, including but not limited to indoor environments ( For example , living rooms, conference rooms, convention halls, home offices, etc.)
[0018] and / or outdoor environments ( For example , patios, rooftops, gardens, etc.).
[0019] The processing unit 112 can be any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), and / or any other type of processing unit or combination of processing units, such as a CPU configured to operate in combination with a GPU. Generally, the processing unit 112 can be any technically feasible hardware unit capable of processing data and / or executing software applications.
[0020] The memory 114 may include random access memory (RAM) modules, flash memory cells, or any other type of memory cells or a combination thereof. The processing unit 112 is configured to read data from the memory 114 and write data to the memory. In various embodiments, the memory 114 includes non-volatile memory, such as optical drives, magnetic drives, flash drives, or other storage devices. In some embodiments, a separate data storage device included in a network (such as an external data storage device (“cloud storage”)) may supplement the memory 114. The processing unit 112 may execute the crosstalk cancellation application 120 within the memory 114 to implement the overall functions of the computing device 110 and thus coordinate the operation of the audio processing system 100 as a whole. In various embodiments, an interconnect bus (not shown) connects the processing unit 112, the memory 114, the speaker 160, the sensor 150, and any other components of the computing device 110.
[0021] The crosstalk cancellation application 120 determines the orientation of the listener within the listening environment and selects parameters (such as one or more transfer functions 132) of one or more filters 138 to generate a sound field for the orientation of the listener. The transfer function 132 is selected to minimize or remove crosstalk. The transfer function 132 causes the filter 138 to generate audio in the sound field such that the left channel is perceived by the left ear of the listener and the crosstalk from the right channel is minimized. Similarly, the transfer function 132 causes the filter 138 to generate audio in the sound field such that the right channel is perceived by the right ear of the listener and the crosstalk from the left channel is minimized. In various embodiments, the crosstalk cancellation application utilizes sensor data from the sensor 150 to identify the position of the listener and, in particular, the position of the listener's head. Based on the position and orientation of the listener, the crosstalk cancellation application 120 selects the appropriate filter 138 and transfer function 132 for processing the audio source 140 for playback. In some embodiments, the crosstalk cancellation application 120 sets the parameters of a plurality of filters 138 corresponding to a plurality of speakers 160. For example, a first transfer function 132 may be used for a first filter 138 that is used for audio played back by a first speaker 160; and a second transfer function 132 is utilized by a second filter 138 that is used for audio played back by a second speaker 160. In other embodiments, a filter network is utilized such that signals for driving each speaker 160 are passed through a network of a plurality of filters. Additionally or alternatively, the crosstalk cancellation application 120 tracks the positions and orientations of multiple listeners.
[0022] The filter 138 includes one or more filters that modify the input audio source 140. In various embodiments, a given filter 138 modifies the input audio signal by modifying the energy within a particular frequency range, adding directional information, and the like. For example, the filter 138 may include a plurality of filter parameters, such as a set of values that modify the operating characteristics ( For example , center frequency, gain, Q factor, cutoff frequency, etc.) of the filter 138. In some embodiments, the filter parameters include one or more digital signal processing (DSP) coefficients that manipulate the generated sound waves in a particular direction. In such instances, the generated filtered audio signal is used to generate sound waves in the direction specified in the filtered audio signal. For example, one or more speakers 160 reproduce audio using one or more filtered audio signals to generate a sound field. In some embodiments, the crosstalk cancellation application 120 sets separate filter parameters, such as selecting different transfer functions 132 for separate filters 138 for different speakers 160. In such instances, one or more speakers 160 use the separate filters 138 to generate a sound field. For example, each filter 138 may generate a filtered audio signal for a single speaker 160 within the listening environment.
[0023] The transfer function 132 includes one or more transfer functions that are used to configure one or more filters 138 selected by the crosstalk cancellation application 120 to process an input signal (such as a channel of the audio source 140) to produce an output signal for driving the speaker 160. Different transfer functions 132 are utilized based on the position and orientation of the listener in three-dimensional space.
[0024] In some embodiments, the dimensional map 134 maps a given position within the three-dimensional space (such as the interior of a vehicle) to filter parameters of one or more filters 138 (such as one or more finite impulse response (FIR) filters). In various embodiments, the crosstalk cancellation application 120 determines the position and orientation of the listener based on data from the sensors 150 and identifies the transfer function 132 or other filter parameters of the filter 138 corresponding to each speaker 160. Then, when the listener's head moves, the crosstalk cancellation application 120 updates the filter parameters for a particular speaker ( For example , the first filter 138(1) for the first speaker 160(1)). For example, the crosstalk cancellation application 120 may first generate filter parameters for a set of filters 138. After determining that the listener's head has moved to a new position or orientation, the crosstalk cancellation application 120 then immediately determines whether any of the speakers 160 requires an update to the corresponding filter 138. The crosstalk cancellation application 120 updates the filter parameters of any filter 138 that requires an update. In some embodiments, the crosstalk cancellation application 120 generates each of the filters 138 independently. For example, after determining that the listener has moved, the crosstalk cancellation application 120 may update the filter 138 ( For example , the filter parameters for a particular speaker 160 ( For example , 138(1) for 160(1)). Alternatively, the crosstalk cancellation application 120 updates multiple filters 138.
[0025] The dimensional map 134 includes a plurality of points representing positions and orientations in the three-dimensional space ( For example, a point in a six - dimensional space identified by x, y, and z position coordinates and three orientations of roll, pitch, and yaw). The dimensional map 134 maps positions relative to a reference position in a given environment. The dimensional map 134 further maps orientations relative to a reference orientation in the environment. The dimensional map 134 can be generated by making acoustic measurements in three - dimensional space for filter parameters (such as transfer function 132) that minimize or remove crosstalk. Then, the dimensional map 134 is stored on the audio processing system 100 and used to configure the filter 138 utilized by the computing device 110 to minimize or remove crosstalk during playback of the audio source 140. In some embodiments, the dimensional map 134 includes specific coordinates relative to a reference point. For example, the dimensional map 134 can store the potential positions and orientations of the listener's head as distances and angles from a specific reference point. In some embodiments, the dimensional map 134 can include additional orientation information, such as pitch, yaw, and roll that characterize the orientation of the listener's head. The dimensional map 134 can also include a set of angles relative to the normal orientation of the listener's head ( For example , ). In such instances, the corresponding positions and orientations defined by the points in the dimensional map 134 are associated with one or more transfer functions 132 for the filter 138. In one example, the dimensional map 134 is structured as a set of points, each of which is associated with a specific position and orientation in the environment. Each of the points is associated with one or more filters 138 and / or transfer functions 132 that can be used in each of the speakers 160 to reduce or remove crosstalk.
[0026] The crosstalk cancellation application 120 selects a transfer function 132 to configure the filter 138, where the transfer function 132 is identified by the dimensional map 134. The transfer function 132 is used to configure the filter 138 that processes the audio source 140. The transfer function 132 is identified based on a mathematical distance (such as the centroid distance) between a set of points in the dimensional map 134 that represent the position and orientation of the listener's head and one or more points from the set of points. In one example, a given position and orientation of the user are characterized by coordinates in a six - dimensional space. In some embodiments, a graph search algorithm (such as Delaunay triangulation) is then used to identify a subset of points in the dimensional map 134 that are closest to the user's coordinates. The weights 180 of each transfer function 132 associated with the subset of closest points are interpolated based on any technically feasible multi - dimensional distance calculation (such as centroid distance or Euclidean distance) between the subset of closest points and the user's coordinates. The weights 180 and the associated transfer functions 132 associated with a subset of closest points in the dimensional map 134 are used in combination to configure the filter 138, which filters the played - back audio signal.
[0027] As another example, a simplified method for identifying the transfer function 132 includes reducing the dimensions of the user's position and orientation considered when identifying a set of transfer functions defined by the dimensionality map 134. As noted above, the dimensionality map 134 includes a set of points in a six-dimensional space to illustrate three parameters representing position and three parameters representing orientation. To reduce mathematical complexity, a reduced set of parameters representing the user's position and orientation can be considered. For example, one or more of the parameters representing orientation can be removed, and a set of nearest points can be identified based on the mathematical distance in the dimensionality map 134 from the coordinates representing the position and orientation of the user's head to one or more points from the set of points. Examples of coordinates that can be removed include yaw angle, pitch angle, and / or roll angle. In one scenario, only the position and yaw angle of the user's head are considered, reducing the complexity to considering four dimensions. As another example, only the position of the user's head and the yaw and pitch angles are considered, reducing the complexity to five dimensions.
[0028] As another example, an alternative simplified method for identifying the transfer function 132 includes reducing the dimensionality of the dimensionality map 134. As noted above, the dimensionality map 134 includes a set of points in a six-dimensional space to illustrate three parameters representing position and three parameters representing orientation. To reduce mathematical complexity, a dimensionality map 134 can be generated and utilized that includes a set of points mapped in a three-dimensional, four-dimensional, or five-dimensional space. For example, the dimensionality map 134 can map only the position of the user's head in three-dimensional space and the yaw angle representing orientation, thereby forming a four-dimensional map. As another example, the dimensionality map 134 maps only the position of the user's head and two parameters representing orientation, reducing the complexity of the dimensionality map 134 to five dimensions.
[0029] Another example of a simplified method for reducing the dimensionality of the dimensionality map 134 is to use multiple dimensionality maps 134, the multiple dimensionality maps including three dimensions representing position in three-dimensional space. Each of the three-dimensional maps is associated with a particular orientation parameter or a series of orientation parameters. For example, each of the three-dimensional maps is associated with a yaw angle or a series of yaw angles. In one scenario, the first three-dimensional map is associated with a yaw angle from 0 degrees to 10 degrees, the second three-dimensional map is associated with a yaw angle greater than 10 degrees to 20 degrees, and so on. By this method, based on the detected yaw angle of the user's head, a three-dimensional map is selected. Then, based on the coordinates of the detected position of the user, a subset of points corresponding to the nearest transfer function 132 within the three-dimensional map is identified, the weight of each transfer function is interpolated based on the centroid distance or Euclidean distance to the detected position of the user, and the weighted transfer function 132 is used to configure the filter 138.
[0030] The sensor 150 includes various types of sensors that acquire data regarding the listening environment. For example, the computing device 110 can include means for receiving several types of sounds ( For example, an auditory sensor for subsonic pulses, ultrasonic waves, voice commands, etc.). In some embodiments, sensor 150 includes other types of sensors. Other types of sensors include optical sensors (such as RGB cameras, time-of-flight cameras, infrared cameras, depth cameras, quick response (QR) code tracking systems), motion sensors (such as accelerometers or inertial measurement units (IMUs)) ( For example , triaxial accelerometers, gyroscopic sensors, and / or magnetometers), pressure sensors, etc. Additionally, in some embodiments, sensor 150 may include wireless sensors, and the wireless sensors include radio frequency (RF) sensors ( For example , sonar and radar); and / or wireless communication protocols, and the wireless communication protocols include Bluetooth, Bluetooth Low Energy (BLE), cellular protocols, and / or near field communication (NFC). In various embodiments, crosstalk cancellation application 120 uses the sensor data acquired by sensor 150 to identify transfer function 132 for filter 138. For example, computing device 110 includes one or more transmitters that emit positioning signals, where computing device 110 includes a detector that generates auditory data including the positioning signals. In some embodiments, crosstalk cancellation application 120 combines multiple types of sensor data. For example, crosstalk cancellation application 120 may combine auditory data with optical data ( For example , camera images or infrared data) to determine the position and orientation of the listener at a given time.
[0031] Figure 2 An example showing how a user observes crosstalk from the input signals generated by one or more speakers 160. When audio source 140 is played back through one or more speakers 160, crosstalk exists in the audio measured at the left ear L and right ear R of listener 202. Without crosstalk cancellation, crosstalk will necessarily occur when the speakers are far from listener 202. Audio source 140a represents the signal desired by the left ear of listener 202 or the left channel of audio source 140. Audio source 140b represents the signal desired by the right ear of listener 202 or the right channel of audio source 140. When audio is played back in an environment, such as through speakers 160 that are far from the ears of listener 202, crosstalk occurs. C 1,1 and C 1,2 represent functions of how the environment affects the audio source when audio source 140a is played back through audio processing system 100. S1 and S2 represent the corresponding parts of audio source 140a that are heard by the left ear and right ear of listener 202 respectively. For example, when audio source 140a is played through the corresponding one or more speakers 160, the environment is according to C 1,1Modify the audio source 140a such that the audio S1 reaches the left ear of the listener 202. Similarly, the environment modifies the audio source 140a according to C1,2 such that the audio S2 reaches the right ear of the listener 202. S2 represents the portion of the audio source 140a that causes crosstalk reaching the right ear of the listener 202. C 2,1 and C 2,2 represent functions that characterize how the environment affects the audio source when the audio source 140b is played back. The audio S3 and S4 respectively represent the corresponding portions of the audio source 140b heard by the left and right ears of the listener 202. For example, when the audio source 140b is played through a corresponding one or more speakers 160, the environment modifies the audio source 140b according to C 2,2 such that the audio S4 reaches the right ear of the listener 202. Similarly, the environment modifies the audio source 140b according to C 2,1 such that the audio S3 reaches the left ear of the listener 202. S3 represents a portion of the audio source 140b. Accordingly, embodiments of the present disclosure utilize the filter 138 to process the signal and then drive one or more speakers 160 with the signal to reduce or remove crosstalk caused by the environment.
[0032] Figure 3 Shows examples of triangulation used during crosstalk cancellation based on the observed position and orientation of a listener in three-dimensional space according to various embodiments of the present disclosure. As Figure 3 shown, a dimensional graph 300 is shown in three dimensions ( For example , the x-dimension, the y-dimension, and the z-dimension), the dimensional graph including a set of points representing different transfer functions (such as the transfer function 132) that effectively minimize or remove crosstalk at a particular position and / or orientation in the environment. For illustrative purposes, Figure 3 the depiction of the dimensional graph 300 in For example shows only a portion of the dimensional graph 300 that includes a subgroup of points (such as points A, B, C, D, and the position of the listener 302) in three dimensions, but is in no way meant to be restrictive. For example, the dimensional graph 300 may include a very large number ( For example , hundreds, thousands, etc.) of points in six dimensions ( For example, Delaunay triangulation) forms polygonal spaces such that the circum-hypersphere of each polygonal space does not contain any other points in the dimensional map 300 and each polygonal space does not overlap. For example, points A, B, C, and D form a tetrahedron, and other points (not shown) in the dimensional map 300 are not inside the tetrahedron. In some embodiments, the triangulated polygonal spaces can be formed by five points, six points, seven points, or any other technically feasible number of points that form the polygonal spaces. Figure 3 The portion of the dimensional map 300 shown in For example , thousands, hundreds, etc.) of tetrahedrons, and each tetrahedron does not overlap with other tetrahedrons. In some embodiments, the triangulated polygonal spaces can be calculated before crosstalk cancellation begins and the triangulated polygonal spaces can be stored in a memory. In some embodiments, two different three-dimensional maps can be utilized instead of a single dimensional map represented in four or more dimensions.
[0033] The crosstalk cancellation application 120 utilizes the dimensional map 300 to identify the transfer function 132 or other filter parameters of the filter 138 corresponding to each speaker 160 based on the position and / or orientation of the listener determined via the sensor 150. For example, if the position and / or orientation of the listener matches the position and / or orientation of point A in the dimensional map 300, the crosstalk cancellation application 120 will identify or select the transfer function associated with point A to be used as the filter parameter. If the position and / or orientation of the listener does not match the position and / or orientation of a single point in the dimensional map 300, the crosstalk cancellation application 120 identifies the subgroup of points in the set of points in the dimensional map 300 that is closest to the position and / or orientation of the user. In some embodiments, the subgroup of closest points can include four to six points in the set of points in the dimensional map 300 that are closest to the position and / or orientation of the listener. For example, the four points in the dimensional map 300 that are closest to the position of the listener 302 are points A, B, C, and D.
[0034] The crosstalk cancellation application 120 identifies a subgroup of points that is closest to the position and / or orientation of the listener by determining that the position and / or orientation of the listener is within one of the previously calculated and stored tetrahedrons. For example, Figure 3The position of the listener 302 is within the tetrahedron formed by points A, B, C, and D, thereby identifying points A, B, C, and D as a subset of points closest to the position of the listener 302. Since the position of the listener does not directly match the points associated with a single transfer function in the transfer function 132 in the dimensional diagram 300, the crosstalk cancellation application 120 determines the weights 180 of each transfer function 132 associated with the corresponding points in a subset of the closest points. The crosstalk cancellation application 120 determines the weights 180 of each transfer function 132 based on the barycentric distance or Euclidean distance from the corresponding points in a subset of the closest points to the position of the listener 302. For example, the weight 180 of each transfer function 132 associated with each point in the subset of the closest points may be equal to the inverse of the distance from the position of the listener 302, and the weights are normalized such that the sum of the weights of all transfer functions 132 is 1.0. Therefore, the closer the transfer function 132 is to the position of the listener 302, the higher the weight 180 of the corresponding transfer function 132. The higher the weight 180, the greater the impact of the corresponding transfer function 132 on the operating characteristics ( For example , center frequency, gain, Q factor, cutoff frequency, etc.) of one or more filters 138. In cases where the weights 180 need to be used again, the crosstalk cancellation application 120 may store the weights 180 in a memory. For example, whenever the position of the listener 302 moves, the weights need to be recalculated. However, if the position of the listener 302 moves back to a previous position, the stored weights 180 associated with the position of the listener 302 may be used.
[0035] In a three-dimensional scenario with four points (such as points A, B, C, and D), the crosstalk cancellation application may determine the 2x2 matrix of transfer function C mxn , where m is the number of speakers 160 and n is the number of the listener's ears:
[0036]
[0037] The crosstalk cancellation application 120 inputs the weights 180 previously determined as w for points A, B, C, and D, such that:
[0038] w = [w A , w B , w C , w D Equation 2
[0039] Using the weights w, the crosstalk cancellation application 120 can calculate the weighted sum of each transfer function in the 2x2 matrix based on the following equation:
[0040]
[0041] The result of the weighted sum isFigure 2 The transfer function C shown in 11 , C 12 , C 21 and C 22 .
[0042] Figure 4 An example of a filter 138 that performs crosstalk cancellation based on the observed position and orientation of a user in a three-dimensional space is shown in accordance with various embodiments of the present disclosure. As Figure 4 shown, one or more speakers 160 play back an audio source 140a corresponding to the left channel of the audio source 140 and an audio source 140b corresponding to the right channel of the audio source 140. As described above in connection with Figure 2 the audio source 140a represents the signal desired by the left ear of the listener 202 or the left channel of the audio source 140. The audio source 140b represents the signal desired by the right ear of the listener 202 or the right channel of the audio source 140. Without filtering, crosstalk may occur when audio is played back in a three-dimensional environment, such as through the speakers 160 that are far from the ears of the listener 202, as Figure 2 described.
[0043] The crosstalk cancellation application 120 determines the position and orientation of the head of the listener 202 based on sensor data from sensors 150 (such as one or more cameras or other devices that detect the position or orientation of the listener 202). The crosstalk cancellation application 120 further determines the distance between the parameters representing the position and orientation of the head of the listener 202 and one or more points within the dimensional map 134 based on the dimensional map 134, as Figure 3 further explained. In one example, the crosstalk cancellation application 120 calculates the mathematical distance between the position and orientation of the head of the listener 202 and the points within the dimensional map 134, such as the centroid distance or the Euclidean distance. Then, the crosstalk cancellation application 120 identifies the transfer function 132 associated with the closest point based on the calculated centroid distance or Euclidean distance.
[0044] The crosstalk cancellation application 120 selects a transfer function for configuring a set of filters for filtering a portion of the audio source 140. As Figure 4 shown, the audio sources 140a and 140b represent the audio played back by one or more speakers 160 to reduce or remove crosstalk from portions of the audio signals Z1, Z2, Z3, and Z4 that reach the left and right ears of the listener 202. As Figure 4 shown, the filters H 1,1 and H 1,2 filter a portion of the audio source 140a and the filters H 2,1 and H 2,2 filter a portion of the audio source 140b such that when in accordance with C 1,1, C 1,2 , C 2,1 and C 2,2 When outputting the audio source 140 in an environment that affects the played-back signal, crosstalk is reduced or eliminated.
[0045] V1 and V2 respectively represent the filtered and output corresponding filtered portions of the audio source 140a to one or more speakers 160 through filters H 1,1 and H 1,2 . V3 and V4 respectively represent the filtered and output corresponding filtered portions of the audio source 140b to one or more speakers 160 through filters H 2,1 and H 2,2 . Therefore, when the environment changes the signals output by the filters and played back by one or more speakers 160 according to C 1,1 , C 1,2 , C 2,1 and C 2,2 , the crosstalk of the signals reaching the ears of the listener 202 is reduced or eliminated. As shown in Figure 4 , H 1,1 and H 1,2 filter the audio source 140a to produce V1 and V2 played back by one or more speakers 160, such that when subjected to the environment through C 1,1 and C 2,1 , the resulting signals Z1 and Z3 reaching the left ear of the listener 202 only correspond to the audio source 140a, i.e., the left channel of the audio source 140. Similarly, H 2,1 and H 2,2 filter the audio source 140b to produce V3 and V4 played back by one or more speakers 160, such that when subjected to the environment through C 1,2 and C 2,2 , the resulting signals Z2 and Z4 reaching the right ear of the listener 202 only correspond to the audio source 140b, i.e., the right channel.
[0046] To determine the correct filters H 1,1 , H 1,2 , H 2,1 and H 2,2 , the crosstalk cancellation application 120 can solve for H based on the weighted sum of C Figure 3 described in mxn using Equation 4 below: mxn :
[0047] CH = B Equation 4
[0048] where C is the transfer function determined in Figure 3 , and
[0049] B = I Equation 5
[0050] where I is the identity matrix. To avoid direct inversion of ill-conditioned systems, the crosstalk cancellation application 120 can use any technically feasible technique to obtain the filter H mxn with the desired behavior, such as pseudo-inverse, regularized inverse, frequency-dependent regularization, least mean square (LMS) filter design with arbitrary penalty functions, or similar techniques. The result is that the filters H 1,1 、H 1,2 、H 2,1 and H 2,2 can be used by the audio sources 140a and 140b to appropriately filter the audio to reduce or remove crosstalk reaching the respective ears of the listener 202.
[0051] Figure 5 FIG. shows a flow chart of method steps for configuring the transfer function of a filter for performing crosstalk cancellation according to one or more embodiments in combination. Although the method steps are described with reference to Figures 1 to 4 embodiments, those skilled in the art will understand that any system configured to implement the method steps in any order is within the scope of the present disclosure.
[0052] Method 500 begins at step 502, where the crosstalk cancellation application 120 determines the position and orientation of the listener 202 within the environment. The environment includes a space in which audio is played back through one or more speakers 160, such as the interior of a vehicle or any other interior or exterior environment. The crosstalk cancellation application 120 determines the position and orientation of the listener 202 based on sensor data obtained from sensors 150 associated with the audio processing system 100. As noted above, the sensors 150 include optical sensors, pressure sensors, proximity sensors, and other sensors that obtain information about the environment and the position and orientation of the listener 202 within the environment. The position of the listener 202 is determined based on the sensor data relative to a reference position within the environment. The orientation of the listener 202 is also determined relative to a reference orientation within the environment. In some embodiments, the crosstalk cancellation application 120 determines the position and orientation of the head and / or ears of the listener 202 based on the sensor data.
[0053] At step 504, the crosstalk cancellation application 120 identifies a subgroup of points in the dimensionality map 300 based on the position and / or orientation of the listener 202 within the environment. In one example, a given position and orientation of the listener 202 are represented by coordinates in a six-dimensional space. Since the position and / or orientation of the listener 202 do not directly match a specific point within the dimensionality map 300, the crosstalk cancellation application 120 identifies a subgroup of points within the dimensionality map 300. For example, the dimensionality map 300 can include six dimensions ( For example, a very large number ( For example , hundreds, thousands, etc.) of points, each of which is associated with the transfer function 132. In some embodiments, a simplified method of identifying points based on the position and orientation of the listener 202 includes reducing the dimensions of the position and orientation of the listener considered when identifying the points associated with the listener 202 in the dimensionality map 300. To reduce mathematical complexity, a reduced set of parameters representing the position and orientation of the listener can be considered. For example, one or more of the parameters representing orientation can be removed, and a set of nearest points can be identified based on the mathematical distance in the dimensionality map 300 from the coordinates representing the position and orientation of the listener in the table to one or more points from the set of points. Examples of coordinates that can be removed include yaw angle, pitch angle, and / or roll angle. As another example, an alternative simplified method of identifying the transfer function 132 includes reducing the dimensionality of the dimensionality map 300. As noted above, the dimensionality map 300 includes a subset of nearest points representing the position of the user in three-dimensional space, independent of orientation. In any of the above scenarios, the crosstalk cancellation application 120 identifies a subset of nearest points in the dimensionality map 300 that are closest to the points representing at least some of the parameters corresponding to the position and orientation of the listener 202.
[0054] In some embodiments, the subset of nearest points can include four to six points or the set of points in the dimensionality map 300 that are closest to the position and / or orientation of the listener. For example, the four points in the dimensionality map 300 that are closest to the position of the listener 302 are points A, B, C, and D. The crosstalk cancellation application 120 identifies a subset of points that are closest to the position and / or orientation of the listener by determining that the position and / or orientation of the listener is within one of the previously calculated and stored polygons. For example, Figure 3 the position of the listener 302 in is located within the tetrahedron formed by points A, B, C, and D, whereby points A, B, C, and D are identified as a subset of points that are closest to the position of the listener 302. Since the position of the listener and the points associated with a single transfer function in the transfer function 132 in the dimensionality map 300 do not directly match, the crosstalk cancellation application 120 determines the weight 180 of each transfer function 132 associated with each corresponding point in the subset of nearest points.
[0055] At step 506, the crosstalk cancellation application 120 combines the transfer functions 132 associated with a subgroup of points. The crosstalk cancellation application 120 determines a weight 180 for each transfer function 132 based on the barycentric distance or Euclidean distance from the corresponding point in a subgroup of nearest points to the position of the listener 302. For example, the weight 180 for each transfer function 132 associated with each point in the subgroup of nearest points may be equal to the inverse of the distance to the position of the listener 302, and the weights are normalized such that the sum of the weights of all transfer functions 132 is 1.0. Thus, the closer the transfer function 132 is to the position of the listener 302, the higher the weight 180 of the corresponding transfer function 132. The higher the weight 180, the greater the influence of the corresponding transfer function 132 on the operating characteristics ( For example , center frequency, gain, Q factor, cut-off frequency, etc.) of one or more filters 138. For example, in a three-dimensional scenario with four points A, B, C, and D such as Figure 3 , the crosstalk cancellation application may determine a 2x2 matrix of transfer function C mxn , where m is the number of speakers 160 and n is the number of the listener's ears: Using equations 1, 2, and 3, the crosstalk cancellation application 120 calculates the weighted sum of transfer function C Figure 2 shown in 11 , C 12 , C 21 , and C 22 .
[0056] At step 508, the crosstalk cancellation application 120 configures one or more filters 138 based on the combined transfer functions C 11 , C 12 , C 21 , and C 22 determined in step 506. To determine the correct filters H 1,1 , H 1,2 , H 2,1 , and H 2,2 , the crosstalk cancellation application 120 may solve for H mxn based on the weighted sum of C mxn using equations 4 and 5. The result is that filters H 1,1 , H 1,2 , H 2,1 , and H 2,2 can be used by the audio sources 140a and 140b to appropriately filter the audio to reduce or remove crosstalk reaching the corresponding ears of the listener 202.
[0057] At step 510, the crosstalk cancellation application 120 generates an audio signal for playback based on a filter 138 configured with the identified combined transfer function 132. The audio signal is generated based on an audio source 140 (such as a song or other audio input provided to the audio processing system 100) played back by the audio processing system 100 within the environment. The audio source 140 includes a left channel and a right channel. The crosstalk cancellation application 120 filters the audio source 140 using the filter 138 configured with the combined transfer function 132, which is selected based on a subgroup of points in the dimensionality map 300 that are closest to the position and orientation of the listener 202. When played back in the environment, the filtered audio signals reach the left and right ears of the listener 202 respectively, such that crosstalk is reduced or eliminated.
[0058] For example, as Figure 4 shown, the audio sources 140a and 140b represent audio played back by one or more speakers 160 to reduce or eliminate crosstalk from portions of the audio signals Z1, Z2, Z3, and Z4 reaching the left and right ears of the listener 202. The newly determined filters H 1,1 and H 1,2 filter a portion of the audio source 140a and the newly determined filters H 2,1 and H 2,2 filter a portion of the audio source 140b such that when the audio source 140 is output in an environment that affects the played-back signal according to C 1,1 、C 1,2 、C 2,1 and C 2,2 crosstalk is reduced or eliminated.
[0059] Thus, when the environment changes the signal output by the filters H 1,1 、H 1,2 、H 2,1 and H 2,2 and played back by one or more speakers 160 according to C 1,1 、C 1,2 、C 2,1 and C 2,2 the crosstalk of the signal reaching the ears of the listener 202 is reduced or eliminated. For example, as Figure 4 shown, H 1,1 and H 1,2 filter the audio source 140a to produce V1 and V2 played back by one or more speakers 160 such that when subjected to the effects of the environment through C 1,1 and C 2,1 the resulting signals Z1 and Z3 reaching the left ear of the listener 202 correspond only to the audio source 140a, i.e., the left channel of the audio source 140. Similarly, H 2,1 and H 2,2Filter the audio source 140b to produce V3 and V4 for playback by one or more speakers 160 such that when subjected to the environment via C 1,2 and C 2,2 the resulting signals Z2 and Z4 arriving at the right ear of the listener 202 correspond only to the audio source 140b, i.e., the right channel.
[0060] At step 512, the crosstalk cancellation application 120 outputs the filtered audio signal to one or more speakers 160 associated with the audio processing system 100. The one or more speakers 160 play back the filtered audio signal in the environment based on the filtered audio signal. The one or more speakers 160 include one or more speakers corresponding to the left channel of the audio processing system 100 and one or more speakers corresponding to the right channel of the audio processing system 100.
[0061] In summary, the crosstalk cancellation application configures a set of filters for performing crosstalk cancellation between the left and right channels of an audio source played back by one or more speakers. The crosstalk cancellation application configures the set of filters by selecting a transfer function for each filter in the set of filters. The transfer function is selected by identifying the position and orientation of the user's head in three-dimensional space using sensor data from one or more sensors. A dimensional map specifies a set of points respectively associated with the transfer functions for configuring the filters. Identify a subgroup of the set of points in the dimensional map that is closest to the position and orientation of the user's head. Interpolate the weights of each transfer function associated with the subgroup of points based on the centroid distance from the dimensional map to the position and orientation of the user's head. Filter one or more signals corresponding to the audio source using the filters with the weighted transfer functions, the one or more signals being used to drive one or more speakers to form a sound field. The one or more speakers play back the respective filtered signals. When altered by the environment, crosstalk is reduced or removed once the filtered signals reach the listener's ears.
[0062] At least one technical advantage of the disclosed technology over the prior art is that through the disclosed technology, the audio processing system forms an improved crosstalk cancellation filter in real time for a user with six degrees of freedom by interpolating a subgroup of filters. By interpolating a subgroup of filters, spectral distortion caused by user movement is reduced without the typical computational cost associated with movement covering all six degrees of freedom. Additionally, the audio intended to be received by the user's left and right ears respectively more accurately represents the audio input output by the audio processing and playback system. These technical advantages provide one or more technological advancements over prior art methods.
[0063] 1. In some embodiments, a computer-implemented method includes: determining a user's position and orientation in an environment; determining a subgroup of nearest points in a dimensionality map based on the user's position and orientation, where the dimensionality map includes a set of points in a multi-dimensional space, each point associated with a corresponding transfer function, and in the dimensionality map, the points in the subgroup of nearest points are closer to the user's position and orientation in the dimensionality map than other points; determining a respective weight for each transfer function associated with the subgroup of nearest points based on the distance between each point in the subgroup of nearest points and the user's position and orientation in the dimensionality map; determining at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; generating a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; and transmitting the plurality of audio signals to the plurality of loudspeakers for output.
[0064] 2. The computer-implemented method of clause 1, wherein determining the user's position and orientation in the environment includes receiving sensor data from a plurality of sensors.
[0065] 3. The computer-implemented method of clause 1 or 2, wherein determining the user's position and orientation in the environment includes calculating three coordinates corresponding to a position relative to a reference position and three coordinates corresponding to an orientation relative to a reference orientation.
[0066] 4. The computer-implemented method of clause 3, wherein the three coordinates corresponding to the orientation relative to the reference orientation correspond to a roll angle, a pitch angle, and a yaw angle.
[0067] 5. The computer-implemented method of any one of clauses 1 to 4, wherein the subgroup of nearest points in the dimensionality map is closer to the user's position and orientation in the dimensionality map than the other points based on: generating a non-overlapping polygonal space for each subgroup of points in the dimensionality map, where each vertex of each non-overlapping polygonal space is a different point included in an associated subgroup of points, and where the circumscribed hypersphere of each non-overlapping polygonal space contains only points within the associated subgroup of points; and determining that the user's position and orientation are located within the non-overlapping polygonal space associated with a subgroup of nearest points.
[0068] 6. The computer-implemented method of any one of clauses 1 to 5, wherein the non-overlapping polygonal space for each subgroup of points in the dimensionality map is generated based on Delaunay triangulation.
[0069] 7. The computer-implemented method according to any one of clauses 1 to 6, wherein determining the respective weights of each transfer function further comprises: determining the weights of each transfer function based on the mathematical distances between the position and orientation of the user within the dimensional map and each point in the subgroup of nearest points.
[0070] 8. The computer-implemented method according to clause 7, wherein the mathematical distance is calculated based on the centroid distance or the Euclidean distance.
[0071] 9. The computer-implemented method according to any one of clauses 1 to 8, wherein the respective weights of each transfer function are inversely proportional to the normalized distance to the position and orientation of the user, and the sum of the respective weights is equal to 1.
[0072] 10. The computer-implemented method according to any one of clauses 1 to 9, wherein the dimensional map is selected from a plurality of dimensional maps, and the dimensional map is selected based on the yaw angle relative to a reference orientation corresponding to the first orientation.
[0073] 11. The computer-implemented method according to clause 10, wherein each of the plurality of dimensional maps is associated with a series of yaw angles relative to the reference orientation.
[0074] 12. The computer-implemented method according to any one of clauses 1 to 11, wherein the plurality of audio signals include a left channel signal and a right channel signal.
[0075] 13. In some embodiments, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: determining the position and orientation of a user in an environment; determining a subgroup of nearest points in a dimensional map based on the position and orientation of the user, wherein the dimensional map includes a set of points in a multi-dimensional space, each point being associated with a corresponding transfer function, and in the dimensional map, the points in the subgroup of nearest points are closer to the position and orientation of the user in the dimensional map than other points; determining the respective weights of each transfer function associated with the subgroup of nearest points based on the distances between each point in the subgroup of nearest points and the position and orientation of the user in the dimensional map; determining at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; generating a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; and transmitting the plurality of audio signals to the plurality of loudspeakers for output.
[0076] 14. One or more non-transitory computer-readable media as described in clause 13, wherein each of the subset of nearest points in the dimensionality map is closer to the user's position and orientation in the dimensionality map than the other points based on: generating non-overlapping polygonal spaces for each subset of points in the dimensionality map, wherein each vertex of each non-overlapping polygonal space is a different point included in the associated subset of points, and wherein the circumscribed hypersphere of each non-overlapping polygonal space contains only points within the associated subset of points; and determining that the user's position and orientation are within the non-overlapping polygonal space associated with a subset of nearest points.
[0077] 15. One or more non-transitory computer-readable media as described in clause 13 or 14, wherein the step of determining the respective weights of each transfer function further comprises: determining the weights of each transfer function based on the mathematical distance between the user's position and orientation within the dimensionality map and each point in the subset of nearest points.
[0078] 16. One or more non-transitory computer-readable media as described in any one of clauses 13 to 15, wherein the step of determining the user's position and orientation in the environment further comprises: calculating three coordinates corresponding to the position relative to a reference position and three coordinates corresponding to the orientation relative to a reference orientation.
[0079] 17. One or more non-transitory computer-readable media as described in any one of clauses 13 to 16, wherein the respective weights of each transfer function are inversely proportional to the normalized distance to the user's position and orientation, and the sum of the respective weights is equal to 1.
[0080] 18. One or more non-transitory computer-readable media as described in any one of clauses 13 to 17, wherein the dimensionality map includes three dimensions representing positions in three-dimensional space and includes the orientation as a series of yaw angles.
[0081] 19. One or more non-transitory computer-readable media as described in any one of clauses 13 to 18, wherein the dimensionality map includes three dimensions representing positions in three-dimensional space and includes two dimensions representing the orientation in two-dimensional space.
[0082] 20. In some embodiments, a system includes: at least one sensor configured to obtain information about a user in an environment; at least one speaker configured to play back audio within the environment; a memory storing a crosstalk cancellation application; and a processor coupled to the memory, the processor executing the crosstalk cancellation application by performing the following steps: determining a location and orientation of the user in the environment; determining a subgroup of nearest points in a dimensionality map based on the location and orientation of the user, wherein the dimensionality map includes a set of points in a multi-dimensional space, each point associated with a corresponding transfer function, and in the dimensionality map, the points in the subgroup of nearest points are closer to the location and orientation of the user in the dimensionality map than other points; determining a respective weight for each transfer function associated with the subgroup of nearest points based on a distance between each point in the subgroup of nearest points and the location and orientation of the user in the dimensionality map; determining at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; generating a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; and transmitting the plurality of audio signals to the plurality of loudspeakers for output.
[0083] Any and all combinations of any claim elements recited in any of the claims and / or any elements described in this application are within the scope and coverage of the present invention and protection in any way.
[0084] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0085] Aspects of the embodiments of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an all-hardware embodiment, an all-software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which embodiments are generally referred to herein as a "module", "system", or "computer". Additionally, any hardware and / or software technologies, processes, functions, components, engines, modules, or systems described in the present disclosure may be implemented as a circuit or a collection of circuits. Further, aspects of the present disclosure may take the form of a computer program product comprising one or more computer-readable media having computer-readable program code thereon.
[0086] Any combination of one or more computer-readable media can be utilized. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, by way of example and not limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following media: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing media. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0087] As described above, aspects of the present disclosure are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed via the processor of the computer or other programmable data processing apparatus, enable the functions / acts specified in one or more blocks of the flowchart and / or block diagram to be implemented. Such a processor can be, by way of example and not limitation, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.
[0088] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0089] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, and such scope is determined by the claims that follow.
Claims
1. A computer-implemented method, comprising: Determining a location and orientation of a user in an environment; Determining a subgroup of nearest points in a dimensionality map based on the location and the orientation of the user, wherein the dimensionality map includes a set of points in a multi-dimensional space, each point being associated with a corresponding transfer function, and in the dimensionality map, the points in the subgroup of nearest points are closer to the location and the orientation of the user in the dimensionality map than other points; Determining a respective weight of each transfer function associated with the subgroup of nearest points based on a distance between each point in the subgroup of nearest points and the location and the orientation of the user in the dimensionality map; Determining at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; Generating a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; And Transmitting the plurality of audio signals to the plurality of loudspeakers for output.
2. The computer-implemented method according to claim 1, wherein determining the location and the orientation of the user in the environment includes receiving sensor data from a plurality of sensors.
3. The computer-implemented method according to claim 1, wherein determining the location and the orientation of the user in the environment includes calculating three coordinates corresponding to a location relative to a reference location and three coordinates corresponding to an orientation relative to a reference orientation.
4. The computer-implemented method according to claim 3, wherein the three coordinates corresponding to the orientation relative to the reference orientation correspond to a roll angle, a pitch angle, and a yaw angle.
5. The computer-implemented method according to claim 1, wherein the subgroup of nearest points in the dimensionality map is closer to the location and the orientation of the user in the dimensionality map than the other points based on: Generating a non-overlapping polygon space for each subgroup of points in the dimensionality map, wherein each vertex of each non-overlapping polygon space is a different point included in an associated subgroup of points, and wherein the circumscribed hypersphere of each non-overlapping polygon space contains only points within the associated subgroup of points; and Determining that the location and the orientation of the user are within a non-overlapping polygon space associated with a subgroup of nearest points.
6. The computer-implemented method according to claim 1, wherein the non-overlapping polygon space for each subgroup of points in the dimensionality map is generated based on Delaunay triangulation.
7. The computer-implemented method according to claim 1, wherein determining the respective weight of each transfer function further includes: Determining a weight of each transfer function based on a mathematical distance between the location and the orientation of the user within the dimensionality map and each point in the subgroup of nearest points.
8. The computer-implemented method according to claim 7, wherein the mathematical distance is calculated based on a centroid distance or an Euclidean distance.
9. The computer-implemented method according to claim 1, wherein the respective weights of each transfer function are inversely proportional to the normalized distance to the position and orientation of the user, and the sum of the respective weights is equal to 1.
10. The computer-implemented method according to claim 1, wherein the dimensionality map is selected from a plurality of dimensionality maps, and the dimensionality map is selected based on the yaw angle relative to a reference orientation corresponding to the first orientation.
11. The computer-implemented method according to claim 10, wherein each of the plurality of dimensionality maps is associated with a series of yaw angles relative to the reference orientation.
12. The computer-implemented method according to claim 1, wherein the plurality of audio signals include a left channel signal and a right channel signal.
13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: Determine the position and orientation of a user in an environment; Determine a subgroup of nearest points in a dimensionality map based on the position and orientation of the user, wherein the dimensionality map includes a set of points in a multi-dimensional space, each point being associated with a corresponding transfer function, and in the dimensionality map, the points in the subgroup of nearest points are closer to the position and orientation of the user in the dimensionality map than other points; Determine the respective weights of each transfer function associated with the subgroup of nearest points based on the distance between each point in the subgroup of nearest points and the position and orientation of the user in the dimensionality map; Determine at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; Generate a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; And Transmit the plurality of audio signals to the plurality of loudspeakers for output.
14. The one or more non-transitory computer-readable media according to claim 13, wherein each of the subgroup of nearest points in the dimensionality map is closer to the position and orientation of the user in the dimensionality map than the other points based on: Generating a non-overlapping polygonal space for each subgroup of points in the dimensionality map, wherein each vertex of each non-overlapping polygonal space is a different point included in an associated subgroup of points, and wherein the circumscribed hypersphere of each non-overlapping polygonal space contains only points within the associated subgroup of points; and Determining that the position and orientation of the user are within the non-overlapping polygonal space associated with a subgroup of nearest points.
15. The one or more non-transitory computer-readable media according to claim 13, wherein the step of determining the respective weights of each transfer function further includes: Determining the weight of each transfer function based on the mathematical distance between the position and orientation of the user within the dimensionality map and each point in the subgroup of nearest points.
16. The one or more non-transitory computer-readable media of claim 13, wherein the step of determining the user's position and orientation in the environment further comprises: Calculating three coordinates corresponding to the position relative to a reference position and three coordinates corresponding to the orientation relative to a reference orientation.
17. The one or more non-transitory computer-readable media of claim 13, wherein the respective weights of each transfer function are inversely proportional to the normalized distance to the user's position and orientation, and the sum of the respective weights equals 1.
18. The one or more non-transitory computer-readable media of claim 13, wherein the dimensionality map includes three dimensions representing a position in three-dimensional space and includes an orientation as a series of yaw angles.
19. The one or more non-transitory computer-readable media of claim 13, wherein the dimensionality map includes three dimensions representing a position in three-dimensional space and includes two dimensions representing an orientation in two-dimensional space.
20. A system, comprising: At least one sensor configured to obtain information about a user in an environment; At least one speaker configured to play back audio within the environment; A memory storing a crosstalk cancellation application; And A processor coupled to the memory, the processor executing the crosstalk cancellation application by performing the following steps: Determining the position and orientation of a user in an environment; Determining a subgroup of nearest points in a dimensionality map based on the user's position and orientation, wherein the dimensionality map includes a set of points in a multi-dimensional space, each point associated with a corresponding transfer function, and in the dimensionality map, the points in the subgroup of nearest points are closer to the user's position and orientation in the dimensionality map than other points; Determining the respective weights of each transfer function associated with the subgroup of nearest points based on the distance between each point in the subgroup of nearest points and the user's position and orientation in the dimensionality map; Determining at least one crosstalk cancellation filter by combining each of the transfer functions associated with the subgroup of nearest points based on the respective weights; Generating a plurality of audio signals for a plurality of loudspeakers based on the at least one crosstalk cancellation filter; And Transmitting the plurality of audio signals to the plurality of loudspeakers for output.