Wearable audio device to achieve immersive spatial audio experience and method thereto
The wearable audio device with a speaker array and crosstalk cancellation filtering generates binaural signals to overcome the lack of spatial audio in existing devices, achieving an immersive 3D audio experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HARMAN INT IND INC
- Filing Date
- 2025-01-26
- Publication Date
- 2026-07-30
AI Technical Summary
Wearable audio devices typically support only mono or stereo sound and lack immersive spatial audio experience due to the absence of complex signal processing algorithms.
A wearable audio device with a speaker array and processors that generate binaural signals through crosstalk cancellation filtering, using ideal HRTFs and measured transfer functions to reproduce multi-channel audio signals, enhancing spatial audio experience.
The device provides an immersive spatial audio experience by accurately simulating sound sources in 3D space, preserving directional sense and spatial content, even with a single wearable device.
Smart Images

Figure CN2025075239_30072026_PF_FP_ABST
Abstract
Description
WEARABLE AUDIO DEVICE TO ACHIEVE IMMERSIVE SPATIAL AUDIO EXPERIENCE AND METHOD THERETOTECHNICAL FIELD
[0001] The present disclosure relates generally to wearable audio devices. More particularly, the present disclosure relates to a wearable audio device to achieve immersive spatial audio experience and a method thereto.BACKGROUND
[0002] Wearable audio devices have been becoming more and more popular over past few years. Many developers and manufacturers intend to embed audio speakers into wearable devices, such as on-ear headphones, eye-glass frames, neckband speakers, or the like, to provide users with audio listening functionality.
[0003] However, such wearable audio devices typically may only support mono or stereo sound, but lack of spatial experience, or reproduce audio through simple signal routing without any complicated signal processing algorithms.
[0004] Therefore, it is necessary to provide a solution for the users to achieve immersive spatial audio experience with a single wearable audio device.
[0005] SUMMARY OF THE PRESENT DISCLOSURE
[0006] In order to overcome the shortcomings and deficiencies in the prior art, the purpose of the present disclosure is to provide a wearable audio device to achieve immersive spatial audio experience and method thereto.
[0007] In one aspect, a wearable audio device to achieve immersive spatial audio experience is provided herein. The wearable audio device may comprise a speaker array with multiple speakers that can be arranged at certain positions on the user’s body. The wearable audio device further comprises one or more processors configured to receive a multi-channel signal from a multi-channel source. Based on at least one channel of the received multi-channel signal, a binaural signal composed of a left-binaural signal and a right-binaural signal may be generated. After processing the binaural signal into speaker signals through crosstalk cancellation filtering, the speaker signals can be replayed with the multiple speakers, to reproduce the multi-channel signal at the user’s ears.
[0008] In another aspect, a method to achieve immersive spatial audio experience with a wearable audio device is provided herein. The method may comprise steps of arranging multiple speakers into a speaker array at certain positions on a user’s body. One or more processors in the wearable audio device may be configured to receive a multi-channel signal from a multi-channel source. Based on the received multi-channel signal, a binaural signal composed of a left-binaural signal and a right-binaural signal may be generated. Then, after processing the binaural signal into speaker signals through crosstalk cancellation filtering, the speaker signals can be replayed to the user with the multiple speakers, to reproduce the multi-channel signal at the user’s ears.
[0009] In yet another aspect, a non-transitory computer-readable medium including instructions which, when executed by a processor, perform the method to achieve immersive spatial audio experience with a wearable audio device is provided herein, as well.
[0010] In one or more embodiments, a pair of ideal HRTFs may be applied to each channel of the multi-channel signal to generate the left-binaural signal and the right-binaural signal, respectively, and wherein the pair of ideal HRTFs can be each retrieved from open-source databases based on multiple spatial contents in the multi-channel signal.
[0011] In one or more embodiments, the crosstalk cancellation filtering may further comprise applying the inverses of measured transfer function matrix to each of the multiple speakers, respectively, to process the binaural signal into the speaker signals. In particular, a pair of inverses of the measured transfer function may be applied to each of the multiple speaker. The measured transfer functions shall be physically measured utilizing a mannequin wearing the speaker array or simulated with parametric spherical head model in designing the arrangement of the multiple speakers.
[0012] In one or more embodiments, a center channel signal of the multi-channel signal may be directly distributed to at least one of the multiple speakers in front of the user’ ears without any binauralization processing and the crosstalk cancellation filtering.
[0013] In one or more embodiments, a bass channel signal of the multi-channel signal may be replayed with at least one of the multiple speakers, without any binauralization processing and the crosstalk cancellation filtering.
[0014] In one or more embodiments, the ear binaural signals above a certain frequency can be replayed with one speaker on each side of the user that has the highest raw channel separation, respectively, without the crosstalk cancellation filtering.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The present disclosure may be better understood from reading the following description of non-limiting embodiments, with reference to the attached drawings. In the figures, like reference numeral designates corresponding parts, wherein below:
[0016] FIG. 1A and 1B illustrate exemplary schematic diagrams of the wearable audio device, with a neckband speaker array being worn by a user, to achieve immersive spatial audio experience, according to one or more embodiments of the present disclosure;
[0017] FIG. 2 illustrates an exemplary flowchart of signal processing to achieve immersive spatial audio experience in the wearable audio device, according to one or more embodiments of the present disclosure;
[0018] FIG. 3 illustrates an exemplary flowchart of the method to achieve immersive spatial audio experience with the wearable audio device, according to one or more embodiments of the present disclosure; and
[0019] FIG. 4 illustrates an exemplary schematic diagram of generating the binaural signals from a 5.1 surround sound, according to one or more embodiments of the present disclosure.DETAILED DESCRIPTION
[0020] The detailed description of the one or more embodiments of the present disclosure is disclosed hereinafter; however, it is understood that the disclosed embodiments are merely exemplary of the present disclosure that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and function details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present disclosure.
[0021] When spatial sounds are recorded, amplified, and replayed, the spatial distribution of the sounds may disappear. However, if the original spatial sense can be largely restored from recording to playback, while reproducing spatial distribution characteristics such as directional hierarchy through a single compact and wearable audio device, it may bring users an immersive spatial experience and very convenient auditory enjoyment.
[0022] In this disclosure, a wearable audio device equipped with multiple speakers is provided. The multiple speakers therein may be arranged as a speaker array, to replay surround sounds or even Dolby Atmos, enabling its user to achieve an immersive spatial auditory experience in the listening space. The wearable audio device may comprise, but not limited to, such as a headset, an eye-glass frame, a neckband speaker, or the like.
[0023] A neckband speaker array can be taken as an example, but not only limited thereto, to illustrate the technical solution disclosed herein to achieve an immersive spatial audio experience with multiple speakers arranged as the speaker array in a wearable audio device.
[0024] FIG. 1A and 1B illustrate exemplary schematic diagrams 100 and 100’ of the wearable audio device, with a neckband speaker array being worn by a user, to achieve immersive spatial audio experience, according to one or more embodiments of the present disclosure. The top view of the user wearing the neckband speaker array is shown in FIG. 1A, and the front view is shown in FIG. 1B. The neckband speaker array with six speakers, comprising speakers 120, 122 and 124 on the left side and speakers 130, 132 and 134 on the right side of the user, are arranged within the wearable audio device being worn on the user’s neck. Accordingly, it can be seen in both FIGs. 1A and 1B that, for the user’s head 110 with his / her front face (not shown) 110’ , and his / her left ear 112 and right ear 114 thereon, the six speakers can be located at different positions on the user’s body, relative to the user’s ears 112 and 114, respectively. The speakers 120, 130 are located on the left front and right front of the user’s ears, the speakers 122, 132 are located on the left side and right side of the user’s ears, and the speakers 124, 134 are located on the left rear and right rear of the user’s ears. The orientation and distance from each speaker therein to the user’s ears are also different.
[0025] In the example shown in FIGs. 1A and 1B, there are even number of speakers distributed symmetrically on the left and right sides of the human head. Nevertheless, in one or more embodiments, the number of the plurality of speakers may be not restricted to a certain number, and the multiple speakers can be distributed as any speaker array, not limited to a symmetric arrangement of even numbered speakers. The form of the wearable audio device with such a speaker array can also be diverse, such as but not limited to, a neckband, an eye-glass frame, a headset, etc.
[0026] In general, the arrangement of multiple speakers in the speaker array should be designed such that when the user is wearing the wearable audio device, the multiple speakers are located in certain positions of the user's body, such as in periphery of or around the user’s head and ears. For example, the multiple speakers may be located in any directions relative to the user’s head or ears, and at different distances from the user’s head, his / her left ear, or right ear, in three-dimensional (3D) space.
[0027] Alternatively or additionally, the number of the multiple speakers in the speaker array can also vary, for example, more speakers may be arranged to reproduce more multi-channel source, and one or more other embodiments with 6, 9, or any actual number of speakers are also possible. The number of the multiple speakers arranged on either side of the user’s head or ears may be more or less than on the other side. In an example, at least one speaker shall be arranged in front of and behind of the user’s ears, respectively.
[0028] Sound source shall be recorded as a multi-channel signal enabling the multiple speakers in the speaker array to reproduce the sound effect of the multi-channel source at the time of recording in the field, allowing the user to obtain an immersive spatial acoustic experience, when the user is listening thereto with the wearable audio device.
[0029] FIG. 2 illustrates an exemplary signal flowchart 200 of signal processing to achieve immersive spatial audio experience in the wearable audio device, according to one or more embodiments of the present disclosure. To virtual render of spatial audio for both ears of a user shall involve the creation of a binaural signal 230. The binaural signal may represent the desired sound arriving at the listener’s left and right ears, respectively, and can be synthesized to simulate a particular audio scene in 3D space, where contains possibly a multitude of sound sources at different locations, referred to as spatial contents herein. The binaural signal 230 may comprise a left-binaural signal 232 and a right-binaural signal 234, which then can be both fed into the multiple speakers 250 of the speaker array through crosstalk cancellation (CTC) filtering 240, to generate speaker signals which shall be transmitted via the multiple speakers to the user’s ears.
[0030] As shown in FIG. 2, one or more processors (not shown) mounted in the wearable audio device may be configured to receive a multi-channel signal 220 from the multi-channel source 210, which may comprise audio signals from multiple channels 1, 2, …, n, and one bass channel. Each channel of the multiple channels may contain and carry its spatial content. As an example, the multi-channel source 210 may comprise, but not limited to, local stored audio files, online music streaming apps, such as movie audio, music, Tidal or the like, and large-scale computer games, such as PS3, XBOX360 or the like, for example.
[0031] In one example, the multi-channel signal 220 can be a surround sound, such as a 5.1 or 7.1 surround sound or other multi-channel surround sound. The surround sound may not only retain the directional sense of the original audio source, but also produce a sound effect that creates a sense of surround and expansion. This effect allows listeners to feel sound coming from all directions, as if surrounded by sound, thereby enhancing the sense of presence and space in the music. Generally, for the 5.1 surround sound, it may comprise five surround channels 1, 2, …, 5, and one bass channel for the low-frequency audio; and for the 7.1 surround sound, it may comprise seven surround channels 1, 2, …, 7, and one bass channel, correspondingly.
[0032] Alternatively or additionally, the audio source may be provided as an object-based signal, such as Dolby Atmos. The Dolby Atmos breaks through the limitations of surround sound and provides an immersive audio experience. In one or more applications, those objects may be sound sources allowed to move freely anywhere in 3D space, and thus may be given by the individual channels of the multi-channel signal. In this case, the object-based signal may be firstly encoded into a channel-based signal to be stored in the multi-channel source 210. For example, for 7.1.4 Dolby Atmos, it can be decoded into the multi-channel signal, comprising seven surround channels, one bass channel, and four height channels.
[0033] Then, for each channel in the multi-channel signal 220, such as Channel 1, Channel 2, …, Channel n, its corresponding spatial content can be rendered into a binaural signal for the left and right ears of the user, respectively. Each ear signal of the binaural signal 230, including the left-binaural signal 232 for the left ear (or the right-binaural signal 234 for the right ear) , may be modeled as the sum of the left- (or right-) channel signals, respectively.
[0034] The spatial content in each channel of the multi-channel signal 220 may be processed with a pair of head-related transfer functions (HRTFs) for both ears, to generate the binaural signal 230. Generally, such HRTFs are meant to represent a set of model HRTFs addressable by position, which may contain both the head-related impulse response and the reverberation information. In this case, many ideal HRTFs already determined from human subjects in a laboratory exist and can be found from open-source databases of high-spatial-resolution HRTF measurements for a number of different subjects, such as the CIPIC database.
[0035] The spatial content included in each of the multiple channels represents the desired position of the specific channel in space relative to the listener, which may be represented in a coordinate system, such as a Cartesian coordinate or a polar system, for example. For each channel in the multi-channel signal, a pair of HRTFs (e.g., HRTF n, left and HRTF n, right, shown in FIG. 2) for the left-binaural and right-binaural signals selected as a function of the spatial content of the channel (e.g., Channel n) can be applied to generate the binaural signal 230, respectively. In one or more examples, the spatial content therein may be varying in time to simulate dynamical movement of the audio source.
[0036] Next, the binaural signal 230 shall be replayed via the multiple speakers 250 after processing it into speaker signals through the CTC filtering 240. In particular, the CTC filtering 240 may eliminate or reduce the natural crosstalk inherent in the speaker playback, so that the left-binaural signal 232 of the binaural signal 230 can be delivered substantially to the left ear only and the right-binaural signal 234 to the right ear of the listener only, for preserving the intention of the binaural signal 230.
[0037] The CTC filtering 240 may be designed based on a model of acoustic transmission from each of the multiple speakers 250 to the listener’s left and right ears, respectively. The binaural signal to be transmitted via each speaker of the multiple speakers 250 to the user’s ears, (e.g., each speaker signal) , shall be filtered with a CTC filter which may comprise a separate linear time-invariant transfer function HS modeling the acoustic transmission from each of the multiple speakers 250 to that ear.
[0038] A set of transfer functions HS shall be modeled using corresponding head-related transfer functions (HRTFs) physically measured as a function of the placement of the multiple speakers 250 in the speaker array with respect to the listener, i.e., the user of the specific wearable audio device here. In this case, such a transfer function HS may be a response that characterizes how an ear receives a sound from a point in space, and a pair of transfer functions HS associated with the user’s two ears, respectively, can be used to synthesize a binaural sound that seems to emanate from a particular point in the space, instead of from any of the multiple speakers.
[0039] Therefore, the speaker signals, which may comprise the binaural signal 230 transmitted by the multiple speakers 250 of the speaker array after being processed through the CTC filtering 240, may reproduce the desired sound arriving at the user’s left and right ears and can be synthesized to simulate a particular audio scene in the space, containing possibly a multitude of sources at different locations. The immersive spatial auditory experiences then can be achieved for the user from the speaker array that may be even better out-of-head sound effects.
[0040] Nevertheless, it shall be noted that, in one and more embodiments, the bass channel signal of the multi-channel signal 220, due to its poor directionality and being unsuitable for the CTC filtering, can be directly fed into and replayed by the multiple speakers 250, as shown in FIG, 2. As an example, the bass channel signal thus can be replayed via at least one of the multiple speakers 250, such as the at least one speaker behind the user’s head or ears.
[0041] In addition, in one and more embodiments, the center channel signal of the multi-channel signal 220 can be further directly distributed to the multiple speakers 250, without any binauralization nor CTC filtering processes, to enhance the externalization of signals from the front, such as speech dialogs. As an example, the center channel signal can be replayed via the at least one of the multiple speaker 250, such as the at least one speaker in front of the user’s head or ears.
[0042] FIG. 3 illustrates an exemplary flowchart 300 of the method to achieve immersive spatial audio experience with the wearable audio device, according to one or more embodiments of the present disclosure. At first, in step S310, for designing a wearable audio device, arranging multiple speakers as a speaker array into the wearable audio device. When a user wearing the wearable audio device, the multiple speakers therein can be located at different positions on the user’s body (or head, or other parts) , and, as noted, those positions may vary depending on the type of the wearable audio device.
[0043] In step S320, a multi-channel signal can be received from a multi-channel source. Firstly, any object-based signals, if provided, shall be decoded into the multi-channel signal.
[0044] In step S330, the multi-channel signal may be processed to a binaural signal that contains spatial information, where a pair of ideal HRTFs shall be applied to each channel therein. For such an HRTF, describing the transmission process of sound waves from the sound source to both ears, may contain both the direct sound and reverberation part, wherein the direct head-related impulse response can be obtained from open-source databases, and the reverberation part can be measured using the sweep sine method, or can be further simulated using the feedback delayed network method or the image source method. As noted, many ideal HRTFs already measured from human subjects in a laboratory exist, such as the CIPIC database, from which the pair of the ideal HRTFs retrieved based on each spatial content contained in a channel of the multi-channel signal can be used as the ideal HRTFs associated with that channel. That’s to say, the spatial content contained in each channel of the multi-channel signal may be processed with a corresponding pair of ideal HRTFs (for left and right ears) .
[0045] For each channel of the multi-channel signal, the spatial content therein may be convolved with their associated HRTF of both ears and generate the binaural signal for each channel. Thereby, the left-and right-binaural signals from each channel are summed together as follow: Pb, L=Pch1*H1, L+Pch2*H2, L+…+Pchn*Hn, L (1) Pb, R=Pch1*H1, R+Pch2*H2, R+…+Pchn*Hn, L (2)
[0046] where Pb, L and Pb, R are the binaural signals for the left and right ear of the user, respectively;
[0047] Pchn each represents the signal played through the nth channel;
[0048] Hn, L and Hn, R represent the HRIRs, denoting the corresponding HRTFs in spatial-domain form, from the nth channel of the multi-channel signal to the left and right ears of the user, respectively.
[0049] For the sake of simplicity and clarity, 5.1 surround sound can be taken as an example, but not only limited thereto, to set forth processing the multi-channel signal to the binaural signal herein. In an ideal multi-channel arrangement, a 5.1 multi-channel signal typically is comprised of center, left, right, left surround and right surround channels.
[0050] FIG. 4 illustrates an exemplary schematic diagram 400 of generating the binaural signals from 5.1 surround sound, according to one or more embodiments of the present disclosure. In the example as shown in FIG. 4, the 5.1 surround sound forms its multi-channel signal with six channels, which comprises a center channel, a left front channel Ch1, a right front channel Ch2, a right surround channel Ch3 and a left surround channel Ch4, each containing its source content positioned around the listener 410, separately, as well as a bass channel, as shown in FIG. 4.
[0051] Therefore, for the listener 410’s left ear 412, a left-binaural signal may be generated by summing each channel Ch 1…4 with its channel signal convoluting with its corresponding HRIR (i.e., the HRTF in spatial-domain form) for the left ear HRIR1L…4L, respectively; and for the listener 410’s right ear 414, a right-binaural signal may be generated by summing each channel Ch 1…4 with its channel signal convoluting with its corresponding HRIR for the right ear HRIR1L…4L, respectively, referring to FIG. 4.
[0052] The position of each channel relative to the listening position, such as the sound source with its spatial content herein, is recorded and carried in the multi-channel signal, so for this multi-channel signal, ready-made ideal HRTFs can be retrieved from the open-source databases according to those corresponding relative positions therein, to generate the binaural signal for the listener’s ears.
[0053] In one or more embodiments, the center channel of multi-channel sources may be directly transmitted to the user’s ears, via replayed by at least one speaker in the speaker array arranged in front of the user, for example, without the binauralization process, as shown in FIG. 4.
[0054] In one or more embodiments, the bass channel of multi-channel sources may be directly transmitted to the user’s ears, via replayed by at least one speaker in the speaker array arranged behind the user, for example, without the binauralization process, as shown in FIG. 4.
[0055] Based on the one or more alternative or additional embodiments mentioned above, the binaural signal, composed of a left-binaural signal and a right-binaural signal, may be generated based on at least one channel of the multi-channel signal. In an example, the binaural signal may be generated from those surround channels, excluding the bass channel and / or the center channel.
[0056] Now return to FIG. 3, then in step S340, after processing the binaural signal into the speaker signals through the CTC filtering, the speaker signals may be replayed to and reproduced the sound sources at the user’s ears by replaying via the multiple speakers of the speaker array. The aim here is to reproduce the right and left ear binaural signal without the interference from the other side channels.
[0057] The CTC filtering may eliminate or reduce the crosstalk inherent in the speaker playback so that the left channels of the binaural signal can be delivered substantially to the left ear only of the user and the right channels to the right ear only, thereby preserving the intention of the binaural signal. In this way, any spatial content brought into the binaural signals from the multi-channel source may be placed “virtually” in the 3D listening space since those speakers in the speaker array are not necessarily physically located at the point from which a rendered sound appears to emanate.
[0058] The CTC filtering can be designed based on a model of acoustic transmission from the speakers for playback of the binaural signal to the user’s ears. Each ear signal may be modeled as the sum of the speaker signals from the speaker array, so each speaker signal can be filtered by an associated CTC filter with transfer function modeling the inverse of the acoustic transmission from each speaker of the speaker array to that ear. Then, a set of transfer functions HS can be modeled using the head related transfer functions (HRTFs) , which shall be physically measured as a function of those speaker placements of the speaker array with respect to the user’s ears. As noted, an HRTF may be a response that characterizes how an ear receives a sound from a point in space. The set of measured transfer functions HS associated with those speakers in the speaker array for the user’s two ears can be used to synthesize the binaural signal that is physically transmitted by the speaker array, but seems to emanate from a particular point in space. The particular point may approximately have the same spatial content as its sound source.
[0059] The CTC filtering procedures herein for the binaural signal have limitations on its frequency band. In the frequency range of the binaural signal suitable for adopting the CTC filtering, usually within a middle range of frequency bandwidth, the associated pair of inversed transfer function each derived by inversing the transfer function HS, can be applied to each speaker. As noted, the transfer function HS here between each speaker and two ears are physically measured. The expression for acoustic transmission between the signals at each of the ears and the speakers for playback may be as follows: PE, L=Ps1*Hs1, L+Ps2*Hs2, L+…+Pn*Hsn, L (3) PE, R=Ps1*Hs1, R+Ps2*Hs2, R+…+Pn*Hsn, R (4)
[0060] where, PE, L and PE, R are signals received at the left and right ear, respectively;
[0061] Psn is signals played by the nth loudspeaker; and
[0062] Hsn, L and Hsn, R are the transfer functions measured between the nth loudspeaker and the user’s left / right ear, respectively. The above Equations (3) and (4) can be expressed in matrix format as follows: PE=HS*Ps (5)
[0063] And accordingly, the speaker array signals can be calculated by inverse calculation as follows:
[0064] Where the matrix is the pseudoinverse of the matrix HS.
[0065] The inversion of the speaker array transmission is underdetermined, therefore, those skilled in the art may conceive that the calculation of the inverted transfer function matrix has several methods including regularization methods.
[0066] As to different speaker arrangements of the speaker arrays for various wearable audio devices, in a practical implementation, when designing the layout of the speaker arrays for those wearable audio devices, a parametric model, such as the mannequin or the spherical head model, wearing the speaker array shall be used for physically constructing the measured HS here for the CTC filtering. Through such physical simulation, the inversed transfer function associated with each specific speaker in the speaker array of the wearable audio device to the user’s ears can be calculated, separately.
[0067] As noted, in one and more embodiments, the bass channel of the multi-channel signal, due to its poor directionality of the bass audio being unsuitable for the CTC filtering, can be directly fed into and replayed by at least one speaker of the multiple speakers, without any binauralization for CTC filtering.
[0068] In one or more embodiments, for the binaural signal above a certain frequency, which depends on the acoustic system, the CTC filtering may be too sensitive to the modeling of the acoustic transmission. Even a small change, for example a head rotation, or a different person, the CTC filtering performance may be degraded a lot. Above this frequency, the binaural signal can be played by the one speaker on each side that has the highest raw channel separation. Nevertheless, alternatively or additionally, it shall be necessary to compensate for the time delay between the user’s left ear and right ear in the binaural signal playback for each channel, for the binaural signal above the certain frequency.
[0069] In addition, in one and more embodiments, again as noted, the center channel of the multi-channel signal can be further directly distributed to the at least one speaker in front of the user, without any binauralization nor CTC filtering processes, to enhance the externalization of signals from the front, such as speech dialogs.
[0070] In the foregoing specification, the present disclosure has been described with reference to specific embodiments thereof. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader scope of the invention. For example, the above-described process flows are described with reference to a particular ordering of process actions. However, the ordering of many of the described process actions may be changed without affecting the scope or operation of the invention. The specification and drawings are, accordingly, to be regarded in an illustrative rather than restrictive sense.
[0071] Any combination of one or more computer-readable media may be used to perform the method provided in one or more embodiments of the present disclosure. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium may include, for example, an electrical connection with one or more wires, portable computer floppy disks, hard disks, random access memory (RAM) , read-read-only memory (ROM) , erasable programmable read-only memory (EPROM or flash memory) , optical fibers, portable compact disc read-only memory (CD-ROM) , optical storage devices, magnetic storage devices, or any suitable combinations of the foregoing. In the context of the present disclosure, the computer-readable storage medium may be any tangible medium that can include or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0072] As used in the present disclosure, an element or step listed in the singular form and preceded by the word “one / a” should be understood as not excluding a plurality of said elements or steps, unless such exception is specifically stated. Furthermore, references to “embodiments” or “examples” of the present disclosure are not intended to be construed as exclusive, also including the existence of other embodiments of the recited features. The terms “first” , “second” , “third” , etc. are used only for identification and are not intended to emphasize a numerical requirement or positioning order of their objects.
[0073] References in the present disclosure to the method for an intelligent anti-interference mode of a speaker system may include the following content:
[0074] Item 1: In one or more embodiments, the present disclosure provides a wearable audio device to achieve immersive spatial audio experience, comprising:
[0075] a speaker array comprising multiple speakers, wherein the multiple speakers are arranged at certain positions on a user’s body to replay speaker signals to the user; and
[0076] one or more processors configured to:
[0077] receive a multi-channel signal from a multi-channel source;
[0078] generate a binaural signal, composed of a left-binaural signal and a right-binaural signal, based on the multi-channel signal; and
[0079] process the binaural signal into the speaker signals through crosstalk cancellation filtering.
[0080] Item 2: The method of item 1, wherein a pair of ideal HRTFs are applied to each channel of the multi-channel signal to generate the binaural signal, respectively, and wherein the pair of ideal HRTFs are each retrieved from open-source databases based on multiple spatial contents in the multi-channel signal.
[0081] Item 3: The method of item 1 or 2, wherein the crosstalk cancellation filtering comprises applying a pair of inverses of measured transfer functions to each of the multiple speakers, respectively, to replay the binaural signal, and wherein the measured transfer functions are measured utilizing a mannequin wearing the speaker array in designing the arrangement of the multiple speakers.
[0082] Item 4. The method of any of items 1 to 3, wherein a center channel signal of the multi-channel signal is directly distributed to at least one of the multiple speakers in front of the user’s ears, without any binauralization processing and the crosstalk cancellation filtering.
[0083] Item 5. The method of any of items 1 to 4, wherein a bass channel signal of the multi-channel signal is replayed with at least one of the multiple speakers, without any binauralization processing and the crosstalk cancellation filtering.
[0084] Item 6. The method of any of items 1 to 5, wherein the binaural signal above a certain frequency are replayed with one speaker on each side of the user that has the highest raw channel separation, respectively, without the crosstalk cancellation filtering.
[0085] Item 7. The method of any of items 1 to 6, wherein the speaker array is arranged as a neckband speaker array.
[0086] Item 8. In one or more embodiments, the present disclosure provides a method to achieve immersive spatial audio experience with a wearable audio device, comprising following steps:
[0087] arranging multiple speakers into a speaker array at certain positions on a user’s body to replay speaker signals to the user;
[0088] receiving, in one or more processors, a multi-channel signal from a multi-channel source;
[0089] generating a binaural signal, composed of a left-binaural signal and a right-binaural signal, based on the multi-channel signal; and
[0090] processing the binaural signal into the speaker signals through crosstalk cancellation filtering.
[0091] Item 9. The method of item 8, further comprising applying a pair of ideal HRTFs to each channel of the multi-channel signal to generate the binaural signal, respectively, and wherein the pair of ideal HRTFs are each retrieved from open-source databases based on multiple spatial contents in the multi-channel signal.
[0092] Item 10. The method of item 8 or 9, further comprising applying a pair of inverses of measured transfer functions to each of the multiple speakers, respectively, to replay the binaural signal, and wherein the measured transfer functions are measured utilizing a mannequin wearing the speaker array in designing the arrangement of the multiple speakers.
[0093] Item 11. The method of any of items 8 to 10, further comprising directly distributing a center channel signal of the multi-channel signal to one or two speakers of the multiple speakers in front of the user’s ears, without any binauralization processing and the crosstalk cancellation filtering.
[0094] Item 12. The method of any of items 8 to 11, further comprising replaying the ear binaural signals above a certain frequency with one speaker on each side of the user that has the highest raw channel separation, respectively, without the crosstalk cancellation filtering.
[0095] Item 13. The method of any of items 8 to 12, further comprising replaying a bass channel signal of the multi-channel signal with at least one of the multiple speakers, without any binauralization processing and the crosstalk cancellation filtering.
[0096] Item 14. The method of any of items 8 to 13, further comprising arranging the speaker array as a neckband speaker array.
Claims
1.A wearable audio device to achieve immersive spatial audio experience, comprising:a speaker array comprising multiple speakers, wherein the multiple speakers are arranged at certain positions on a user’s body to replay speaker signals to the user; andone or more processors configured to:receive a multi-channel signal from a multi-channel source;generate a binaural signal, composed of a left-binaural signal and a right-binaural signal, based on the multi-channel signal; andprocess the binaural signal into the speaker signals through crosstalk cancellation filtering.2.The wearable audio device of claim 1, wherein a pair of ideal HRTFs are applied to each channel of the multi-channel signal to generate the binaural signal, respectively, and wherein the pair of ideal HRTFs are each retrieved from open-source databases based on multiple spatial contents in the multi-channel signal.3.The wearable audio device of claim 1, wherein the crosstalk cancellation filtering comprises applying a pair of inverses of measured transfer functions to each of the multiple speakers, respectively, to process the binaural signal into the speaker signals, and wherein the measured transfer functions are measured utilizing a mannequin wearing the speaker array in designing the arrangement of the multiple speakers.4.The wearable audio device of claim 1, wherein a center channel signal of the multi-channel signal is directly distributed to at least one of the multiple speakers in front of the user’s ears, without any binauralization processing and the crosstalk cancellation filtering.5.The wearable audio device of claim 1, wherein a bass channel signal of the multi-channel signal is replayed with at least one of the multiple speakers, without any binauralization processing and the crosstalk cancellation filtering.6.The wearable audio device of claim 1, wherein the binaural signal above a certain frequency are replayed with at least one speaker on each side of the user that has the highest raw channel separation, respectively, without the crosstalk cancellation filtering.7.The wearable audio device of any one claims 1-6, wherein the speaker array is arranged as a neckband speaker array.8.A method to achieve immersive spatial audio experience with a wearable audio device, comprising following steps:arranging multiple speakers into a speaker array at certain positions on a user’s body to replay speaker signals to the user;receiving, in one or more processors, a multi-channel signal from a multi-channel source;generating a binaural signal, composed of a left-binaural signal and a right-binaural signal, based on the multi-channel signal; andprocessing the binaural signal into the speaker signals through crosstalk cancellation filtering.9.The method of claim 8, further comprising applying a pair of ideal HRTFs to each channel of the multi-channel signal to generate the binaural signal, respectively, and wherein the pair of ideal HRTFs are each retrieved from open-source databases based on multiple spatial contents in the multi-channel signal.10.The method of claim 8, further comprising applying a pair of inverses of measured transfer functions to each of the multiple speakers, respectively, to process the binaural signal into the speaker signals, and wherein the measured transfer functions are measured utilizing a mannequin wearing the speaker array in designing the arrangement of the multiple speakers.11.The method of claim 8, further comprising directly distributing a center channel signal of the multi-channel signal to one or two speakers of the multiple speakers in front of the user’s ears, without any binauralization processing and the crosstalk cancellation filtering.12.The method of claim 8, further comprising replaying the ear binaural signals above a certain frequency with one speaker on each side of the user that has the highest raw channel separation, respectively, without the crosstalk cancellation filtering.13.The method of claim 8, further comprising replaying a bass channel signal of the multi-channel signal with at least one of the multiple speakers, without any binauralization processing and the crosstalk cancellation filtering.14.The method of any one of claims 8-13, further comprising arranging the speaker array as a neckband speaker array.