System and method for providing three-dimensional immersive sound
The system uses narrowband loudspeakers and psychoacoustic direction-determining bands to enhance 3D immersive sound localization and reduce computational load, addressing limitations in existing technologies.
Patent Information
- Application Number
- JP2022006915
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-01
- Filing Date
- 2022-01-20
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2042-01-20
AI Technical Summary
Current wideband loudspeaker placements and digital signal processing techniques for 3D immersive sound suffer from limited sound image localization, computational intensity, and reliance on room geometry, while headphones lack low-frequency reproduction and are limited in use scenarios.
A system utilizing narrowband loudspeakers and psychoacoustic direction-determining bands, including controllers that store and analyze psychoacoustic measures in subbands to generate loudspeaker drive signals, enhancing sound localization and minimizing computational load.
The system provides 3D immersive sound with improved spatial perception and reduced distortion, independent of loudspeaker placement, and reduces the number of channels and DSP computational requirements.
Smart Images

Figure 0007814949000001 
Figure 0007814949000002 
Figure 0007814949000003
Abstract
Description
[Technical Field]
[0001] Aspects disclosed herein generally relate to systems and methods for three-dimensional (3D) immersive sound. In one example, the systems and methods for providing 3D immersive sound may be based on at least one of psychoacoustic direction-determining bands and narrowband loudspeakers. These and other aspects are described in more detail herein. [Background technology]
[0002] Current wideband loudspeaker placements have many drawbacks. One drawback is limited sound image localization, which is consistent with respect to where the loudspeakers are positioned. For example, front loudspeakers may be localized in front of the listener's position, rear loudspeakers may be localized behind the listener's position, etc. Another drawback is that many digital signal processing (DSP) techniques used to achieve virtual height effects either have a limited listener sweet spot and are computationally intensive, or they rely on sound field obstructions and room geometry to reflect sound sources.
[0003] For narrowband loudspeaker placement, the auditory system forms a sound sensation with a direction that depends only on the frequency of the signal. The psychoacoustic relationship between signal frequency and direction of sound sensation can be described by the Blauert direction decision band (BDB).
[0004] Headphones are also another way to create 3D immersive sound, but their use is limited and / or prohibited in certain situations, such as while driving a car. Additionally, headphones are not capable of reproducing the low-frequency vibrations generated by loudspeakers, especially subwoofers. Summary of the Invention [Means for solving the problem]
[0005] In one embodiment, a system for providing three-dimensional (3D) immersive sound is provided. The system includes a loudspeaker and at least one controller. The loudspeaker transmits an audio output signal in a listening environment. The at least one controller is programmed to store a plurality of direction determination bands, each direction determination band defined by a narrowband frequency interval, and to store at least a psychoacoustic measure for each direction determination band, the psychoacoustic measure including subbands. The at least one controller is further programmed to determine energy in the subbands and, based on the energy in the at least subbands, generate a loudspeaker drive signal to drive the loudspeaker and transmit the audio output signal.
[0006] In at least another embodiment, a computer program product embodied in a non-transitory computer-readable medium is provided that is programmed to provide three-dimensional (3D) immersive sound. The computer program product includes instructions for transmitting an audio output signal in a listening environment and instructions for storing a plurality of direction determination bands, each direction determination band defined by a narrowband frequency interval. The computer program product includes instructions for storing at least psychoacoustic measures including subbands in each direction determination band and instructions for determining the energy of the subbands. The computer program product includes instructions for generating a loudspeaker drive signal based on the energy of the subbands to drive the loudspeaker to transmit the audio output signal.
[0007] In at least another embodiment, a method for providing three-dimensional (3D) immersive sound is provided. The method includes transmitting an audio output signal in a listening environment and storing a plurality of direction determination bands, each direction determination band defined by a narrowband frequency interval. The method includes storing at least psychoacoustic measures including subbands in each direction determination band and determining energy of the subbands. The method includes generating loudspeaker drive signals based on at least the energy of the subbands to drive loudspeakers to transmit the audio output signal.
[0008] The embodiments of the present disclosure are pointed out with particularity in the appended claims. However, other features of the various embodiments will become more apparent and will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings. For example, the present application provides the following: (Item 1) 1. A system for providing three-dimensional (3D) immersive sound, the system comprising: a loudspeaker for transmitting an audio output signal in the listening environment; at least one controller, wherein the at least one controller: storing a plurality of direction determination bands, each direction determination band being defined by a narrowband frequency interval; storing at least a psychoacoustic measure including subbands in each direction determination band; determining the energy of said subbands; generating a loudspeaker drive signal based on the energy in at least the subband to drive the loudspeaker and transmit the audio output signal; The system is programmed to perform the following steps: (Item 2) The system described in the preceding item, wherein the at least one controller is further programmed to determine a difference between the energy of the subband and a masked hearing threshold. (Item 3) 10. The system of claim 9, wherein the masked hearing threshold corresponds to an audible signal audible by a listener. (Item 4) 20. The system of claim 19, wherein the at least one controller is further programmed to compare the difference with one or more threshold values. (Item 5) 10. The system of claim 9, wherein the at least one controller is further programmed to apply a gain to the loudspeaker drive signal based on a comparison of the difference to the one or more thresholds. (Item 6) 2. The system of claim 1, wherein the gain performs one of increasing the directivity of the audio output signal or minimizing distortion of the audio output signal. (Item 7) 10. The system of claim 9, wherein the plurality of direction determination bands correspond to a plurality of Blauert direction determination bands. (Item 8) 10. The system of claim 1, wherein the at least one psychoacoustic measure is at least one Bark measure. (Item 9) 1. A computer program product embodied in a non-transitory computer-readable medium that is programmed to provide three-dimensional (3D) immersive sound, the computer program product including instructions that: transmitting an audio output signal in a listening environment; storing a plurality of direction determination bands, each direction determination band being defined by a narrowband frequency interval; storing at least a psychoacoustic measure including subbands in each direction determination band; determining the energy of said subbands; generating a loudspeaker drive signal based on the energy in at least the subband to drive the loudspeaker and transmit the audio output signal; The computer program product. (Item 10) 10. The computer program product of claim 9, further comprising instructions for determining the difference between the energy of the subband and a masked hearing threshold. (Item 11) 10. The computer program product of claim 1, wherein the masked hearing threshold corresponds to an audible signal that is audible by a listener. (Item 12) 10. The computer program product of claim 1, further comprising instructions for comparing the difference to one or more threshold values. (Item 13) 10. The computer program product of claim 1, further comprising instructions for applying a gain to the loudspeaker drive signal based on the comparison of the difference to the one or more thresholds. (Item 14) 2. The computer program product of claim 1, wherein the gain performs one of increasing the directionality of the audio output signal or minimizing distortion of the audio output signal. (Item 15) 10. The computer program product of claim 1, wherein the plurality of direction determination bands correspond to a plurality of Blauert direction determination bands. (Item 16) 10. The computer program product of claim 1, wherein the at least one psychoacoustic measure is at least one Bark scale. (Item 17) 1. A method for providing three-dimensional (3D) immersive sound, the method comprising: transmitting an audio output signal in a listening environment; storing a plurality of direction determination bands, each direction determination band being defined by a narrowband frequency interval; storing at least a psychoacoustic measure including subbands in each direction determination band; determining the energy of said subbands; generating a loudspeaker drive signal based on the energy in at least the subband to drive the loudspeaker and transmit the audio output signal; The above method. (Item 18) 10. The method of claim 9, further comprising instructions for determining the difference between the energy of the subband and a masked hearing threshold. (Item 19) 10. The method of any one of the preceding claims, further comprising instructions for comparing the difference to one or more threshold values. (Item 20) 10. The method of claim 1, further comprising instructions for applying a gain to the loudspeaker drive signal based on the comparison of the difference to the one or more thresholds. (Summary) In one embodiment, a system for providing three-dimensional (3D) immersive sound is provided. The system includes a loudspeaker and at least one controller. The loudspeaker transmits an audio output signal in a listening environment. The at least one controller is programmed to store a plurality of direction determination bands, each direction determination band defined by a narrowband frequency interval, and to store at least a psychoacoustic measure for each direction determination band, the psychoacoustic measure including subbands. The at least one controller is further programmed to determine energy in the subbands and, based on the energy in the at least subbands, generate a loudspeaker drive signal to drive the loudspeaker and transmit the audio output signal. [Brief explanation of the drawings]
[0009] [Figure 1] 1 shows the corresponding listener's 3D immersive sound sensation plane, which is divided into a median plane and an upper part of the median plane. [Figure 2] A schematic diagram of narrowband sound localization in the median plane, independent of the sound source position, is shown. [Figure 3A]1 shows various example placements of psychoacoustic loudspeakers, subwoofers, and tweeters in a first configuration of the listening environment. [Figure 3B] 10 shows various example placements of psychoacoustic loudspeakers, subwoofers, and tweeters in a second configuration of the listening environment. [Figure 4] 1 shows the relationship between the Blauert direction decision band and the critical subband. [Figure 5] 1 shows the psychoacoustic Bark scale including critical subbands and frequency ranges. [Figure 6] 1 illustrates a system for providing three-dimensional immersive sound based on at least one psychoacoustic direction determining band and narrowband loudspeakers, according to one embodiment. [Figure 7] 10 shows a plot illustrating an example of a smoothing filter for a selected BDB band that enhances frequencies within the range of the BDB while attenuating frequencies outside the range of the BDB, according to one embodiment. [Figure 8] 1 illustrates a method for providing three-dimensional immersive sound based on at least one psychoacoustic direction determining band and narrowband loudspeakers, according to one embodiment. [Figure 9] 1 illustrates an example of a system for providing three-dimensional immersive sound based on at least one psychoacoustic direction determining band and narrowband loudspeakers, according to one embodiment. [Figure 10] 1 illustrates another example of a system for providing three-dimensional immersive sound based on at least one psychoacoustic direction determining band and narrowband loudspeakers, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Where necessary, detailed embodiments of the present invention are disclosed herein; however, it should be understood that the disclosed embodiments are merely examples of the invention, which may be embodied in various and alternative forms. The figures are not necessarily to scale. Some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not limiting, but should merely be construed as a representative basis for teaching those skilled in the art how to variously use the present invention.
[0011] It will be appreciated that the controllers / devices disclosed herein may include any number of microprocessors, integrated circuits, memory devices (e.g., flash memory, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other suitable variants), and software that interact to perform the operation(s) disclosed herein. Additionally, such disclosed controllers utilize one or more microprocessors to execute computer programs embodied in non-transitory computer-readable media that are programmed to perform any number of the disclosed functions. Additionally, the controller(s) provided herein include a housing, various numbers of microprocessors, integrated circuits, and memory devices (e.g., flash memory, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) positioned within the housing. Additionally, the disclosed controller(s) each include hardware-based inputs and outputs for sending and receiving data to and from other hardware-based devices described herein. Although various system, block, and / or flow diagrams described herein refer to the time domain, frequency domain, etc., it is recognized that such system, block, and / or flow diagrams may be implemented in one or more of the time domain, frequency domain, etc.
[0012] Current technologies for delivering 3D immersive sound all around and around the listener's position fall into two categories: For example, the first category may use multiple loudspeakers utilizing surround sound technologies such as 5.1 and 7.1. These corresponding surround sound technologies add height channels to the system. Consequently, fully immersive 3D sound is possible by adding loudspeakers in the ceiling and upward-facing speakers, bouncing the sound off higher surfaces. Newer configurations such as 11.2 or 22.4 are examples of such arrangements.
[0013] The second category for delivering 3D immersive sound includes sound bars. For example, existing sound bar technology relies on multiple loudspeakers arranged in a linear array. Some loudspeakers point directly across the median plane, while others are aimed beyond the listening position and rely on sound reflected off surfaces and around the listener's position. Furthermore, some sound bars may include additional digital signal processing (DSP) techniques, such as phase and magnitude correction, to direct individual channels of sound to specific locations around the listening position.
[0014] Unlike the current technologies described above, the embodiments disclosed herein provide 3D immersive sound while, among other things, minimizing the number of loudspeaker channels, being independent of loudspeaker placement and sound directionality, and minimizing DSP computational load. Furthermore, the embodiments disclosed herein may generally rely on psychoacoustic concepts such as critical subbands (CSBs) (or subbands of the Bark scale (or psychoacoustic measure)), Blauert direction decision bands (BDBs) (or direction decision bands), masking thresholds, substantially enhanced acoustic imaging, etc. These and other embodiments are described in more detail below.
[0015] 1 shows a 3D immersive sound sensation plane 100 for a listener (or user) 102 divided into various planes (or sectors) 104a-104c. For example, plane 104a may be defined as the rear-superior median plane (or RU plane) relative to the listener 102, plane 104b may be defined as the upper median plane (or TOP plane) relative to the listener 102, and plane 104c may be defined as the front-superior median plane (or FU plane) relative to the listener 102. Generally, 3D immersive sound provides the listener(s) 102 with an improved perception of spatial dimensions over mono, stereo, and surround mixes. On the other hand, sound localization in mono, stereo, and surround mixes may be limited to within ±15 degrees of horizontal relative to the median plane 106 of the listener 102. The 3D immersive sound sensation is distributed above the median plane 106 (eg, planes 104a-104c) in addition to the horizontal median plane.
[0016] FIG. 2 shows a schematic diagram 120 of narrowband sound localization in the median plane 106, regardless of the location of the sound source. Psychoacoustic studies have shown that narrowband sound localization can be perceived as coming from a specific direction, regardless of the location of the sound source. In other words, the human auditory system creates a sound sensation with a direction that depends on the frequency of the audio signal. The psychoacoustic function between signal frequency and the direction of the sound sensation can be explained by Blauert's direction determination band, as shown in FIG. 2 below (see also J. Blauert, "Sound Localization in the Median Plane," Acta Acustica 22(4), pp. 205-13, November 1969, and H. Fastl and E. Zwicker, "Psychoacoustics Facts and Models," Third Edition, Springer 2007).
[0017] For example, when a narrowband sound having a center frequency of 300 Hz or 3 kHz is presented to the listener 102, the sound stage is perceived by the listener 102 at the FU plane 104c of the median plane 106. For example, a narrowband sound centered at 8 kHz is perceived as coming from the TOP plane 104b of the median plane 106, even if the sound source is in front of the listener 102. For example, a narrowband sound centered at 1 kHz or 10 kHz is perceived to originate at the RU plane 104a of the median plane 106, regardless of the actual location of the sound source.
[0018] 3A shows an example implementation 150 of various arrangements or locations of psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a, subwoofer 158, and tweeter 160 in listening environment 161. Generally, the number of psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a implemented is based at least on the number of Blauert direction determination bands (BDBs). The psychoacoustic loudspeakers 152a, 152b may be oriented to provide sound to listener 102 at FU side 104c of listening environment 161. The psychoacoustic loudspeakers 154a, 154b may be oriented to provide sound to listener 102 at RU side 104a of listening environment 161. Psychoacoustic loudspeaker 156a may be oriented to provide sound to TOP surface 104b of listening environment 161. Subwoofer 158 and tweeter 160 complement psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a to provide sound in the low frequency range (e.g., subwoofer range) and high frequency range (e.g., tweeter range), respectively. For clarity, psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a will be identified as actual physical loudspeakers. Sound source 159 may be positioned within listening environment 161 and transmit sound to the various psychoacoustic loudspeakers 152a-152b, 154a-154b, 156a, subwoofer 158, and tweeter 160 for reproduction in listening environment 161.
[0019] In general, the placement or location of one or more psychoacoustic loudspeakers 152a-152b, 154a-154b, 156a can be independent of the location of the desired sound source (or audio source 159). This is further illustrated in embodiment 170 of FIG. 3B, in which psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a are all positioned in front of listener 102. In contrast, in FIG. 3A, psychoacoustic loudspeakers 152a and 154a are positioned behind listener 102a and behind psychoacoustic loudspeakers 152b, 154b, and 156a. Due to its omnidirectionality, subwoofer 158 can be placed anywhere in the room enclosure (or listening environment 161). Tweeter 160 can be placed in front of listener 102 due to the directionality of its focused beam. Generally, both implementations 150, 170 each produce a comparable 3D immersive effect.
[0020] Psychoacoustic speakers 152a-152b, 154a-154b, and 156a may be a combination of individual narrowband speakers that include a psychoacoustic critical subband scale, such as the Bark scale or the Equivalent Rectangular Bandwidth (ERB) scale or the Mel scale. Additionally or alternatively, any one of psychoacoustic speakers 152a-152b, 154a-154b, and 156a may be a single loudspeaker that covers the BDB frequency range.
[0021] FIG. 4 illustrates the relationship between Blauert direction determination bands (BDBs) and critical subbands (CSBs) for various psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a. FIG. 5 illustrates the corresponding Blauert direction determination bands and frequencies, which are referenced in connection with the description of FIG. 4 below. The CSBs are designated as Bark numbers (e.g., 1-25), and the corresponding BDBs include groups of CSBs that define frequency ranges. Generally, as shown in psychoacoustic loudspeaker 152a (e.g., an FU1-based loudspeaker), psychoacoustic loudspeaker 152a may include four separate narrowband loudspeakers covering Bark bands 3, 4, 5, and 6 (see FIGS. 4 and the values under the "Bark" heading in FIG. 5), or one loudspeaker with a programmable center frequency ranging from 250 Hz to 570 Hz (see the values under the "Center Frequency (Hz)" heading in FIG. 5), or a combination of any group of these four Bark bands. The psychoacoustic loudspeaker 154a (e.g., an RU1-based loudspeaker) includes seven separate narrowband speakers covering Bark bands 7, 8, 9, 10, 11, 12, and 13 (see Figures 4 and 5 under the heading "Bark"), or one loudspeaker with a programmable center frequency ranging from 700 Hz to 1850 Hz (see Figure 5 under the heading "Center Frequency (Hz)"), or a combination of any group of these seven Bark bands.
[0022] Psychoacoustic loudspeaker 152b (e.g., an FU2-based loudspeaker) includes eight separate narrowband loudspeakers covering Bark bands 14, 15, 16, 17, 18, 19, 20, and 21 (see FIGS. 4 and 5 under the heading "Bark"), or one loudspeaker with a programmable center frequency in the range of 2150 Hz to 7000 Hz (see FIG. 5 under the heading "Center Frequency (Hz)"), or any grouping combination of these eight Bark bands. Psychoacoustic loudspeaker 156a (e.g., a TOP loudspeaker) includes a single narrowband loudspeaker covering Bark band 22 (see FIG. 4 and 5 under the heading "Bark"), or a single loudspeaker with a programmable center frequency in the range of 8500 Hz (see FIG. 5 under the heading "Center Frequency (Hz)").
[0023] Psychoacoustic loudspeaker 154b (e.g., RU2 loudspeaker) comprises two narrowband loudspeakers covering Bark bands 23 and 24 (see Figures 4 and 5 under the heading "Bark"), or a single loudspeaker with a programmable center frequency in the range of 10,500 Hz to 13,500 Hz (see Figure 5 under the heading "Center Frequency (Hz)"). Loudspeaker 158 (e.g., subwoofer) comprises two narrowband loudspeakers covering Bark bands 1 and 2 (see Figures 4 and 5 under the heading "Bark"), or a single loudspeaker with a programmable center frequency in the range of 50 Hz to 150 Hz (see Figure 5 under the heading "Center Frequency (Hz)"). The loudspeaker 160 (e.g., a tweeter loudspeaker) comprises a single narrowband loudspeaker covering the Bark band 25 (see FIG. 4 and the values under the heading "Bark" in FIG. 5) or a loudspeaker with a programmable center frequency in the range of 17,750 Hz (see the values under the heading "Center Frequency (Hz)" in FIG. 5). Generally, embodiments disclosed herein provide, without limitation, systems and methods for modifying the energy of CSB and BDB to increase the directivity factor while minimizing any additional distortion. For example, the spectral components of CSB and DBD can enhance the perceived sound image without using loudspeakers at physical height.
[0024] 6 illustrates a system 300 for providing three-dimensional immersive sound based on at least one psychoacoustic direction-determining band and narrowband loudspeaker, according to one embodiment. The system 300 includes at least one controller 302 (hereinafter "controller 302") operably coupled to a plurality of loudspeakers 304 (e.g., psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a, a subwoofer 158, and a tweeter 160). It will be appreciated that the controller 302 may include any number of digital signal processors (DSPs) and is generally programmed to provide input audio signals to the plurality of loudspeakers 304 for playback by the listener 102 in the listening environment 161.
[0025] The controller 302 includes a first filter bank 304, a mixing matrix block 306, a crossover network 308 (e.g., a Blauert crossover network 308), a psychoacoustic modeling block 310, a gain block 312, and a second filter bank 314. An input audio signal may be split into a right channel and a left channel, and both channel signals are provided to the first filter bank 304. The first filter bank 304 converts the channel signals from the time domain to the frequency domain. The first filter bank 304 may map the frequency-domain channel signals to a set of M critical subbands (CSBs) according to the Bark scale, the Mel scale, or the ERB scale. For example, the mapping performed by the first filter bank 304 may be a linear transformation of discrete frequencies in the Hertz scale to discrete subbands in the Bark scale, the Mel scale, or the ERB scale.
[0026] The mixing matrix block 306 may reduce or increase the number of input channels to match the number N of loudspeakers by applying various scaling factors. In the example of FIG. 6, the N output channels from the mixing matrix block 306 may be equal to a linear combination of the left and right input channels from the analysis filter block 304 in the case of a stereo input signal. For example, channel 1 = 0.5*inputR + 0.5*inputL, and similarly for the other N-1 channels. In this example, the 0.5 scaling factor is real, but the scaling factor may also be complex. The crossover network 308 groups the BDBs into the various loudspeakers 152a-152b, 154a-154b, 156a, 158, and 160 according to a pre-defined mapping of CSBs as illustrated in the example shown in FIG. 4. As discussed in connection with FIG. 4, the CSBs are designated as Bark numbers (e.g., 1-25), and the corresponding BDBs contain groups of CSBs that define frequency ranges.
[0027] The psychoacoustic modeling block 310 calculates the energy, the masked hearing threshold, and the difference (or delta (Δ)) between the energy and the masked hearing threshold of each CSB in the BDB. The energy of a CSB is the squared magnitude of the complex number associated with the CSB calculated by the filter bank block 304. The masked hearing threshold of a CSB in the BDB is the sound level below which any CSB energy is inaudible, and above which any energy level is audible to humans. The calculation of the masked threshold may be based on the psychoacoustic model described in H. Fastl and E. Zwicker, "Psychoacoustics Facts and Models," Third Edition, Springer 2007, introduced above. The psychoacoustic modeling block 310 calculates the delta (Δ) (or the difference between the energy and the masked hearing threshold) of each CSB in the BDB. The gain block 312 applies gain to the N channels from the crossover network block 308 to either amplify or attenuate the energy of the CSB. By either amplifying or attenuating the amount of energy in each CSB within the BDB, this embodiment can increase the directivity factor of a particular loudspeaker while minimizing any additional distortion. This embodiment is described in more detail in connection with FIG.
[0028] The second filter bank 314 converts the BDB loudspeaker channels from the frequency domain back to the time domain, where the second filter bank 314 also applies smoothing filters. The smoothing filter for a given BDB band is chosen to boost frequencies within the range of the BDB while attenuating frequencies outside the range of the BDB. This is further illustrated in FIG. 7, which shows an example of a single CSB #22 and a BDB with a center frequency of 8.5 kHz. Generally, the BDD loudspeaker channels correspond to the various channels associated with the psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a (e.g., loudspeakers transmitting audio on the FU1, FU2, RU1, RU2, and TOP planes). Time-domain based narrowband signals (or loudspeaker drive signals) are used to drive the multiple loudspeakers 304 with the available amplification.
[0029] 8 illustrates a method 400 for providing three-dimensional immersive sound based on at least one psychoacoustic direction-determining band and narrowband loudspeakers, according to one embodiment. At operation 402, the controller 302 loops through the various BDB groups stored in its memory (e.g., the BDB groups of associated psychoacoustic loudspeakers 152a-152b, 154a-154b, and 156a, subwoofer 158, and tweeter 160). Similarly, at operation 404, the controller 302 loops through the various CSB (or Bark Scale) groups for each BDB group.
[0030] In operation 406, the controller 302 calculates the energy of each CSB. Similarly, the controller 302 calculates the difference (or delta (Δ)) between the calculated energy and the masked hearing threshold for each CSB in the BDB group. In operation 408, the controller 302 compares delta (Δ) with a first threshold T1 and a second threshold T2. It is recognized that the first threshold T1 and the second threshold T2 correspond to predetermined values and may vary based on desired criteria of a particular implementation. If the controller 302 determines that delta (Δ) is greater than the first threshold T1 and less than the second threshold T2, the method 400 proceeds to operation 416. Otherwise, the method proceeds to operation 410 and operation 412.
[0031] At operation 410, the controller 302 determines whether delta (Δ) is less than a first threshold T1. If this condition is true, the method 400 proceeds to operation 414, whereby the controller 302 applies a first gain G1 via the gain block 312 to the CSB that satisfies the condition set forth in operation 410 (e.g., the audio output corresponding to the CSB (or Bark Scale #) that includes a lower frequency limit, an upper frequency limit, a center frequency, and a bandwidth). At operation 414, the controller 302 applies the first gain G1 to a single CSB in the BDB group. It will be appreciated that the first gain G1 may correspond to an attenuation gain (a reduction gain) or a gain that increases the audio output (or a gain that attenuates (a reduction gain) or increases the audio output of a single CSB in the BDB group). Thus, the net result of applying a first gain G1 to a single CSB in a BDB group results in the generation of a drive signal for driving the corresponding psychoacoustic loudspeaker 152a-152b, 154a-154b, or 156a, which outputs sound at the center frequency specified by the CSB with such gain. After applying all gains to the frequency-domain CSB, the controller 302 converts the N-channel signal to the time domain via the second filter bank block 314 and applies a smoothing filter at the selected center frequency as described above. It will be further recognized that the first gain G1 may correspond to a real and / or complex number. As described above, increasing the gain applied to a corresponding CSB (e.g., the first gain G1, the second gain G2, and the third gain G3) may increase the directivity factor of that CSB. Conversely, decreasing the gain applied to a corresponding CSB may decrease the distortion of that CSB.
[0032] At operation 412, the controller 302 also determines whether delta (Δ) is greater than a second threshold T2. If this condition is true, the method 400 proceeds to operation 418, whereby the controller 302 applies a third gain G3 via the gain block 312 to the CSB that satisfies the condition set forth in operation 412 (e.g., the audio output corresponding to the CSB (or Bark Scale #) that includes a lower frequency limit, an upper frequency limit, a center frequency, and a bandwidth). At operation 418, the controller 302 applies the third gain G1 to a single CSB in the BDB group. It will be appreciated that the third gain G3 may correspond to an attenuation gain (reduction gain) or a gain that increases the audio output (or a gain that attenuates (reduction gain) or increases the audio output of a single CSB in the BDB group). Thus, the net result of applying the first gain G3 to a single CSB in a BDB group results in the generation of a drive signal for driving the corresponding psychoacoustic loudspeaker 152a-152b, 154a-154b, or 156a that outputs sound at the center frequency specified by the CSB at such gain. It is further recognized that the third gain G3 may correspond to a real and / or complex number.
[0033] In operation 416, the controller 302 applies, via the gain block 312, a second gain G2 to the CSB that satisfies the conditions set forth in operation 408 (e.g., the audio output corresponding to the CSB (or Bark Scale #) including the lower limit frequency, the upper limit frequency, the center frequency, and the bandwidth). In operation 416, the controller 302 applies a third gain G3 to a single CSB in the BDB group. It is recognized that the second gain G2 may correspond to an attenuation gain (reduction gain) or a gain that increases the audio output. It is recognized that the second gain G2 may correspond to an attenuation gain (reduction gain) or a gain that increases the audio output (or a gain that attenuates (reduction gain) or increases the audio output of a single CSB in the BDB group). Thus, the net result of applying the second gain G2 to a single CSB in a BDB group results in the generation of a drive signal for driving the corresponding psychoacoustic loudspeaker 152a-152b, 154a-154b, or 156a that outputs sound at the center frequency specified by the CSB at such gain. It is further recognized that the second gain G2 may correspond to a real and / or complex number.
[0034] In operation 420, the controller 302 determines whether all CSBs (i.e., Bark measures) of a particular BDB have been verified with respect to analysis for delta (Δ), comparison with thresholds T1, T2, and T3, and application of first gain G1, second gain G2, and third gain G3. If all CSBs of a particular BDB have been verified, the method 400 proceeds to operation 422. Otherwise, the method 400 returns to operation 404 and loops to the next CSB that needs to be verified.
[0035] In operation 422, the controller 302 determines whether all BDBs have been verified. If all BDBs have been verified, the method 400 stops. If all BDBs have not been verified, the method 400 returns to operation 402 to verify the next BDB.
[0036] FIG. 9 illustrates an exemplary system 500 for providing three-dimensional immersive sound based on at least one psychoacoustic direction-determining band and narrowband loudspeakers, according to one embodiment. The system 500 illustrated in connection with FIG. 9 is substantially similar to the system 300 illustrated in connection with FIG. 6. However, the system 500 illustrates that the audio input signal is a single input audio signal. In this case, the mixing matrix block 306 upmixes the single mono input channel into N output channels corresponding to the number of loudspeakers. The Nth output channel is given as a scaled version of the single input channel, e.g., Channel1 = A1 * InputR (where A1 corresponds to a multiplication factor, and A2 through A7 are also applied). The mixing matrix block 306 illustrated in FIG. 9 illustrates that the amplitude for the left channel is zeroed, assuming that the system 500 receives only a mono input audio signal. The crossover network block 308 illustrates that, for example, a Bark scale of 25 (see FIG. 5) is applied to the mono input audio signal. As mentioned above, one or more of the 25 Burke scales (or CSBs) are grouped into the BDB.
[0037] FIG. 10 illustrates an exemplary system 600 for providing three-dimensional immersive sound based on at least one psychoacoustic direction-determining band and narrowband loudspeakers, according to one embodiment. The system 600 illustrated in connection with FIG. 10 is substantially similar to the system 300 illustrated in connection with FIG. 6. The system 600 also illustrates that the audio input signal is a stereo input audio signal. In this case, the mixing matrix block 306 illustrated in FIG. 9 indicates the amplitudes of the left and right channels, assuming the system 600 receives a stereo input audio signal. The mixing matrix block 306 upmixes the dual stereo input channels into N output channels corresponding to the number of loudspeakers. The N output channels are provided as scaled versions of the stereo input channels, e.g., Channel 1 = A1 * InputR + B1 * InputL, Channel 2 = A2 * InputR + B2 * InputL, where A1 through A7 and B1 through B7 correspond to multiplication factors. In the crossover network block 308, for example, 25 Bark scales (see FIG. 5) are shown to be applied to the mono input audio signal. As mentioned above, one or more of the 25 Bark scales (or CSBs) are grouped into BDBs.
[0038] While exemplary embodiments are described above, these embodiments are not intended to describe all possible forms of the invention. Rather, the words used in the specification are words of description rather than limitation, and it is understood that various changes can be made without departing from the spirit and scope of the invention. Furthermore, the features implementing various embodiments can be combined to form further embodiments of the invention.
Claims
1. A system for providing three-dimensional (3D) immersive sound, said system comprising: a loudspeaker system including N loudspeakers for transmitting audio output signals in a listening environment; at least one controller, wherein the at least one controller: receiving an input audio signal; storing a plurality of Brauart direction determination bands associated with the input audio signal, each Brauart direction determination band being defined by a narrowband frequency interval regardless of the location of a sound source; storing at least a psychoacoustic measure including at least one subband in each Blauert direction determining band; determining the energy of each subband within the Blauert direction determination band; generating a loudspeaker drive signal based on the energy of at least the subband to drive the loudspeaker and transmit the audio output signal; It is programmed to The at least one controller a first filter bank that maps frequencies of the input audio signal to a set of critical subbands that includes the at least one subband; a Brauart crossover network that groups the plurality of Brauart direction determination bands to the N loudspeakers according to the critical subband mapping; a psychoacoustic modeling block that calculates the energy as the magnitude of the square of a complex number associated with a critical subband of the set of critical subbands; Including, the system.
2. The system of claim 1 , wherein the at least one controller is further programmed to determine a difference between the energy of the subband and a masked hearing threshold.
3. The system of claim 2 , wherein the masked hearing threshold corresponds to an audible signal audible by a listener.
4. The system of claim 2 , wherein the at least one controller is further programmed to compare the difference to one or more threshold values.
5. 5. The system of claim 4, wherein the at least one controller is further programmed to apply a gain to the loudspeaker drive signal based on a comparison of the difference to the one or more threshold values.
6. The system of claim 5 , wherein the gain one of: increases directionality of the audio output signal; or minimizes distortion of the audio output signal.
7. The system of claim 1 , wherein the at least one psychoacoustic measure is at least one Bark measure.
8. 1. A computer program for providing three-dimensional (3D) immersive sound, the computer program causing at least one controller to perform steps, the steps comprising: receiving an input audio signal; storing a plurality of Blauert direction determination bands, each Blauert direction determination band defined by a narrowband frequency interval regardless of the location of a sound source; storing at least a psychoacoustic measure including at least one subband in each Blauert direction determining band; mapping frequencies of the input audio signal to a set of critical subbands including the at least one subband; grouping the plurality of Blauert direction determination bands into N loudspeakers according to the critical subband mapping; determining an energy of each subband of each Blauert direction decision band, said determining comprising calculating a magnitude of a square of a complex number associated with a critical subband of the set of critical subbands; generating a loudspeaker drive signal based on the energy of at least each subband to drive the loudspeaker and deliver an audio output signal in a listening environment; a computer program comprising:
9. The computer program of claim 8, wherein the procedure further comprises determining the difference between the energy of the subband and a masked hearing threshold.
10. 10. The computer program of claim 9, wherein the masked hearing threshold corresponds to an audible signal audible by a listener.
11. The computer program of claim 9, wherein the procedure further comprises comparing the difference with one or more thresholds.
12. The computer program of claim 11, wherein the procedure further comprises applying a gain to the loudspeaker drive signal based on a comparison of the difference with the one or more thresholds.
13. The computer program product of claim 12 , wherein the gain one of: increases directionality of the audio output signal; or minimizes distortion of the audio output signal.
14. The computer program of claim 8 , wherein the at least one psychoacoustic measure is at least one Bark scale.
15. 1. A method for providing three-dimensional (3D) immersive sound, said method comprising: At least one controller receiving an input audio signal; the at least one controller storing a plurality of Brauart direction determination bands, each Brauart direction determination band defined by a narrowband frequency interval regardless of the location of a sound source; the at least one controller storing at least a psychoacoustic measure including at least one subband in each Blauert direction determination band; the at least one controller mapping frequencies of the input audio signal to a set of critical subbands including the at least one subband; the at least one controller grouping the plurality of Blauert direction determination bands into N loudspeakers according to the critical subband mapping; the at least one controller determining an energy of each subband of each Blauert direction decision band, said determining including calculating a magnitude of a square of a complex number associated with a critical subband of the set of critical subbands; the at least one controller generating a loudspeaker drive signal based on the energy in at least each subband to drive the loudspeaker and deliver an audio output signal in a listening environment; A method comprising:
16. 16. The method of claim 15, further comprising instructions for determining a difference between the energy of the subband and a masked hearing threshold.
17. The method of claim 16 , further comprising instructions for comparing the difference to one or more threshold values.
18. The method of claim 17, further comprising instructions for applying a gain to the loudspeaker drive signal based on a comparison of the difference with the one or more threshold values.
Citation Information
Patent Citations
Stereoscopic acoustic processing unit using linear prediction coefficient
JP1997247799A