Method for outputting sound and speaker

The method generates multiple audio sub-signals for speaker transducers to simulate a 3D sound field, addressing the limitations of conventional systems by adapting to dynamic acoustic spaces and enhancing sound reproduction.

JP7701753B2Active Publication Date: 2025-07-02CLANG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023522355
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-07
Filing Date
2021-09-24
Publication Date
2025-07-02
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Conventional speaker systems fail to account for the dynamic changes in acoustic spaces due to movement of objects and people, leading to suboptimal sound reproduction in various environments and inability to provide rich 3D audio cues.

Method used

A method involving the generation of multiple audio sub-signals representing different frequency intervals, each output by a speaker transducer, which are directed in various directions within a room to simulate a 3D sound field, using a combination of band-pass and high/low-pass filters to create spatial audio cues.

Benefits of technology

Enables accurate sound localization and immersive 3D audio experience by adapting to the listening environment, providing rich sound reproduction regardless of listener position and environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701753000001
    Figure 0007701753000001
  • Figure 0007701753000002
    Figure 0007701753000002
  • Figure 0007701753000003
    Figure 0007701753000003
Patent Text Reader

Abstract

A method for converting an audio signal into signals for a plurality of speaker transducers, wherein the audio signal is divided into audio sub-signals, each representing a particular frequency interval, and the signal for each speaker transducer includes a portion of each audio sub-signal that varies over time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for outputting sound, and more particularly to a method for imparting spatial information to a sound signal.

[0002] Known speaker systems are in a stereo setup, a surround setup, or an omnidirectional setup, and even if the speakers are composed of speaker transducers for different frequency bands, the same speaker transducer receives at least substantially all of the electrical audio signals within its band and always outputs at least substantially all of the sound in that band. In this sense, stationary speakers output "stationary" audio signals.

[0003] An omnidirectional speaker system reflects sound radially over 360 degrees from a central point, and the dispersion of the sound is substantially within a vertical plane. There are different strategies for dispersing monaural or stereo sound, including an omnidirectional system with the driver facing straight up or diagonally, and a system with the driver radiating upward onto a curved or conical reflector. Although called omnidirectional, none of them are true spherical speaker systems, and the aim is to radiate the desired waveform in a fixed or stationary state.

[0004] Conventional surround systems aim to enrich the fidelity and depth of the reproduced sound by using a plurality of speaker oscillators arranged in front of, to the sides of, and behind the listener. There are various types and numbers of speaker transducers in surround sound systems, but they all aim to emit the desired waveform in a fixed or stationary manner. This may be regardless of the listening environment of a huge number of different acoustic spaces installed, or based on an automated or user-defined process of adapting the sound to a specific listening environment as a customizable sound field. What these systems have in common is that they aim to ignore, deny, or interrupt the influence of the listening environment on reproduction, and these fixed, customizable, or user-definable sound fields, once established, remain stable.

[0005] As a result, these conventional systems operate in an "optimal" playback in a certain installation layout and an "ideal" listening position in a certain listening environment. As a result, there is a significant difference between relatively poor music playback by speakers and the complex and rich sound diffusion by acoustic performances, and this difference has plagued the audio system industry from the beginning. Also, such systems cannot provide richness to other sound fields that are not acoustically produced, such as studio recordings and digitally produced music content. Furthermore, the acoustic space is not completely constant due to the minute movements of people, objects, and other elements within the space, and it gives the sound subtle changes that are important for the overall perceived quality of the sound. This audio system can also take this fact into account in the process of providing or procuring additional 3D audio cues to the input audio signal, whereby the listener can listen to sound playback in a three-dimensional manner as if the listener were in the same space as the sound source. This is in contrast to the two-dimensional method, where the listener will hear the sound as if it were coming from outside the listening space, unless in a highly defined listening position and conditions.

Summary of the Invention

[0006] A first aspect of the present invention relates to a method of outputting audio based on an audio signal, the method comprising: receiving an audio signal; generating a plurality of audio sub-signals from the audio signal, each audio sub-signal representing an audio signal within a frequency interval of 100 - 8000 Hz, and the frequency interval of one audio sub-signal not being completely included in the frequency interval of another audio sub-signal; providing a speaker comprising a plurality of sound output drivers or loudspeaker transducers each capable of outputting sound at least at intervals of 100 - 8000 Hz, the speaker transducer being disposed within a room or venue. generating an electrical sub-signal for each loudspeaker transducer, each electrical sub-signal constituting a predetermined portion of each audio sub-signal; supplying the electrical sub-signal to the speaker transducer; comprising; generating the electrical sub-signal includes varying a predetermined portion of the audio sub-signal of each electrical sub-signal over time.

[0007] In this specification, an audio signal can be received in any format such as analog or digital. The signal may contain any number of channels, such as a monaural signal, a stereo audio signal, a surround sound signal, etc. Audio signals are often encoded by codecs such as FLAC, ALAC, APE, OFR, TTA, WV, MPEG. Often, an audio signal consists of all or most of the frequencies in the audible frequency range of 20 Hz to 20 kHz, even if the audio signal is more suitable for a narrower frequency interval such as 40 Hz to 15 kHz.

[0008] An audio signal usually corresponds to a desired physical or audible output, where correspondence means that the audio signal has the same frequency components as the sound, often the same relative signal strength, at least within the desired frequency band. Such components and relative signal strengths often change over time, but the correspondence preferably does not change.

[0009] An audio signal may be carried wirelessly or via a wire such as a cable (optical or electrical). An audio signal may be received from a streaming or live session, or from any type of storage.

[0010] It is desired to output an audio signal or at least a sound signal corresponding to its frequency interval. The present invention focuses on the sound in the frequency band where the human ear can determine the direction from which the sound arrives and the interaction of the sound within this frequency interval in a room or a venue. This frequency interval can be regarded as a frequency interval of 100 to 8000 Hz, but if desired, it can be selected, for example, between 300 to 7 kHz, 300 to 6 kHz, 400 to 4 kHz, or 200 to 6 kHz.

[0011] The auditory system uses several cues for sound source localization, such as the interaural time - level difference (or intensity - loudness difference), spectral information, timing analysis, correlation analysis, pattern matching, etc. The interaural level difference occurs in the range of 1500 Hz to 8000 Hz, and the level difference is highly frequency - dependent and increases as the frequency increases. The interaural time difference is dominant in the range of 800 - 1500 Hz, and the interaural phase difference is in the range of 80 - 800 Hz.

[0012] At frequencies below 400 Hz, the head size (ear distance of 21.5 cm, corresponding to an interaural time difference of 625 μs) is smaller than 1 / 4 of the sound wave length, so the confusion of the interaural phase difference begins to be a problem. Below 200 Hz, the interaural level difference becomes very small, and it becomes almost impossible to accurately evaluate the input direction only with the ILD. Below 80 Hz, all of the phase difference, ILD, and ITD become small, and it becomes impossible to determine the direction of the sound.

[0013] Similarly, considering the head size, at frequencies above 1600 Hz, the head size becomes larger than the sound wave length, so the phase information becomes ambiguous. However, the ILD becomes larger, and at higher frequencies, the group delay becomes prominent. That is, if there is an onset or transient of the sound, the interaural delay of this onset can be used to determine the input direction of the corresponding sound source. This mechanism becomes particularly important in a reverberant environment.

[0014] According to the present invention, a number of audio sub-signals are generated from an audio signal, each audio sub-signal representing the audio signal within a frequency interval of 100 to 8000 Hz, and the frequency interval of a certain sub-signal is not completely included in the frequency interval of another sub-signal. Therefore, the sub-signals represent the audio signal within the frequency interval. The sub-signals may be desired to constitute relevant portions of the audio signal. The sub-signals may be generated by applying a band-pass filter and / or one or more high-pass and / or low-pass filters to the audio signal in order to select a desired frequency interval. The sub-audio signals may be the same as the audio signal within the frequency interval, but the filters are often not ideal at their ends (extreme frequencies), for example, the filters often lose quality so as to pass frequencies below the center frequency of the high-pass filter to some extent.

[0015] None of the audio sub-signals has a frequency interval that is completely included in the frequency interval of another audio sub-signal. Therefore, all of the audio sub-signals represent different frequency intervals of the audio signal. Therefore, for each frequency within the range of 100 to 8000 Hz, their representations in the audio sub-signals will not be the same. A frequency may fall within one or more frequency intervals of the audio sub-signal and not in others. Of course, the frequency intervals may overlap. The filtering efficiency (Q value) can be selected as desired. The filtering may be performed by discrete components, DSP, a processor, etc.

[0016] To output sound defined by sound or at least an audio sub-signal, a speaker is provided that includes a plurality of sound output speaker transducers capable of outputting sound in a desired frequency range of at least 100 to 8000 Hz. The speaker transducers may be the same, or may have the same characteristics such as the same impedance curve. Alternatively, the speaker transducers may be of different types. It is preferable that the same signal, such as an audio signal or an audio sub-signal, generates the same sound when output from each loudspeaker transducer. Nevertheless, speaker transducers having different types or different characteristics can be used. For example, when an electrical sub-signal for a speaker transducer is adapted to the speaker transducer and all speaker transducers output at least substantially the same sound, that is, each has the same relationship between the sound output, such as one or more frequencies, and the signal adapted and input to the speaker transducer to generate the sound, the speaker transducer is like this.

[0017] The loudspeaker transducer is arranged in a room or a venue and may be directed in at least three different directions. The room or venue can have one or more walls, a ceiling, and a floor. The room or venue preferably has one or more sound reflecting elements such as walls / ceilings / floors / pillars.

[0018] The combination of speaker transducers may be selected to represent a 180-degree sphere like half of a sphere away from a flat surface. Such a flat surface may be a keyboard surface, a laptop surface, or a screen surface.

[0019] The direction of the speaker transducer may be the main direction of the sound wave output by the speaker transducer. The speaker transducer can have an axis such as a symmetry axis around which the highest sound intensity is output or the sound intensity profile is more or less symmetric.

[0020] The speaker transducers are directed in at least three different directions. The directions may be different if there is an angle of at least 5°, for example at least 10°, for example at least 20° between them, such as when projected onto a vertical or horizontal plane or when transformed to intersect. The angle between two directions may be the smallest possible angle between the two directions. The two directions may extend along the same axis in opposite directions. Clearly, more than three different directions may be preferred, such as when four, five, six, seven, eight, ten or more speaker transducers are used.

[0021] A particularly interesting embodiment is one in which one loudspeaker transducer is provided on each face of a cube and is directed to output sound in a direction away from the cube. In this embodiment, six different directions are used. In another embodiment, the loudspeaker transducers are arranged on walls, ceilings, and floors and are directed in such a way as to send sound into the space between the loudspeaker transducers.

[0022] For each loudspeaker transducer, an electrical sub-signal is generated. In this way, each speaker transducer can operate independently of the other speaker transducers. Clearly, when a large number of speaker transducers are used, the plurality of speaker transducers may be driven or operated identically. Such identically driven speaker transducers can have the same or different directions.

[0023] In this context, the electrical sub-signal is the signal intended for the loudspeaker transducer. This signal may be supplied directly to the loudspeaker transducer or may be adapted to the loudspeaker transducer by means such as amplification and / or filtering. Further, the electrical sub-signal may be in any form, such as optical, wireless, or electrical wiring. The electrical sub-signal may optionally be encoded using any codec and may be digital or analog. The speaker transducer may include a step-down transformer, filter, amplifier, receiver, DAC, etc. to receive the electrical sub-signal and drive the speaker transducer.

[0024] Each electrical sub-signal can be adapted in any desired way before being supplied to the speaker transducer. In one embodiment, the electrical sub-signal is amplified before being supplied to the speaker transducer. In that or another embodiment, the electrical sub-signal may be adapted, such as by filtering or equalization, so that its frequency characteristics are adapted to those of the loudspeaker transducer. Different amplification and adaptation may be desired for different loudspeaker transducers.

[0025] Each electrical sub-signal constitutes or represents a predetermined portion of each audio sub-signal. This portion may be zero in some audio sub-signals. Next, each audio sub-signal may be multiplied by a weight or coefficient, and then all the resulting audio sub-signals may be summed to form an electrical sub-signal. Clearly, this process may be performed by a computer, processor, controller, DSP, FPGA, etc., and this computer then outputs the electrical sub-signal or each electrical signal for supply to the speaker transducer or for conversion / reception / adaptation / amplification before supply to the speaker transducer.

[0026] Of course, an electrical sub-signal and / or an audio sub-signal may be stored between its generation and its supply to the speaker transducer. Thus, in addition to, or instead of, the actual audio signal, a new audio format in which such signals are stored may be seen.

[0027] When an electrical sub-signal is supplied to the speaker transducer, sound is output.

[0028] Preferably, the sum of the audio sub-signals is at least substantially identical to the portion of the audio signal provided within the outer frequency spacing of the audio sub-signals. Thus, the audio sub-signals may be selected to represent that portion of the audio signal. The portions of the audio signal outside of this overall frequency spacing may be processed in different ways. In this context, the intensity of the sum of the audio sub-signals may be within 10%, for example within 5%, of the energy / loudness of the corresponding portion of the audio signal. Alternatively, the energy / loudness in each frequency interval of a predetermined width, such as 100 Hz, 50 Hz or 10 Hz, of the synthesized audio sub-signals may be within 10%, such as within 5%, of the energy / loudness in the same frequency interval of the audio signal.

[0029] Of course, as an overall desire, scaling or amplification may be allowed so as not to obscure the frequency components within that frequency spacing of the audio signal. Thus, for each set of one, two, three, a plurality or two frequencies within the frequency spacing, it may be desired that the intensity of the summed audio sub-bands at that frequency be within 10%, for example within 5%, of the intensity of the audio signal. Thus, it is desired that the relative frequency intensities be maintained.

[0030] Similarly, the sum of the electrical sub-signals is preferably at least substantially the same as the portion of the audio signal provided within the outer frequency spacing of the electrical sub-signals. Thus, the electrical sub-signals can represent that portion of the audio signal. The portion of the audio signal outside this overall frequency spacing may be handled by other transducers. In this context, the intensity of the sum of the electrical sub-signals may be within 10%, for example within 5%, of the energy / loudness of the corresponding portion of the audio signal. Alternatively, the energy / loudness in each frequency interval of a predetermined width, such as 100 Hz, 50 Hz or 10 Hz, of the sum of the electrical sub-signals may be within 10%, such as within 5%, of the energy / loudness in the same frequency interval of the audio signal.

[0031] Of course, as an overall desire, scaling or amplification may be allowed so as not to obscure the frequency components within that frequency spacing of the audio signal. Thus, for each set of one, two, three, a plurality, or two frequencies within the frequency spacing, it may be desired that the intensity of the summed electrical sub-bands at that frequency be within 10%, for example within 5%, of the intensity of the audio signal. In this way, it is desirable that the relative frequency intensity be maintained from the audio signal to the sound output.

[0032] Obviously, the electrical sound signal is desirably adjusted so that the sound outputs from all speaker transducers correlate so that the audio signal is correctly represented. Thus, the generation of the audio sub-signals, electrical sub-signals, and any adaptation / amplification preferably preserves the adjustment and phase of the signals.

[0033] According to the present invention, the generation of the electrical sub-signals consists of varying over time a predetermined portion of the audio sub-signal in each electrical sub-signal. Thus, the generation of each electrical sub-signal, returning to the above mathematical parlance, is performed such that the weight(s) applied to the audio sub-signal vary with time and the proportion of a given audio sub-signal varies with time in the electrical sub-signal.

[0034] The way in which a portion or ratio changes over time can be selected in many of the ways described below. As one way, an audio sub-signal can be considered as a virtual loudspeaker transducer that outputs a sound corresponding to that particular signal. One or more of the actual loudspeaker transducers output a portion of the sound of the virtual loudspeaker transducer depending on where the actual loudspeaker transducer is placed and how it is directed. This type of abstraction is also seen in a standard stereo setup, where the position of a virtual sound source such as a string section of a classical orchestra may be represented by a sound that appears to come from that virtual position even if it is placed at a location away from the actual loudspeaker transducers of the stereo setup.

[0035] In this way, the portion of the audio sub-signal contained in the electrical sub-signal can be determined by the correlation between the desired position and potential direction of the virtual loudspeaker transducer corresponding to the audio sub-signal, and the position and potential direction of the actual loudspeaker transducer. The closer the positions and, if relevant, the more aligned the directions, the greater the likelihood that a large portion of the audio sub-signal will be seen in the electrical sub-signal of that loudspeaker transducer.

[0036] This determination can be made, for example, by simulating the positions of the actual and virtual speaker transducers on a geometric shape such as a sphere, where the actual speaker transducer has a fixed position but the virtual speaker transducer is allowed to move on the shape. Then, the portion of the audio signal of the virtual loudspeaker transducer in the electrical signal of the actual loudspeaker transducer can be determined based on the distance between the virtual loudspeaker transducer and the virtual actual loudspeaker transducer.

[0037] In one embodiment, the step of receiving an audio signal includes receiving a stereo signal. In this situation, the step of generating an audio sub-signal may consist of generating a plurality of audio sub-signals for each channel of the stereo audio signal.

[0038] And a large number of audio sub-signals may be associated with the right channel, and a large number of audio sub-signals may be associated with the left channel. There exists a pair of one audio sub-signal of the left channel and one sub-signal of the right channel having at least substantially the same frequency interval, and it may be desirable that such a pair of virtual speaker transducers are not directed at least substantially in opposite directions or at least in the same direction. This can be obtained by appropriately selecting portions of the electrical sub-signals while knowing the position and potential direction of the loudspeaker transducers. Also, it may be desirable that each pair of audio sub-signals is more independent and does not have cooperation, or that the cooperation is related to avoiding a perfect match in the direction between the left and right channels of the same sub-band.

[0039] In one embodiment, the step of receiving an audio signal includes receiving a monaural signal and generating, from the audio signal, a second signal that is at least substantially phase-inverted with respect to the monaural signal. In this situation, the step of generating an audio sub-signal may consist of generating a plurality of audio sub-signals for each of the monaural audio signal and the second signal.

[0040] Then, these two signals are treated as the left signal and the right signal of the stereo signal, and a number of audio sub-bands can be associated with the monaural signal, and a number of audio sub-bands can be associated with other channels. There exists a pair of one audio sub-band of the monaural signal and one sub-band of the other signal having at least substantially the same frequency interval, and it may be desirable that the virtual speaker transducers of such a pair do not face at least substantially in opposite directions or at least in the same direction. This can be obtained by appropriately selecting portions in the electrical sub-signals while knowing the positions and potential directions of the loudspeaker transducers.

[0041] The sub-bands in the central band with spatial audio cues can be generated or defined by several means, and generally, the larger the number of sub-bands, the better the results. Also, setting the frequency boundaries logarithmically is also advantageous, and one sub-band division can be into 3 bands with boundaries (Hz) at 100, 300, 1,200, 4,000. In another division, here with 6 bands, it can have boundaries (Hz) at 100, 200, 400, 800, 1,600, 3,200, 6,400. Such a low number of sub-bands can be given to 1, 2, 3 or more virtual drivers, and the same sub-band is distributed to 1, 2, 3 or more simultaneous virtual drivers at different positions on the virtual sphere. Thereby, the number of virtual drivers greatly contributes to the smoothness of the resulting audio sphere, improving the results.

[0042] The sub-band division can also follow other concepts such that equal distances correspond to perceptually equal distances, such as the Bark scale, which is a psychoacoustic scale. When dividing into 18 sub-bands on the Bark scale, the sub-band boundaries (Hz) are set at 100, 200, 300, 400, 510, 630, 770, 920, 1080, 1270, 1480, 1720, 2000, 2320, 2700, 3150, 3700, 4400.

[0043] For multiple sub-bands, the division into 1 / 3 octaves is also successful, and the sub-band boundaries (Hz) are 111, 140, 180, 224, 281, 353, 449, 561, 707, 898, 1122, 1403, 1795, 2244, 2805, 3534, 4488, 5610, 7069.

[0044] Also, sub-bands can be formed by subtraction. In the 5-sub-band subtraction method, the sub-band boundaries (Hz) are given as 100, 200, 400, 800, 1,600, 3,200, and the sub-bands of each virtual driver consist of the combinations band1 + band3, band1 + band4, band2 + band4, band2 + band5, band3 + band5.

[0045] Furthermore, since a smooth rendering of the incident sound to the sound sphere can be provided, a dynamic boundary approach is also possible, which will be described in detail elsewhere in this book.

[0046] Examples of the methods for determining the above sub-band boundaries all provide slightly different results in that the timbre, or "flavor", of the sound sphere changes to some extent. However, all of these are acceptable and conceptually consistent methods for preparing for the addition or procurement of spatial audio cues in the audio sphere.

[0047] Once the sub-band boundaries are determined using any number of bands as described above, it is possible to calculate an estimated value of the energy, power, loudness, or intensity of the signal in each sub-band. This typically involves non-linear, time-averaging operations such as sum of squares or logarithmic operations, and further smoothing, resulting in sub-band quantities that can be compared to each other, or quantities of a target signal such as pink noise. This comparison makes it possible to adjust by multiplying the sub-band quantity by a certain gain factor. This gain can be determined by 1) a theoretical signal or noise model such as pink noise, 2) dynamically estimated by storing the highest gain measured in real-time operation within a predetermined level, 3) by machine learning of the gain observed in the past during training, etc. Another way to adjust the sub-band quantity is to dynamically change the boundary frequencies, as discussed in depth elsewhere in this document.

[0048] One embodiment further includes deriving, from the audio signal, its low-frequency portion having frequencies below a first threshold frequency such as 100 Hz, and including the low-frequency portion at least substantially equally in all electrical sub-signals, or proportionally in the sub-signals of the same virtual driver. In this aspect, the audio signal having a low frequency is output by all audio sub-signals and / or all electrical sub-signals. Alternatively, it may be desirable to provide this low-frequency signal only to some audio sub-signals and / or some electrical sub-signals.

[0049] As an alternative, it is conceivable to provide this low frequency by one or more separate speaker transducers rather than by the speaker transducer.

[0050] One embodiment further includes deriving a high-frequency portion of the audio signal having a frequency exceeding a second threshold frequency, such as 8000 Hz, and including this high-frequency portion at least substantially equally in all electrical sub-signals or proportionally to the sub-signals of the same virtual driver. In this way, an audio signal having a high frequency is output by all audio sub-signals and / or all electrical sub-signals. Alternatively, it may be desirable to provide this high-frequency signal only to some audio sub-signals and / or some electrical sub-signals.

[0051] As an alternative, it is conceivable to provide this high frequency by one or more separate speaker transducers rather than by a speaker transducer.

[0052] As described above, the selection of the portions of the audio sub-signals represented by each electrical sub-signal can be performed based on many considerations.

[0053] In certain situations, it may be desirable for the sound energy, loudness, or intensity in each audio sub-signal and / or electrical sub-signal to be the same or at least substantially the same. On the other hand, it may be desired that the overall sound output corresponds to the audio signal such that, for example, the correspondence seen between the intensities / loudnesses of different frequency sets is the same or at least substantially the same in the audio signal and the sound output. Thus, the energy or loudness in an audio sub-band can be increased by increasing its intensity / loudness at one, more, or all frequencies within that frequency interval, which may not be desired. Alternatively, the intensity / loudness within a frequency interval can be increased by widening the frequency interval. Such a dynamic boundary approach can also be used to determine the two outer frequency boundaries of the combined frequency band for the low-frequency and high-frequency components. This may be calculated before the individual frequency bands are calculated, and these outer frequency boundaries may be calculated such that the coherence of the combined signal radiated by the combined speaker transducer has a desired degree of correspondence or similarity to the input sound.

[0054] In this context, the energy, loudness, or intensity of a sound or signal can be determined in many ways. One way would be to calculate the spectral envelope by means of a Fourier transform that returns the magnitude of each frequency bin of the transform corresponding to the amplitude of a particular frequency band. Then, by integrating the resulting envelope as a weight in the frequency domain and segmenting the result into a number of equal sizes corresponding to the number of sub-bands, new frequency boundaries for the sub-bands are obtained because the boundaries coincide with the intersections on the frequency axis of each segment obtained from the integration.

[0055] As another method, it is conceivable to calculate the spectral envelope by filter bank analysis, where the filter bank divides the input sound into several separate frequency bands and returns the amplitude of each band. This can be achieved by a large number of band-pass filters, such as 512 or more, or less, and the resulting band centers and loudness are integrated in a similar way as in the previous example.

[0056] Another variation of the filter bank example is to use a non-uniform filter bank where the number of filter bands is the same as the number of sub-bands in a particular implementation. The gradient and center frequency of each filter in the filter bank can be used to calculate the width of the sub-bands, from which the frequency boundaries between sub-bands can be derived.

[0057] Furthermore, there is also a method of using a bank of octave band filters, performing static weighting, and then performing the above integration step.

[0058] As another alternative, it is also possible to use music similarity measurements developed in music information retrieval (MIR). This method deals with extracting and inferring meaningful computable features from an audio signal. Given a collection of such features and appropriate segmentation into frequency sub-bands, the category of the music being played by the system can be determined by a simple look-up process, and the frequency bands can be set dynamically accordingly.

[0059] Finally, statistical methods such as machine learning by features can be used to make predictions and decisions regarding the appropriate frequencies of sub-band boundaries for a given audio input, where the algorithm is pre-trained on a large collection of sample audio data.

[0060] Accordingly, the step of generating the audio sub-signals may include selecting one or more frequency intervals of the audio sub-signals such that the combined energy in each audio sub-signal is within 10% of a predetermined energy / loudness value. Accordingly, all audio sub-signals have an energy / loudness within 10% of this value. Of course, the predetermined energy / loudness value may be the average value of the energy / loudness values of the audio sub-signals. Alternatively, the energy / loudness may be determined, for example, for the audio signal itself or for its channel. This energy / loudness may be divided by the number of audio sub-signals desired for the audio signal or channel. For example, if the energy / loudness in an audio signal in the interval 100 - 8000 Hz is determined and three audio sub-signals are desired, it may be divided by 3. And it is desirable that the energy / loudness of each audio signal be between 90% and 110% of this determined energy / loudness. And the frequency intervals can be adapted to achieve this energy / loudness. It is recognized that the frequency intervals may overlap.

[0061] It is reiterated that the above consideration of energy / loudness may relate to the audio sub-signals and / or electrical sub-signals.

[0062] In a particularly interesting embodiment, the portion of the audio sub-signal represented by one - or each - electrical sub-signal rather varies significantly. Accordingly, it may be desirable that the step of generating the electrical sub-signals includes generating the electrical sub-signals for one or more electrical sub-signals such that the portion of the audio sub-band represented by the electrical sub-band increases or decreases by at least 5% per second. Accordingly, that portion, which may be the ratio of energy / loudness / intensity of the audio sub-band, varies by more than 5% per second. Accordingly, if the ratio is 50% at t = 0, at t = 1 (second), the ratio is 47.5% or less or 52.5% or more.

[0063] In particular, when a speaker transducer is provided on the outer surface of an enclosure such as a speaker cabinet of any desired size and shape, the audio sub-signal can be regarded as individual virtual speaker transducers moving around within the cabinet, on the surface of the cabinet, or on a predetermined geometric shape. Their positions, and optionally their directions if not assumed to be in a predetermined direction, are correlated with the position of the actual loudspeaker transducer and are used to calculate the portions or weights. By simulating the rotation or movement of the individual virtual speaker transducers within or on the shape, a temporal change in the portions can be obtained.

[0064] Obviously, the sound output by the virtual loudspeaker transducer is the sound output by the actual loudspeaker transducer that receives a part of the audio sub-signal forming the virtual loudspeaker transducer. The whole sound output from the virtual loudspeaker transducer is determined by the portions supplied to each loudspeaker transducer, the position of the loudspeaker transducer, and optionally its direction. Repositioning or rotating the virtual loudspeaker transducer is done by changing the corresponding sound intensity / loudness in the individual loudspeaker transducers, i.e., by changing the portion of the audio sub-signal or the electrical sub-signal in the loudspeaker transducer.

[0065] A second aspect of the present invention relates to a system for outputting sound based on an audio signal, the system comprising an input for receiving the audio signal, a speaker including a plurality of sound output speaker transducers each capable of outputting sound at intervals of at least 100 - 8000 Hz, the speaker transducers being arranged within a room or a venue, A controller configured to generate a plurality of audio sub-signals from an audio signal and to generate an electrical sub-signal for each loudspeaker transducer, each audio sub-signal representing an audio signal within a frequency interval of 100 - 8000 Hz, the frequency interval of one audio sub-signal not being completely included in the frequency interval of another audio sub-signal, and each electrical sub-signal including a predetermined portion of each audio sub-signal, means for supplying the electrical sub-signal to a loudspeaker transducer and comprising the controller is configured to generate each of the electrical sub-signals such that a predetermined portion of the audio sub-signal of each electrical sub-signal changes over time.

[0066] As used herein, a system may be a combination of discrete elements or a single unitary element. The input, controller, and speaker may be a single element configured to receive an audio signal and output sound.

[0067] Alternatively, the controller may be separate or separable from the speaker and configured to remotely generate an electrical sub-signal or an audio signal from the speaker and supply it to the speaker.

[0068] Obviously, the controller may be one or more elements configured to communicate. Thus, the audio sub-signal may be generated by one controller and the electrical sub-signal may be generated by another controller. As will be described later, a new codec or encapsulation may be generated, whereby the audio sub-signal or electrical sub-signal is transferred to the controller or speaker in a controlled and standardized manner and then interpreted to output sound.

[0069] As described above, the audio signal may be in any format, such as either a known codec or encoding format. The audio signal may be received from a live performance, streaming, or storage.

[0070] The input may be configured to receive signals from a wireless source, an electrical cable, an optical fiber, storage, etc. The input may be configured to perform any desired or necessary signal processing, conversion, error correction, etc. to reach the audio signal. Thus, the input may be an input of another chip such as an antenna, a connector, a controller, or a MAC.

[0071] The speaker is configured to receive a signal and output sound. In this context, the speaker includes a plurality of speaker transducers configured to output sound. The loudspeaker transducer directs sound in at least three different directions as described above.

[0072] If a plurality of loudspeaker transducers are required to cover all of the frequency intervals covered, for example, by the frequency intervals of the audio sub-signals, the plurality of loudspeaker transducers may be directed in the same direction. If this frequency interval is wide and the speaker transducers have a narrower operating frequency interval, a large number of different speaker transducers may be required for each direction.

[0073] Also, if the directivity of the speaker transducer is too narrow, it may be desirable to provide a plurality of such speaker transducers having only slightly diverging directions to cover a specific angular interval with the audio sub-signal. As described above, more directions may be used.

[0074] The electrical sub-signals are to be supplied to the loudspeaker transducer. A controller or a part thereof that generates the electrical sub-signals may be provided in the speaker so that they do not need to be carried to the speaker. Alternatively, an input section for receiving these signals may be provided in the speaker. It is clear that this input must be configured to receive such signals and, if necessary, process the received signals to obtain the signals for each loudspeaker transducer. This processing may derive the electrical sub-signals from a general-purpose signal or a composite signal received by the speaker input.

[0075] The frequency interval is at least 100 - 8000 Hz, but may be narrower.

[0076] The controller is configured to generate a number of audio sub-signals from the audio signal. This process is further described above.

[0077] Note that the number of audio sub-signals does not need to correspond to the number of electrical sub-signals.

[0078] As described above, the same or a different controller may generate the electrical sub-signals from the audio signal in such a way that the portion of the audio sub-signal in each electrical sub-signal changes over time.

[0079] In one embodiment, the input is configured to receive a stereo signal. And the controller may be configured to generate a plurality of audio sub-signals for each channel of the stereo audio signal. And the audio sub-signals corresponding to the same frequency interval are supplied to a predetermined speaker transducer, and may be supplied over time so that two signals are not supplied to the same speaker transducer at a portion that is too high (included in the same electrical sub-signal).

[0080] In another embodiment, the input is configured to receive a monaural signal. And the controller can be configured to generate, from the audio signal, a second signal that is at least substantially phase-inverted with respect to the monaural signal, and to generate a plurality of audio sub-signals for each of the monaural audio signal and the second signal. And the audio sub-signals corresponding to the same frequency interval are supplied to a predetermined speaker transducer, and can be supplied over time so that the two signals are not supplied to the same speaker transducer (included in the same electrical sub-signal) in the part where they are too high.

[0081] In one embodiment, the controller is further configured to derive from the audio signal its low-frequency portion having a frequency less than a first threshold frequency that can be 100 Hz, 200 Hz, 300 Hz, 400 Hz or any frequency therebetween, and to include the low-frequency portion in all electrical sub-signals at least substantially uniformly. Alternatively, the speaker can be composed of a separate speaker transducer to which this low-frequency signal is supplied.

[0082] In one embodiment, the controller is further configured to derive from the audio signal its high-frequency portion having a frequency exceeding a second threshold frequency that is 4000 Hz, 5000 Hz, 6000 Hz, 7000 Hz or 8000 Hz or any frequency therebetween, and to include this high-frequency portion in all electrical sub-signals at least substantially equally. Alternatively, the speaker can also be composed of a separate speaker transducer to which this high-frequency signal is supplied.

[0083] In one embodiment, the controller is further configured to select one or more frequency intervals of the audio sub-signals such that the combined energy, e.g., combined loudness, in each audio sub-signal is within 10% of a predetermined energy / loudness value. As described above, it may be preferable for the energy, loudness, or intensity in each audio sub-signal to be the same. To achieve this, the frequency intervals of each audio sub-signal can be adapted. The predetermined energy value may be, for example, the total energy or average energy or loudness value of all audio sub-signals in a channel, or the ratio of the energy / loudness of the audio signal within the overall frequency interval of the audio sub-signals.

[0084] In one embodiment, the controller is further configured to generate an electrical sub-signal for one or more electrical sub-signals such that a portion of the audio sub-band represented by the electrical sub-band increases or decreases by at least 5% per second. In this aspect, the portion of the audio sub-signal in the electrical sub-signal varies quite a lot.

Brief Description of the Drawings

[0085] Unless otherwise specified, the accompanying drawings illustrate the aspects of the technological innovation described in this specification. Referring to the drawings, in several figures and throughout this specification, like numerals refer to like parts, and several embodiments of the presently disclosed principles are illustrated by way of example and not limitation.

[0086]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7a

Figure 7b

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0087] DETAILED DESCRIPTION In the following, various innovative principles for a system for providing a sound sphere with a smoothly varying or constant three-dimensional air transition are explained. For example, certain aspects of the disclosed principles relate to audio devices configured to project a desired sound sphere or an approximation thereof throughout a listening environment.

[0088] Embodiments of such systems described in the context of method acts are merely selected as convenient examples of the contemplated systems, which are only specific examples of the contemplated systems. One or more of the disclosed principles can be incorporated into various other audio systems to achieve any of the corresponding various system characteristics.

[0089] Accordingly, systems having attributes different from the specific examples discussed herein can embody one or more of the presently disclosed innovative principles and can be used for applications not described in detail herein. Accordingly, such alternative embodiments are also within the scope of this disclosure.

[0090] In some embodiments, the innovations disclosed herein generally relate to systems and related techniques for providing a three-dimensional sound sphere having a plurality of beams that are combined to provide smoothly varying sound localization information. For example, some of the disclosed audio systems can project subsections in the frequency band of the sound with a subtly changing or constant phase relationship and independent amplitudes onto loudspeaker transducers. Thereby, the audio system can render added or procured spatial information to any input audio across the listening environment.

[0091] As an example, an audio device can have an array of loudspeaker transducers, each constituting an independent full-range transducer. The audio device includes a processor and a memory that, when executed by the processor, contains instructions to render a three-dimensional waveform as a 360-degree sphere with weighted combinations such as individual virtual shape components that are slowly moved along the loudspeaker transducers by a panning process of the audio signal, adjustment pairs of shape components, etc. for the audio device. For each speaker transducer, the audio device can filter the received audio signal according to a specified procedure. When performing a dynamic sound sphere, the audio device retains the original sound across the combined sphere components when they are summed in the acoustic space. Thus, for the listener, the resulting sound retains the frequency envelope of the original sound but has the addition or procurement of dynamic or constant three-dimensional audio spatialization.

[0092] The present disclosure can combine its three-dimensional audio rendering with the total signal above and below two specified thresholds, and the audio signals outside the thresholds do not hold information regarding sound localization that can be recognized by a cognitive listening device. These two ranges are separately summed into two monophonic audio signals and can be transmitted simultaneously to all speaker transducers. Thereby, the audio device can provide complete three-dimensional spatialization recognizable by the cognitive listening device along with independent control of all speaker transducers in the low and high frequencies.

[0093] The present disclosure can manage one monaural signal input on one audio device with an equal number of independent spherical components as the number of speaker transducers of the device, or a different number of virtual spherical components from the number of speaker transducers of the device. Each spherical component can be a subset of the frequency range, and all components can be evenly distributed along the range as a balanced sum of the components. These components can be panned independently at all speaker transducers on the plane of the geometric solid, or can be changed as a polarity-inverted pair at opposite points of the geometric solid, or in other ways, and they can be placed at any point between adjacent planes. Such a system is used in a stereo configuration paired with two devices, provides separate three-dimensional spatialization for each of the monophonic audio channels, renders the left and right channels separately to the two audio devices, resulting in a three-dimensional stereophonic audio rendering system. Also, it is possible to pan the stereo pair individually, and no correlation can be observed at opposite points.

[0094] The present disclosure can manage one stereo signal on one audio system with an equal number of independent iterations equal to half the number of unit speaker transducers. Each pair is a subset of the frequency range of the stereo signal and can be placed at opposing points on a geometric solid or at any point between adjacent planes of the solid. The stereo pairs are evenly panned and can perform a satisfactory rendering of the input stereo signal with one audio device. As a result, a point sound source, three-dimensional stereophonic audio rendering system is realized.

[0095] Instructions stored in the processor memory can create an adaptable division of frequency bands that can observe equal loudness between bands if so desired. This avoids sudden direction changes due to energy / loudness changes in very local frequency ranges.

[0096] I. Overview Referring now to FIGS. 1 and 2, an audio device (or speaker) 10 can be placed in a room 20. A three-dimensional sound sphere 30 is rendered by the audio device 10, and the optimal listening area for the listener coincides with the sphere 30.

[0097] FIGS. 3 and 4 show other exemplary representations of the positioning of the device 10. The audio device 10 can correspond to the position of one or more reflective boundaries, such as walls 22a, 22b, relative to the device 10, and the likely positions 26a, 26a of the listener that coincide with the sound spheres 30a, 30b. The rendered three-dimensional sound spheres 30a, 30b are enhanced as the waveform is reflected from the wall.

[0098] As will be described in more detail below, a three-dimensional sound sphere can be composed of a combination of spherical components. The three-dimensional sound sphere depends on changes in amplitude, phase, and time along different audio frequencies, or frequency bands. Methodologies can be devised to manage such dependencies, and the disclosed audio devices can apply these methods to acoustic or digital signals containing audio content to render them as three-dimensional sound spheres.

[0099] In Section II, the principles related to such audio devices will be explained with reference to the device depicted in FIG. 1. In Section III, the principles related to the desired three-dimensional sound sphere will be explained, and in Section IV, the principles related to decomposing audio content into a combination of virtual and real spherical components and recombining them in the acoustic space will be explained. In Section V, the principles of directivity related to the three-dimensional nature of the audio device and its variation with frequency will be disclosed. Section VI describes the principles related to an audio processor suitable for rendering an approximation of the desired three-dimensional sound field from an input audio signal on input 51 containing audio content. In Section VII, the principles related to a computing environment suitable for implementing the disclosed processing methods will be explained. This will include, for example, an example of a machine-readable medium that, when executed, causes a processor 50 of the computing environment to execute one or more of the disclosed methods. Such instructions can be embedded in software, firmware, or hardware. Further, the disclosed methods and techniques can also be implemented in software, firmware, or hardware in various forms of signal processors.

[0100] II. Audio Device FIG. 1 shows an audio device 10 including a loudspeaker cabinet 12 incorporating a loudspeaker array including a plurality of individual loudspeaker transducers or loudspeaker transducers S1, S2, ..., S6.

[0101] In general, the loudspeaker array can have any number of individual loudspeaker transducers, even though the illustrated array has six loudspeaker transducers. The number of speaker transducers shown in FIG. 1 is chosen for the sake of illustration. Other arrays can have more or less than six transducers, and there can be more or less than three axes of transducer pairs, where an axis can have only one transducer. For example, one embodiment of an array for an audio device can have 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or more loudspeaker transducers.

[0102] In FIG. 1, the cabinet 12 is generally cubic in shape, defining a central axis z disposed at opposite corners 16 of the cubic cabinet.

[0103] Each loudspeaker transducer S1, S2, ..., S6 within the loudspeaker array of the figure is evenly distributed on the plane of the cube at a uniform radial distance, pole, and azimuth angle at a constant or substantially constant position relative to the center of the axis. In FIG. 1, the loudspeaker transducers are arranged at spherical intervals of approximately 90 degrees from each other.

[0104] Other arrangements of the speaker transducers are possible. For example, the speaker transducers within the array can be evenly or unevenly distributed within the loudspeaker cabinet 12. Also, the speaker transducers S1, S2, ..., S6 can be arranged at various selected spherical positions measured from the axis center rather than at a constant distance as shown in FIG. 1. For example, each speaker transducer can be distributed from more than two axis points.

[0105] Each of the transducers S1, S2, ..., S6 may be an electrodynamic or other type of loudspeaker transducer specially designed for the output of sound in a specific frequency band, such as a woofer, tweeter, midrange, full range, etc. The audio device 10 can be combined with a seventh speaker transducer S0, which can complement the output from the array. For example, the auxiliary loudspeaker transducer S0 can be configured to radiate a selected frequency, such as a low-end frequency as a subwoofer. The supplementary loudspeaker transducer S0 can be built into the audio device 10 or housed in a separate cabinet. Also, the speaker transducer of S0 can be used for high-frequency output.

[0106] The loudspeaker cabinet 12 is shown as a cube, but other embodiments of the loudspeaker cabinet 12 have different shapes. For example, some loudspeaker cabinets can be arranged in, for example, a general prism structure, a tetrahedral structure, a spherical structure, an elliptical structure, a toroidal structure, or any other desired three-dimensional shape.

[0107] III. 3D Sound Sphere Referring again to FIG. 2, the audio device 10 can be placed in the center of the room. In such a situation, as described above, a 3D sound sphere is evenly distributed around the audio device 10.

[0108] By projecting acoustic energy onto a three-dimensional sphere, the user's listening experience can be improved compared to a two-dimensional audio system. This is because, in contrast to the prior art in one-dimensional and two-dimensional sound fields, the three-dimensional listening cues provided by the disclosure are spatial and thus immersive, similar to sound cues in the physical world.

[0109] Furthermore, since the added spatial audio queue does not operate based on an ideal listening position as long as the entire listening field, i.e., the sphere, contains an even or nearly even balance of the prominent features of the original sound input, the disclosed listening space provides an infinite number of listening positions around device 10.

[0110] Figure 3 shows audio device 10 in a different position than shown in Figure 2. In Figure 2, the sound field 30 is circular and directs little or no acoustic energy towards wall 22. The three-dimensional sound sphere shown in Figure 3 is different from that shown in Figure 2, but the sound sphere shown in Figure 3 can fit well with the illustrated position of the loudspeaker compared to wall 22 and a listening position that could coincide with the sound sphere 30 shown in Figure 3, which is now partially folded, since the reflection of wall 22 is not incompatible with sound sphere 30. This is to avoid a constant forcing of a particular frequency, i.e., a frequency band, since the components of the sphere are constantly shifted along the loudspeaker transducers. Similarly, Figure 4 shows audio device 10 in yet another position in the room and a three-dimensional sound sphere 30 that coincides with the listening position, again folded correspondingly by the position of wall 22 and the room layout compared to the position of the audio device 10 shown in Figure 2. In this particular arrangement, the same situation regarding the projection of sound sphere 30 occurs by shifting the components of the sphere as in Figure 3, and no particular frequency, or frequency band, is constantly forced.

[0111] In some embodiments of the audio device, when the proximity of the audio device to the wall 22 is extreme or very prominent, the cubic sound field can be changed. For example, by expressing the three-dimensional sound sphere 30 using polar coordinates with the z-axis of the audio device 10 as the origin, the user can "draw" the directional scaling of the amplitude of the speaker transducer with respect to the z-axis of the audio device 10 on the touch screen, thereby modifying the sound sphere 30 from a sphere to an asymmetric triaxial ellipsoid.

[0112] In yet other embodiments, the user can select from a plurality of three-dimensional asymmetric triaxial ellipsoids stored by the audio device 10 or remotely. When stored remotely, the audio device 10 can load the selected triaxial asymmetric ellipsoid via a communication connection. In yet another embodiment, the user can "draw" the desired triaxial asymmetric ellipsoid contour or the boundary of the existing room on a smartphone or tablet as described above, and the audio device 10 can receive the desired asymmetric triaxial ellipsoid or the representation of the room boundary directly or indirectly from the user's device via a communication connection. As will be described later in connection with the computer environment, other forms of user input other than the touch screen can be used.

[0113] IV. Mode Decomposition and Reconstruction of Three-Dimensional Sound Sphere FIG. 5 shows the frequency range from 40 (located at 100 Hz) to 45 (located at 3 kHz) as a subset of the full frequency range of the listener's hearing for spatial sound source localization in three-dimensional hearing. Clues for sound source localization include the interaural time difference and level difference, spectral information, timing analysis, correlation analysis, pattern matching, etc. In the present disclosure, this knowledge of the auditory system is used to divide the frequency range from 40 to 45 into several bands (arrows), and by processing these bands, spatial information is added to or obtained from the input sound. The number of bands can be half the number of loudspeaker transducers, and the number of transducers can be increased or decreased.

[0114] As a means of an example rather than all possible embodiments, in FIG. 6, the high-pass filter 50, band-pass filters 51, 52, and 53, and the low-pass filter 54 separate the audio stream into five sub-streams or audio sub-signals. The high-pass filter removes signal components above 4 kHz, and the low-pass filter removes signal components below 100 Hz. The audio streams from filters 50 and 54 are outside the three-dimensional audible range and are sent equally to all loudspeaker transducers S1, S2,..., S6 according to different methods, or are sent to the loudspeaker transducer S0. Copies of the signals from each frequency band from filters 51, 52, and 53 can be changed by applying a degree of phase shift or by polarity inversion, and then, as the sum of the individual signals, a signal changed to a different point such as a point 180 degrees opposite to the original signal of the audio device 10 is sent to reach the signals of the loudspeaker transducers S1 - S6. The resulting audio output is a monophonic sound with an independent spatial cue added to three sets of connected spherical components, becoming a monophonic three-dimensional sound sphere. In a variation of this example, the audio streams from filters 51, 52, and 53 are sent individually to the loudspeaker transducers S1, S2,..., S6 and are moved in a randomly or semi-randomly adjusted manner. This also provides a spatial clue for a monophonic three-dimensional sound sphere, but is of a substantially different nature from the previous example.

[0115] Figure 7a represents the same scenario, but with a stereo signal input. As an example, but not in all possible embodiments, in Figure 7a, the high-pass filter 60, band-pass filters 61, 62, 63, and low-pass filter 64 separate the audio into five audio streams. The audio streams from filters 60 and 64 are outside the three-dimensional audible range and are evenly transmitted to all loudspeaker transducers S1, S2, ..., S6. This provides little or no spatial information, either as the sum monaural signal of the low-pass and high-pass audio, or as two separate audio streams of the left and / or right channels of the low-pass and high-pass audio, and is transmitted before emission. The audio streams from filters 61, 62, and 63 that are within the three-dimensional audible range are transmitted individually, but are currently transmitted in pairs to the speaker transducers [S1, S2], [S3, S4], [S5, S6], or any axial point between the transducers. The resulting audio output is stereo with the addition or procurement of a spatial cue to provide a three-dimensional sound field of point-source stereo.

[0116] Figure 7b shows a scenario where the stereo signal input is treated as separate monaural channels. As an example, not all possible embodiments, in Figure 7b, the high-pass filter 70, band-pass filters 71A, 71B, 72A, 72B, 73A, 73B, and low-pass filter 74 separate the audio into eight audio streams. The audio streams from filters 70 and 74 are outside the three-dimensional audible range and are evenly transmitted to all loudspeaker transducers S1, S2, ..., S6. This provides little or no spatial information as the combined monaural signal of the low-pass and high-pass audio, or is transmitted before emission as two separate audio streams for the left and / or right channels of the low-pass and high-pass audio. The audio streams from filters 71A, 71B, 72A, 72B, 73A, 73B that are within the three-dimensional audible range are individually transmitted to the speaker transducers [S1, S2, S3, S4, S5, S6], or any axial point between the transducers. The resulting audio output is a plurality of unidirectional sounds, adding or procuring a spatial cue to provide a plurality of unidirectional three-dimensional sound fields of point sources. Therefore, compared to Figure 7a, there is no need for a correlation relationship in the angle of the direction in which the corresponding audio sub-signals (related to the same sub-band) are output.

[0117] V. Consideration of Directivity Figure 8 shows the 10 directivity coefficient aspects of the sound device. The range of the directivity coefficient is 1 - ∞, indicating the ability of the loudspeaker transducer (or any other arbitrary sound generator) to confine the applied energy to a spherical section. The audio device has different degrees of directivity throughout the audible frequency range (e.g., from about 20 Hz to about 20 kHz). Generally, as the frequency approaches 20 Hz, the directivity coefficient decreases, and as the frequency increases, the directivity coefficient increases. Considering that the 10 directivity coefficients of the publicly available audio devices are evenly or nearly evenly distributed in a regular geometric solid, they are 1 or close to 1 across the entire frequency range. The directivity coefficients of the 10 individual loudspeaker transducers of the disclosed audio device are close to 2 at low frequencies and vary depending on the frequency range, but tend towards higher values at high frequencies. When the directivity coefficient is 8, each transducer has a spherical portion and, in combination with the 6 transducers on the aforementioned cube cabinet, is coupled to form a complete sphere for the audio device 10. The directed energy for a single loudspeaker transducer, as a range of selected angular positions at a fixed radius with the loudspeaker at the origin to determine a defined listening window, decreases the user's listening experience as the user's position relative to the loudspeaker changes. This disclosure with a much lower directivity coefficient has an infinite or far more desired listening positions than the prior art in the second-order sound field.

[0118] In order to realize a desired sound sphere or a smoothly varying spherical component (or pattern) across all frequencies, the above spherical component can undergo equalization, so that each spherical component provides a corresponding sound field with a desired frequency response throughout. In other words, a filter can be designed to give the desired frequency response to the entire spherical component. Then, by combining the equalized spherical components, a sound sphere can be rendered in which the spherical components transition smoothly across audible frequencies and / or selected frequency bands within the audible frequency range.

[0119] VI. Audio Processor A block diagram of an audio rendering processor for playing audio content (e.g., music works, movie soundtracks) on an audio device 10 is shown in FIG. 9. The audio rendering processor 50 may be a special-purpose processor such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a series of hardware logic structures (e.g., filters, arithmetic logic units, dedicated state machines). In some cases, the audio rendering processor can be implemented using a combination of machine-executable instructions that, when executed by the processor, cause the audio device to process one or more input channels as described. The rendering processor 50 receives input channels of the content of the sound program from the input audio source 51.

[0120] The input audio source 51 can provide a digital input or an analog input. The input audio source or input 51 can include a programmed processor executing a media player application program and can include a decoder that generates a digital audio input to the rendering processor. To do this, the decoder can decode the encoded audio signal, which may be encoded using a suitable audio codec such as Advanced Audio Codec (AAC), MPEG Audio Layer II, MPEG Audio LAYER III, and Free Lossless Audio Codec (FLAC). Alternatively, the input audio source can include a codec that converts an analog or optical audio signal to a digital format for the audio rendering processor 50, for example, from a line input. Alternatively, there may be multiple input audio channels, such as a two-channel input of the left and right channels of a stereo recording of a musical work, or there may be multiple input audio channels, such as the entire audio soundtrack of a movie or a movie in 5.1 surround format. Other examples of audio formats include 7.1 and 9.1 surround formats.

[0121] The array of loudspeaker transducers 58 can render a sphere of the desired sound (or an approximation thereof) based on a combination of spherical component segments 52a... 52N applied to the audio content by the audio rendering processor 50. The rendering processor 50 according to FIG. 9 can conceptually be divided into a spherical component domain and a loudspeaker transducer domain. In the component region, segment processing 53a... 53N for each of the constituent spherical components 52a... 52N can be applied to the audio content corresponding to the desired spherical component as described above. The equalizers 54a.... 54N provide equalization to each of the spherical components 52a... 52N to adjust for the variations in the directivity coefficients resulting from the particular audio device 10 as described above, and for any spherical adjustments towards the desired asymmetric ellipsoidal sphere profile.

[0122] In the loudspeaker transducer region, a spherical region matrix can be applied to various spherical region signals to provide the signals reproduced by each of the loudspeaker transducers in the array 58. Generally, the matrix is a matrix of size M×N, where N is the number of loudspeaker transducers and M=(2×N)+(2×O), where O represents the number of virtual spherical components. The equalizers 56a... 56N can provide equalization to each of the spherical components 57a... 57N to adjust for the variations in the directivity coefficients resulting from the particular audio device 10 as described above, and for any spherical adjustments towards the desired ellipsoidal sphere profile.

[0123] It should be understood that the audio rendering processor 50 can perform other signal processing operations to render the input audio signal for playback by the transducer array 58 in a desired manner. In another embodiment, the audio rendering processor can use an adaptive filter process to determine a constant or varying boundary frequency in order to determine how to modify the loudspeaker transducer signal. FIG. 10 is a block diagram of an audio device 10 for rendering a synthesized sound (e.g., a digital keyboard, a digital audio workstation (DAW)), or an audio rendering processor of an electric and / or acoustic musical instrument.

[0124] VII. Computing Environment FIG. 10 shows a generalized example of a suitable computing environment 100 that can include the operation of the controller 50, where, for example, methods, embodiments, techniques, and technologies related to procedurally generating a sound sphere are described. The computing environment 100 is not intended to suggest any limitation as to the scope or functionality of the technologies disclosed herein, since each technology can potentially be implemented in a variety of general-purpose or special-purpose computing environments. For example, each of the disclosed technologies can be implemented in other computer system configurations including, but not limited to, wearable and handheld devices, mobile communication devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, embedded platforms, network computers, minicomputers, mainframe computers, smartphones, tablet computers, data centers, etc. Each of the disclosed technologies may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications connection or network, or incorporated into digital or analog musical instruments. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.

[0125] Computing environment 100 includes at least one central processing unit 110 and a memory 120. In FIG. 10, this most basic configuration 130 is included within the dashed lines. The central processing unit 110 executes computer-executable instructions and can be either a physical processor or a virtual processor. In a multiprocessing system, multiple processing units can execute computer-executable instructions to enhance processing power, enabling multiple processors to operate simultaneously. The memory 120 can be volatile memory (e.g., registers, cache, RAM, etc.), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or a combination thereof. The memory 120 stores software 180a that can implement, for example, one or more of the innovative technologies described herein when executed by the processor.

[0126] The computing environment may have additional features. For example, computing environment 100 includes storage 140, one or more input devices 150, one or more output devices 160, and one or more communication connections 170. Interconnection mechanisms (not shown), such as buses, controllers, networks, etc., interconnect the components of computing environment 100. Typically, an operating system software (not shown) provides an operating environment for other software executed in computing environment 100 and coordinates the activities of the components of computing environment 100.

[0127] The storage 140 can be removable or non-removable and includes selected forms of machine-readable media such as magnetic disks, magnetic tapes or cassettes, non-volatile solid-state memories, CD-ROMs, CD-RWs, DVDs, magnetic tapes, optical data storage devices, and carrier waves, or other machine-readable media that can be used to store information and can be accessed within computing environment 100. The storage 140 stores instructions of software 180b that can implement the technologies described herein.

[0128] Storage 140 can also be distributed over the network so that software instructions are stored and executed in a distributed manner. In other embodiments, some of these operations may be performed by certain hardware components including hardwired logic. These operations can also be performed by any combination of programmed data processing components and fixed hardwired circuit components.

[0129] Input device 150 can be a keyboard, keypad, mouse, pen, touch screen, touch pad, trackball, or other touch input device, voice input device, scanning device, or other device that provides input to computing environment 100. In the case of audio, input device 150 can include a microphone or other transducer (e.g., a sound card or similar device that accepts voice input in analog or digital form), or a computer-readable media reader that provides audio samples to computing environment 100.

[0130] Output device 160 can be a display, printer, speaker transducer, DVD writer, or other device that provides output from computing environment 100.

[0131] Communication connection 170 enables communication to another computing entity via a communication medium (e.g., a connected network). The communication medium transmits information such as computer-executable instructions, compressed graphics information, processed signal information (including processed audio signals), or other data as a modulated signal.

[0132] Thus, the disclosed computing environment is suitable for performing the disclosed direction estimation and audio rendering processes as disclosed herein.

[0133] A machine-readable medium is any available medium that can be accessed within computing environment 100. In computing environment 100, by way of example and not limitation, machine-readable media include memory 120, storage 140, communication media (not shown), and any combination of the foregoing. Tangible machine-readable (or computer-readable) media exclude transitory signals.

[0134] As described above, some of the disclosed principles can be embodied in a tangible, non-transitory machine-readable medium (such as a microelectronic memory) that stores instructions for programming one or more data processing components (commonly referred to herein as “processors”) to perform the digital signal processing operations described above, including estimation, adaptation, calculation, computation, measurement, adjustment (by audio processor 50), sensing, measurement, filtering, addition, subtraction, inversion, comparison, and decision-making. In other embodiments, some of these operations (of the machine process) may be performed by specific electronic hardware components that include wired logic (such as a dedicated digital filter block). These operations can also be performed by any combination of programmed data processing components and fixed hard-wired circuit components.

[0135] Audio device 10 can include a loudspeaker cabinet 12 configured to generate sound. Audio device 10 can also include a processor and a non-transitory machine-readable medium (memory) storing instructions that, when executed by the processor, automatically perform a three-dimensional sphere construction process and processes to support it, as described herein.

[0136] The above examples generally relate to an apparatus, method, and related system for rendering audio, and more particularly, for providing a desired three-dimensional sphere pattern. Nevertheless, embodiments other than those described in detail above are contemplated in accordance with the principles disclosed herein, along with attendant modifications in the configuration of each apparatus described herein.

[0137] Directions and other relative references (e.g., up, down, above, below, left, right, behind, in front, etc.) can be used herein to facilitate the discussion of the drawings and principles, but are not intended to be limiting. For example, specific terms such as "up," "down," "above," "below," "horizontal," "vertical," "left," "right," etc. may be used. Such terms are used where applicable to clarify the description to some extent when dealing with relative relationships, particularly with respect to the specifically illustrated embodiments. However, such terms do not mean absolute relationships, positions, and / or directions. For example, with respect to an object, simply turning the object over causes the "upper" surface to become the "lower" surface. Nevertheless, it is the same surface, and the object remains the same. As used herein, "and / or" means "and," "or," as well as "and" and "or." Further, all patent documents and non-patent documents cited herein are hereby incorporated by reference in their entirety for all purposes.

[0138] The principles described above in connection with any particular example can be combined with the principles described in connection with another example described herein. Accordingly, this detailed description is not to be construed in a limiting sense, and after review of this disclosure, one of ordinary skill in the art will be able to evaluate a variety of signal processing and audio rendering techniques that can be devised using the various concepts described herein.

[0139] Furthermore, those skilled in the art will understand that the exemplary embodiments disclosed herein can be adapted to various configurations and / or uses without departing from the disclosed principles. Applying the principles disclosed herein, a wide variety of systems adapted to provide a desired three-dimensional spherical sound field can be provided. For example, the modules identified as part of a particular computing engine in the above description or drawings can be divided differently than described herein, distributed among one or more modules, or omitted entirely. Similarly, such modules can be implemented as part of a different computing engine without departing from some of the disclosed principles.

[0140] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the disclosed innovation. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the present disclosure. Accordingly, the invention as claimed is not intended to be limited to the embodiments shown in this specification, but should be accorded the full scope consistent with the claim language, for example, the use of the articles "a" or "an" to refer to an element in the singular is not intended to mean "only one" unless specifically stated otherwise, but rather "one or more". All structural and functional acts corresponding to the features and method acts of the various embodiments described throughout the disclosure, which are known or later become known to a person of ordinary skill in the art, are intended to be encompassed by the features described and claimed herein. Furthermore, what is disclosed herein is not intended to be provided to the public regardless of whether such disclosure is explicitly recited in the claims. A claim shall not be construed unless the claim recites the phrase "means for" or "step for" explicitly.

[0141] Accordingly, considering the many possible embodiments to which the disclosed principles can be applied, we reserve the right to claim any combination of features and techniques described herein, including, for example, all those within the scope of the technology, as would be understood by one of ordinary skill in the art.

Claims

1. A method for outputting sound based on an audio signal, comprising: receiving an audio signal; generating a plurality of audio sub-signals from the audio signal, each audio sub-signal representing an audio signal within a frequency interval of 100 - 8000 Hz, and the frequency interval of one audio sub-signal not being completely included in the frequency interval of another audio sub-signal; providing a speaker comprising a plurality of sound output loudspeaker transducers each capable of outputting sound at intervals of at least 100 - 8000 Hz, the loudspeaker transducers being arranged in a room or a venue; generating an electrical sub-signal for each loudspeaker transducer, each electrical sub-signal constituting a predetermined portion of each audio sub-signal; supplying the electrical sub-signal to the transducer of the speaker; and the step of generating the electrical sub-signal comprises: changing a predetermined portion of the audio sub-signal of each electrical sub-signal over time; providing the same or at least substantially the same sound energy, loudness, or intensity to the electrical sub-signal; and a method.

2. The step of receiving an audio signal includes receiving a stereo signal, and the step of generating audio sub-signals includes generating a plurality of audio sub-signals for each channel of the stereo audio signal. The method according to claim 1.

3. The step of receiving an audio signal includes receiving a monaural signal and generating from the audio signal a second signal that is at least substantially phase-inverted with respect to the monaural signal, and the step of generating audio sub-signals includes generating a plurality of audio sub-signals for each of the monaural audio signal and the second signal. The method according to claim 1.

4. further comprising deriving a low-frequency portion from the audio signal, the low-frequency portion having a frequency below a first threshold frequency, and all electrical sub-signals including at least substantially equally the low-frequency portion. The method according to any one of claims 1 to 3.

5. further comprising deriving a high-frequency portion from the audio signal, the high-frequency portion having a frequency exceeding a second threshold frequency. All electrical sub-signals include at least substantially equally said high-frequency portion, The method according to any one of claims 1 to 4.

6. The step of generating an audio sub-signal includes selecting a frequency interval of one or more audio sub-signals such that the energy / loudness of each audio sub-signal is within 10% of a predetermined energy / loudness value. The method according to any one of claims 1 to 5.

7. The step of generating an electrical sub-signal includes generating an electrical sub-signal such that for one or more electrical sub-signals, a part of the audio sub-band represented by the electrical sub-band increases or decreases by at least 5% per second. The method according to any one of claims 1 to 6.

8. In a system for outputting sound based on an audio signal, An input for receiving an audio signal, A speaker including a plurality of sound output loudspeaker transducers each capable of outputting sound at intervals of at least 100 - 8000 Hz, the loudspeaker transducers being arranged in a room or a venue, A controller configured to generate a plurality of audio sub-signals from the audio signal and to generate an electrical sub-signal for each loudspeaker transducer, each audio sub-signal representing an audio signal within a frequency interval of 100 - 8000 Hz, the frequency interval of one audio sub-signal not being completely included in the frequency interval of another audio sub-signal, and each electrical sub-signal including a predetermined portion of each audio sub-signal, Means for supplying the electrical sub-signal to the loudspeaker transducer Including, The controller is Each electrical sub-signal is configured to generate each electrical sub-signal such that a predetermined portion of the audio sub-signal changes over time so that the sound energy, loudness, or intensity of the electrical sub-signal is the same or at least substantially the same. Configured to generate each of the electrical sub-signals, System.

9. The input is configured to receive a stereo signal, and the controller is configured to generate a plurality of audio sub-signals for each channel of the stereo audio signal. The system according to claim 8.

10. The input is configured to receive a monaural signal, and the controller is configured to generate, from the audio signal, a second signal that is at least substantially phase-inverted with respect to the monaural signal, and to generate a plurality of audio sub-signals for each of the monaural audio signal and the second signal. The system according to claim 8.

11. The controller further derives a low-frequency portion having a frequency equal to or lower than a first threshold frequency from the audio signal, and includes the low-frequency portion in all the electrical sub-signals at least substantially equally. The system according to any one of claims 8 to 10, which is configured as described above.

12. The controller further derives a high-frequency portion having a frequency exceeding a second threshold frequency from the audio signal, and includes the high-frequency portion in all the electrical sub-signals at least substantially equally. The system according to any one of claims 8 to 11, which is configured as described above.

13. The controller is further configured to select a frequency interval of one or more audio sub-signals such that the energy / loudness in each audio sub-signal is within 10% of a predetermined energy / loudness value. The system according to any one of claims 8 to 12.

14. The controller is further configured to generate an electrical sub-signal for one or more electrical sub-signals such that a part of the audio sub-bands represented by the electrical sub-bands increases or decreases by at least 5% per second. The system according to any one of claims 8 to 13.

Citation Information

Patent Citations

  • Speaker apparatus

    JP2005198048A

  • Speaker array apparatus and signal processing method

    JP2008205822A

  • Omni-azimuth frequency-directional acoustic apparatus

    JP2009100354A

  • Method and apparatus for applying dynamic range compression to high-order ambisonics signals

    JP2017513367A

  • A rotationally symmetric speaker array

    US20170238090A1