A method of outputting sound and a loudspeaker

CN116569566BActive Publication Date: 2026-09-15科朗公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180075477.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-07
Filing Date
2021-09-24
Publication Date
2026-09-15
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

这导致通过扩音器相对较差的音乐再现与音乐声学表演的复杂而丰富的声音扩散之间存在显著差异,这种差异从一开始就困扰着音频系统行业

Benefits of technology

[0085] In one embodiment, the controller is further configured to generate an electrical sub-signal for one or more electrical sub-signals, such that a portion of the audio sub-band represented in the electrical sub-band increases or decreases by at least 5% per second. In this way, the portion of the audio sub-signal in the electrical sub-signal changes considerably.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116569566B_ABST
    Figure CN116569566B_ABST
Patent Text Reader

Abstract

A method of converting an audio signal into signals for a plurality of loudspeaker transducers, wherein the audio signal is divided into audio sub-signals, each audio sub-signal representing a specific frequency interval, and wherein the signal for each loudspeaker transducer comprises a time-varying portion of each audio sub-signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for outputting sound, and more particularly to a method for assigning spatial information to a sound signal. Background Technology

[0002] Known loudspeaker systems are stereo, surround, or omnidirectional setups, wherein a static loudspeaker outputs a “static” audio signal in the sense that the loudspeaker may include amplifier transducers for different frequency bands but the same amplifier transducer will receive at least substantially all electrical audio signals within its frequency band and will always output at least substantially all sounds within that frequency band.

[0003] Omnidirectional loudspeaker systems reflect sound radially from a center point in 360 degrees, dispersing the sound essentially within a vertical plane. These systems can employ different strategies to disperse mono and stereo sound; some omnidirectional systems have drivers facing directly upwards or at an angle, while others use drivers that radiate upwards into curved or conical reflectors. Although claimed to be omnidirectional, they are not true spherical loudspeaker systems, and they are all designed to emit the desired waveform in a fixed or static manner.

[0004] Conventional surround sound systems aim to enrich the fidelity and depth of sound reproduction by using multiple amplifier transducers positioned in front of, to the sides of, and behind the listener. The format and number of amplifier transducers vary among surround sound systems, but they are all designed to emit a desired waveform in a fixed or static manner. This can be independent of the listening environment of the various acoustic spaces in which they are installed, or it can be based on automated or user-defined processing of a customizable sound field tailored to a specific listening environment. What these systems have in common is that they ignore, or are designed to negate or suspend, the influence of the listening environment on playback, and once established, these fixed, customizable, or user-defined sound fields remain stable.

[0005] Therefore, these conventional systems operate for “optimal” playback in a given installation setup and an “ideal” listening position in a given listening environment. This results in a significant difference between the relatively poor musical reproduction through a loudspeaker and the complex and rich sound diffusion of a musical acoustic performance—a difference that has plagued the audio systems industry from the outset. These systems also fail to provide any enrichment for other constructed sound fields, such as studio recordings and digitally created, or otherwise non-acoustic musical or other audio content. Furthermore, the acoustic space is never perfectly constant due to the minute movements of people, objects, and other elements within it, which introduces subtle variations in sound that are important for the overall perceived quality of the sound. This audio system takes this fact into account in its processing of incoming audio signals to produce or acquire additional three-dimensional audio cues, allowing the listener to hear the sound reproduction in three dimensions, as if the listener were in the same space as the sound source. This contrasts with the two-dimensional approach, where, unless the listener is in a highly defined listening position and under specific conditions, the sound seems to enter the listening space from the outside. Summary of the Invention

[0006] A first aspect of the present invention relates to a method for outputting sound based on an audio signal, the method comprising: - Receive audio signals, - Generate multiple audio sub-signals from the audio signal. Each audio sub-signal represents an audio signal within a frequency range of 100-8000Hz. The frequency range of one sub-signal may not be completely included in the frequency range of another sub-signal. - Provides a loudspeaker comprising multiple sound output drivers or amplifier transducers, each capable of outputting sound in a range of at least 100-8000Hz, the amplifier transducers being positioned within a room or venue. - Generate an electrical sub-signal for each loudspeaker transducer, each electrical sub-signal comprising a predetermined portion of each audio sub-signal, and - Feed the electrical signals to the loudspeaker transducer. The generation of electrical sub-signals includes changing a predetermined portion of the audio sub-signal in each electrical sub-signal over time.

[0007] In this context, audio signals can be received in any format, such as analog or digital. The signal can include any number of channels, such as mono, stereo, surround sound, etc. Audio signals are often encoded by codecs such as FLAC, ALAC, APE, OFR, TTA, WV, MPEG, etc. Audio signals typically encompass all or most of the audible frequency range of 20Hz-20kHz, although audio signals can be applied to narrower frequency ranges, such as 40Hz-15kHz.

[0008] Audio signals typically correspond to a physical or sonic desired output, where correspondence means that the audio signal has the same frequency components as the sound, at least within the desired frequency band, and often has the same relative signal strength. These components and relative signal strength often change over time, but correspondence preferably does not change over time.

[0009] Audio signals can be transmitted wirelessly or via wires such as cables (optical fiber or electrical cable). Audio signals can be received from streaming or live sessions or from any kind of storage device.

[0010] The desired output is a sound signal corresponding to an audio signal or at least its frequency range. This invention focuses on sound within a frequency band from which the human ear can determine the direction of sound arrival, and the interaction of sound within this frequency range in a room or space. This frequency range can be considered as 100-8000 Hz, but can be selected, for example, between 300 and 7 kHz, 300 and 6 kHz, 400 and 4 kHz, or 200 and 6 kHz, if desired.

[0011] The auditory system uses several cues for sound source localization, including interaural temporal and horizontal differences (or intensity / loudness differences), spectral information, timing analysis, correlation analysis, and pattern matching. Interaural horizontal differences occur in the 1.500Hz–8000Hz range, where the horizontal difference is highly frequency-dependent and increases with increasing frequency. Interaural temporal differences are primarily in the 800–1.500Hz range, while interaural phase differences occur in the 80–800Hz range.

[0012] For frequencies below 400Hz, the head dimension (21.5cm ear-to-ear distance, corresponding to a 625µs interaural time delay) is less than a quarter wavelength of sound, so confusion arising from phase delay between ears begins to occur. Below 200Hz, the interaural horizontal difference becomes so small that it is almost impossible to accurately assess the direction of the input based solely on ILD. Below 80Hz, the phase difference, ILD, and ITD all become so small that it is impossible to determine the direction of the sound.

[0013] For the same head size, at frequencies above 1,600 Hz, the head dimension is larger than the wavelength of the sound wave: therefore, phase information becomes blurred. However, the ILD becomes larger, and the group delay becomes more pronounced at higher frequencies; that is, if there is a sound initiation, a transient, then the delay between the ears at this initiation can be used to determine the input direction of the corresponding sound source. This mechanism becomes particularly important in reverberant environments.

[0014] According to the present invention, a plurality of audio sub-signals are generated from an audio signal, each audio sub-signal representing an audio signal within a frequency range of 100-8000 Hz, wherein the frequency range of one sub-signal is not entirely contained within the frequency range of another sub-signal. Therefore, the sub-signals represent audio signals within a frequency range. It may be desirable for the sub-signals to include relevant portions of the audio signal. Sub-signals can be generated by applying a bandpass filter and / or one or more high-pass and / or low-pass filters to the audio signal to select the desired frequency range. The sub-audio signals can be identical to the audio signal within the frequency range, but the filters are often not ideal at their edges (extreme frequencies), where the filters often lose quality, thus allowing frequencies below the center frequency of the high-pass filter to pass through to some extent, for example.

[0015] No audio sub-signal's frequency range is completely contained within another audio sub-signal's frequency range. Therefore, each audio sub-signal represents a different frequency range of the audio signal. Consequently, for each frequency within the 100-8000Hz range, their representation in the audio sub-signals will be different. A frequency may fall within the frequency range of one or more audio sub-signals but not within the frequency ranges of other audio sub-signals. Naturally, frequency ranges can overlap. The filtering efficiency (Q value) can be selected as desired. Filtering can be performed in discrete components, in a DSP, in a processor, etc.

[0016] To output sound, or at least sound defined by an audio sub-signal, a loudspeaker is provided comprising a plurality of sound output amplifier transducers, each amplifier transducer capable of outputting sound in a desired frequency range of at least 100-8000 Hz. The amplifier transducers may be identical or have identical characteristics, such as identical impedance profiles. Alternatively, the amplifier transducers may be of different types. Preferably, the same signal, such as an audio signal or audio sub-signal, generates the same sound when output from each amplifier transducer. Nevertheless, amplifier transducers of different types or with different characteristics may still be used, such as when an electrical sub-signal for an amplifier transducer is adapted for the associated amplifier transducer so that all amplifier transducers output at least substantially the same sound; that is, each amplifier transducer has the same relationship between its sound output (e.g., for one or more frequencies) and the signal adapted and fed into the amplifier transducer to generate sound.

[0017] The loudspeaker transducer is positioned within a room or space and can be pointed in at least three different directions. The room or space may have one or more walls, ceilings, and floors. The room or space preferably has one or more sound-reflecting elements, such as walls / ceilings / floors / pillars, etc.

[0018] Alternatively, a combination of loudspeaker transducers can be chosen to represent a 180-degree sphere, such as a hemisphere detached from a flat surface. Such a flat surface could be a keyboard surface, a laptop computer surface, or a screen surface.

[0019] The direction of a loudspeaker transducer can be the main direction of the sound waves output by the loudspeaker transducer. A loudspeaker transducer can have an axis such as an axis of symmetry, along which the highest sound intensity is output or the sound intensity distribution is more or less symmetrical about the axis.

[0020] The loudspeaker transducers point in at least three distinct directions. The directions can be different if, for example, when projected onto a vertical or horizontal plane, or when translated to intersect, there is an angle of at least 5°, such as at least 10°, or such as at least 20° between two directions. The angle between two directions can be the minimum possible angle between the two directions. The two directions can be along the same axis and extend in opposite directions. Obviously, more than three distinct directions may be preferred, such as if more than 4, 5, 6, 7, 8, or 10 loudspeaker transducers are used.

[0021] A particularly interesting embodiment is one in which a loudspeaker transducer is provided on each side of the cube and oriented to output sound in a direction away from the cube. In this embodiment, six different directions are used. In another embodiment, the loudspeaker transducers are positioned on the walls, ceiling, and floor—and oriented to feed sound into the space between the loudspeaker transducers.

[0022] An electrical sub-signal is generated for each megaphone transducer. In this way, each megaphone transducer can operate independently of the others. Obviously, if a large number of megaphone transducers are used, multiple megaphone transducers can be driven or operated in exactly the same way. Such identically driven megaphone transducers can have the same or different directions.

[0023] In this context, the electrical sub-signal is a signal used by a loudspeaker transducer. This signal can be fed directly to the loudspeaker transducer or adapted to the loudspeaker transducer, for example, through amplification and / or filtering. Furthermore, the electrical sub-signal can be in any form, such as optical, wireless, or in wires. The electrical sub-signal can be encoded using any codec, or it can be digital or analog. The loudspeaker transducer may include decompression, filtering, amplification, a receiver, a DAC, etc., to receive the electrical sub-signal and drive the loudspeaker transducer.

[0024] Each electrical sub-signal can be adapted in any desired manner before being fed to the loudspeaker transducer. In one embodiment, the electrical sub-signal is amplified before being fed to the loudspeaker transducer. In this or another embodiment, the electrical sub-signal can be adapted, such as by filtering or equalization, to make its frequency characteristics suitable for the frequency characteristics of the relevant loudspeaker transducer. Different amplification and adaptation may be desired for different loudspeaker transducers.

[0025] Each electrical subsignal includes or represents a predetermined portion of each audio subsignal. For some audio subsignals, this portion may be zero. Each audio subsignal is then multiplied by a weight or factor using the stated mathematical method, and all the resulting audio subsignals are summed to form the electrical subsignal. This processing can obviously occur in a computer, processor, controller, DSP, FPGA, etc., which will then output the electrical subsignal or each electrical subsignal to be fed to a loudspeaker transducer or converted / received / adapted / amplified before being fed to the loudspeaker transducer.

[0026] Naturally, electrical and / or audio sub-signals can be stored between their generation and being fed to the amplifier transducer. Thus, a new audio format can be seen in which such signals are stored in addition to or in place of the actual audio signal.

[0027] When the electrical signal is fed to the amplifier transducer, sound is output.

[0028] Preferably, the sum of the audio sub-signals is at least substantially identical to the portion of the audio signal provided in the frequency ranges outside the audio sub-signals. Therefore, the audio sub-signals can be selected to represent that portion of the audio signal. The portions of the audio signal outside this total frequency range can be treated differently. In this context, the intensity of the sum of the audio sub-signals can be within 10% of the energy / loudness of the corresponding portion of the audio signal, such as within 5%. And, or alternatively, the energy / loudness in each frequency range of a predetermined width (such as 100Hz, 50Hz, or 10Hz) of the combined audio sub-signals can be within 10% of the energy / loudness in the same frequency range of the audio signal, such as within 5%.

[0029] Naturally, scaling or amplification is permissible, such that the overall expectation is to maintain the frequency components within that frequency range of the audio signal without obscuring it. Therefore, it can be expected that, for one, two, three, or more pairs of frequencies within a frequency range, the summed intensity of the audio subband is within 10% of the intensity of the audio signal at that frequency, such as within 5%. Thus, maintaining relative frequency intensity is desired.

[0030] Similarly, it is preferable that the sum of the electrical sub-signals is at least substantially identical to a portion of the audio signal provided in the outer frequency range of the electrical sub-signals. Thus, the electrical sub-signals can represent that portion of the audio signal. The portion of the audio signal outside this overall frequency range can be handled by other transducers. In this context, the intensity of the sum of the electrical sub-signals can be within 10% of the energy / loudness of the corresponding portion of the audio signal, such as within 5%. And, or alternatively, the energy / loudness in each frequency range of a predetermined width of the combined electrical sub-signals (such as 100Hz, 50Hz, or 10Hz) can be within 10% of the energy / loudness in the same frequency range of the audio signal, such as within 5%.

[0031] Naturally, scaling or amplification is permissible, such that the overall expectation is to maintain the frequency components within that frequency range of the audio signal without obscuring them. Therefore, it may be desirable that, for one, two, three, or more pairs of frequencies within a frequency range, the summed electrical subband at that frequency has an intensity within 10% of the audio signal intensity, such as within 5%. Thus, it is desirable to maintain the relative frequency intensity from the audio signal to the sound output.

[0032] Clearly, the electrical sound signals are expected to be coordinated so that the sounds output from all the amplifier transducers are correlated, thus correctly representing the audio signal. Therefore, the generation of audio sub-signals, electrical sub-signals, and any adaptation / amplification preferably maintains the coordination and phase of the signals.

[0033] According to the present invention, the generation of electrical sub-signals includes changing a predetermined portion of the audio sub-signal in each electrical sub-signal over time. Therefore, returning to the mathematical approach described above, the generation of each electrical sub-signal is performed by multiplying the weight of the audio sub-signal by a change over time, such that the proportion of the predetermined audio sub-signal in the electrical sub-signal changes over time.

[0034] The way in which a portion or proportion changes over time can be chosen in several ways, as described below. In one way, the audio sub-signal can be considered as a virtual amplifier transducer, each virtual amplifier transducer outputting sound corresponding to that particular signal. One or more of the real amplifier transducers then output a portion of the sound from the virtual amplifier transducer, depending on the location of the virtual amplifier transducer and, possibly, how it is oriented compared to the real amplifier transducers. This type of abstraction is also seen in standard stereo setups, where the location of a virtual sound generator (such as the string section in a classical orchestra) can be far removed from the location of the real amplifier transducers in the stereo setup, but is still represented by sound that sounds as if it comes from that virtual location.

[0035] Therefore, the portion of the audio sub-signal provided in the electrical sub-signal can be determined by the correlation between the expected position and potential direction of the virtual loudspeaker transducer corresponding to the audio sub-signal and the position and potential direction of the real loudspeaker transducer. The closer the positions and the more aligned the directions (if correlated), the larger portion of the audio sub-signal can be seen in the electrical sub-signal of that loudspeaker transducer.

[0036] This determination can be made, for example, by simulating the positions of a real and a virtual megaphone transducer on a geometry such as a sphere, where the real megaphone transducer has a fixed position but allows the virtual megaphone transducer to move in shape. The portion of the audio signal from the virtual megaphone transducer that constitutes the electrical signal for the real megaphone transducer can then be determined based on the distance between the relevant virtual and virtual-real megaphone transducers.

[0037] In one embodiment, the step of receiving an audio signal includes receiving a stereo signal. In this case, the step of generating audio sub-signals may include generating multiple audio sub-signals for each channel of the stereo audio signal.

[0038] Then, multiple audio sub-signals can be associated with the right channel and multiple audio sub-signals can be associated with the left channel. It may be desirable for an audio sub-signal from the left channel and a sub-signal from the right channel to exist in pairs, having at least substantially the same frequency range, and for the virtual amplifier transducers of such pairs to point at least substantially opposite directions or at least not in the same direction. This is achieved by accordingly selecting portions of the electrical sub-signals, knowing the positions of the amplifier transducers, and their potential directions. It may also be desirable for each pair of audio sub-signals to have more independence, and for them to be uncoordinated, or for coordination to involve avoiding complete overlap of directions between the left and right channels of the same sub-band.

[0039] In one embodiment, the step of receiving an audio signal includes receiving a mono signal and generating a second signal from the audio signal that is at least substantially out of phase with the mono signal. In this case, the step of generating audio sub-signals may include generating a plurality of audio sub-signals for each of the mono audio signal and the second signal.

[0040] These two signals can then be considered as the aforementioned left and right signals of a stereo signal, such that multiple audio sub-bands can be associated with the mono signal and multiple audio sub-bands can be associated with the other channel. It is desirable that one audio sub-band of the mono signal and one sub-band of the other signal exist in pairs, having at least substantially the same frequency range, and that the virtual amplifier transducers of such pairs point at least substantially opposite or at least not in the same direction. This is achieved by accordingly selecting portions of the electrical sub-signals, knowing the location of the amplifier transducers, and their potential directions.

[0041] Subbands of the center frequency band containing spatial audio cues can be generated or defined in several ways, with a generally better number of subbands providing better results. Setting frequency boundaries logarithmically can also be advantageous, and a subband division can be in three frequency bands with boundaries (Hz) of 100, 300, 1.200, and 4000. Another division, here in six frequency bands, can have boundaries (Hz) at 100, 200, 400, 800, 1.600, 3.200, and 6.400. This smaller number of subbands can be assigned to one, two, three, or more virtual drivers, such that the same subband is assigned to one, two, three, or more simultaneous virtual drivers at different locations on the virtual sphere. This enhances the results because the number of virtual drivers significantly contributes to the smoothness of the resulting audio sphere.

[0042] Subband division can also follow other concepts, such as the Bark scale, a psychoacoustic scale on which equal distances correspond to perceptually equal distances. The 18 subband divisions on the Bark scale would set subband boundaries (Hz) at 100, 200, 300, 400, 510, 630, 770, 920, 1080, 1270, 1480, 1720, 2000, 2320, 2700, 3150, 3700, and 4400.

[0043] For a large number of subbands, dividing them into 1 / 3 octaves would also be successful, with subband boundaries (Hz) at 111, 140, 180, 224, 281, 353, 449, 561, 707, 898, 1122, 1403, 1795, 2244, 2805, 3534, 4488, 5610, and 7069.

[0044] Subbands can also be constructed by subtraction, so the 5-subband subtraction method will give subband boundaries (Hz) at 100, 200, 400, 800, 1.600, and 3.200, and the subbands for each virtual driver will consist of combined band 1 + band 3, band 1 + band 4, band 2 + band 4, band 2 + band 5, and band 3 + band 5.

[0045] In addition, the dynamic boundary method is also possible because it can render the incoming sound onto the sound sphere more smoothly, which will be discussed in more detail elsewhere in this document.

[0046] The examples of methods described above for determining subband boundaries all provide slightly different results because the timbre or "flavor" of the sound sphere varies to some extent. However, they are all acceptable and conceptually consistent ways of preparing for adding or acquiring spatial audio cues within the audio sphere.

[0047] Once the subband boundaries are determined using any number of frequency bands as described above, it becomes possible to calculate estimates of the energy, power, loudness, or intensity of the signal in each subband. This typically involves nonlinear, time-averaged operations (such as sum-of-squares or logarithmic operations, and smoothing) and produces a number of subbands that can be compared to each other or to a target signal (such as pink noise). Through this comparison, it becomes possible to adjust the number of subbands by multiplying them by a constant gain factor. These gains can be 1) determined by a theoretical signal or noise model (such as pink noise), 2) dynamically estimated by storing the highest gain measured in real-time operation within a predetermined level, or 3) through machine learning of gains previously observed during training. Another way to adjust the number of subbands is to dynamically change the frequencies of the boundaries, as discussed in detail elsewhere in this document.

[0048] One embodiment further includes the step of: deriving a low-frequency portion of the audio signal whose frequency is below a first threshold frequency (such as 100 Hz), and including the low-frequency portion at least substantially uniformly in all electrical sub-signals or proportionally to the sub-signals in the same virtual driver. In this way, the audio signal with the low frequency is output by all audio sub-signals and / or all electrical sub-signals. Alternatively, it may be desirable to provide this low-frequency signal only in some audio sub-signals and / or some electrical sub-signals.

[0049] The alternative would be to provide this low frequency not through the aforementioned loudspeaker transducer, but through one or more separate loudspeaker transducers.

[0050] One embodiment further includes the step of: deriving a high-frequency portion of the audio signal whose frequency is above a second threshold frequency (such as 8000 Hz), and including the high-frequency portion at least substantially uniformly in all electrical sub-signals or proportionally to sub-signals in the same virtual driver. In this way, all audio sub-signals and / or all electrical sub-signals output a high-frequency audio signal. Alternatively, it may be desirable to provide this high-frequency signal only in some audio sub-signals and / or some electrical sub-signals.

[0051] An alternative is to provide this high frequency not through the aforementioned loudspeaker transducer, but through one or more separate loudspeaker transducers.

[0052] As mentioned above, the selection of the portion of the audio sub-signal represented in each electrical sub-signal can be performed based on a variety of considerations.

[0053] In one scenario, it is desirable that the acoustic energy, loudness, or intensity in each audio sub-signal and / or electrical sub-signal be the same or at least substantially the same. On the other hand, it may be desirable that the overall sound output corresponds to the audio signal such that, for example, the correspondence seen between the intensity / loudness of pairs of different frequencies should be the same or at least substantially the same in both the audio signal and the sound output. Therefore, the energy or loudness in an audio sub-band can be increased by increasing its intensity / loudness at one, several, or all frequencies in the relevant frequency range, but this may not be desirable. Alternatively, the intensity / loudness within a frequency range can be increased by widening the frequency range. This dynamic boundary method can also be used to determine two external frequency boundaries of the combined frequency bands, involving low-frequency and high-frequency components. This can be calculated before calculating the individual frequency bands, and these external frequency boundaries can be calculated such that the coherence of the combined signal emitted by the combined amplifier transducers has a desired correspondence or similarity to the input sound.

[0054] In this context, the energy, loudness, or intensity of a sound or signal can be determined in several ways. One way is by calculating the spectral envelope using a Fourier transform, which returns the magnitude of each frequency interval of the transform, corresponding to the amplitude of a specific frequency band. The resulting envelope is then integrated as a weight in the frequency domain, and the result is divided into a number of segments equal in size to the number of sub-bands. This provides new frequency boundaries for the sub-bands, as the boundaries coincide with the intersections on the frequency axes of each segment derived from the integral.

[0055] Another approach would be to use filter bank analysis to calculate the spectral envelope, where the filter bank divides the incoming sound into several independent frequency bands and returns the amplitude of each band. This can be achieved using a large number of bandpass filters, which can be 512 or more, or fewer, and the resulting band centers and loudness are integrated in a manner similar to the previous example.

[0056] Another variation of the filter bank example would be to use a non-uniform filter bank, where the number of filter bands is the same as the number of subbands in a particular implementation. The slope and center frequency of each filter in the filter bank can be used to calculate the width of the subbands, from which the frequency boundaries between the subbands can be derived.

[0057] A further variant would be to use a set of octave band filters and static weighting, followed by the integration steps outlined above.

[0058] A different approach is to use music similarity measurements developed in Music Information Retrieval (MIR), which processes the extraction and inference of meaningful and computable features from audio signals. With such a set of features, and appropriately segmented into sub-bands, a simple lookup process can determine the category of the music the system is playing and dynamically set the bands accordingly.

[0059] Finally, statistical methods (such as feature-based machine learning) can be used to make predictions and decisions about the appropriate frequencies for subband boundaries of a given audio input, where the algorithm is pre-trained with a large set of sample audio data.

[0060] Therefore, the step of generating audio sub-signals may include selecting one or more frequency ranges for the audio sub-signals such that the combined energy in each audio sub-signal is within 10% of a predetermined energy / loudness value. Thus, the energy / loudness of all audio sub-signals is within 10% of this value. Naturally, the predetermined energy / loudness value can be the average of the energy / loudness values ​​of the audio sub-signals. Alternatively, for example, the energy / loudness of the audio signal itself or its channels can be determined. This energy / loudness can be divided into the desired number of audio sub-signals for the audio signal or channels. For example, if three audio sub-signals are desired, the energy / loudness of the audio signal in the 100-8000Hz range can be determined and divided by three. Then, the energy / loudness of each audio signal should be between 90% and 110% of this calculated energy / loudness. The frequency ranges can then be adapted to achieve this energy / loudness. Again, overlapping frequency ranges are permissible.

[0061] It is reiterated that the above energy / loudness considerations may involve audio sub-signals and / or electrical sub-signals.

[0062] In embodiments of particular interest, the portion of the audio sub-signal represented in one or each electrical sub-signal varies considerably. Therefore, it may be desirable to generate the electrical sub-signal by generating, for one or more electrical sub-signals, an electrical sub-signal such that a portion of the audio sub-band represented in the electrical sub-band increases or decreases by at least 5% per second. Thus, the percentage of the energy / loudness / intensity of the audio sub-band may vary by more than 5% per second. Therefore, if the percentage is 50% at t=0, the percentage is 47.5% or less or 52.5% or more at t=1s.

[0063] Especially when the amplifier transducers are mounted on the outer surface of the enclosure, such as a speaker enclosure of any desired size and shape, the audio sub-signals can be considered as individual virtual amplifier transducer enclosures moving around within the enclosure or on the enclosure surface or a predetermined geometry. Their positions, and optionally their orientations (if not assumed to be in a predetermined orientation), are related to the positions and potential orientations of the real amplifier transducers and are used to calculate said portions or weights. The changes of these portions over time can then be obtained by simulating the rotation or movement of the individual virtual amplifier transducers within or on the shape.

[0064] Clearly, the sound output by the virtual loudspeaker transducer is the sound output by the real loudspeaker transducer receiving a portion of the audio sub-signal that forms the virtual loudspeaker transducer. The portion fed to each loudspeaker transducer, as well as the position of the loudspeaker transducer, and potentially its orientation, will determine the overall sound output from the virtual loudspeaker transducer. The virtual loudspeaker transducer can be repositioned or rotated by changing the intensity / loudness of the corresponding sound in each loudspeaker, thereby changing the portion of that audio sub-signal in the loudspeaker transducer or electrical sub-signal.

[0065] A second aspect of the present invention relates to a system for outputting sound based on an audio signal, the system comprising: - Input terminal used to receive audio signals. - A loudspeaker, comprising multiple sound output amplifier transducers, each capable of outputting sound in the range of at least 100-8000Hz, the amplifier transducers being positioned within a room or venue. - The controller is configured as follows: - Generate multiple audio sub-signals from the audio signal. Each audio sub-signal represents an audio signal within a frequency range of 100-8000Hz. The frequency range of one sub-signal may not be completely contained within the frequency range of another sub-signal. - Generate an electrical sub-signal for each loudspeaker transducer, each electrical sub-signal comprising a predetermined portion of each audio sub-signal, and - A component used to feed electrical signals to the loudspeaker transducer. The controller is configured to generate each electrical sub-signal such that a predetermined portion of the audio sub-signal in each electrical sub-signal changes over time.

[0066] In this context, the system can be a combination of independent components or a single component. Inputs, controllers, and speakers can be single components configured to receive audio signals and output sound.

[0067] Alternatively, the controller can be detachable or separable from the speaker, allowing electrical or audio signals to be generated remotely from the speaker and then fed to it.

[0068] Clearly, a controller can be one or more elements configured for communication. Therefore, an audio sub-signal can be generated in one controller while an electrical sub-signal is generated in another. As mentioned below, new codecs or encapsulations can be generated, thereby forwarding the audio or electrical sub-signals to the controller or speaker in a controlled and standardized manner, which can then interpret them and output sound.

[0069] As mentioned above, audio signals can be in any format, such as any known codec or encoding format. Audio signals can be received from live performances, streaming, or storage devices.

[0070] The input can be configured to receive signals from wireless sources, cables, optical fibers, storage devices, etc. The input can include any desired or necessary signal processing, conversion, error correction, etc., to arrive at an audio signal. Therefore, the input can be an antenna, connector, controller, or the input of another chip (such as a MAC).

[0071] The loudspeaker is configured to receive signals and output sound. In this context, the loudspeaker includes multiple amplifier transducers configured to output sound. The amplifier transducers direct sound in at least three different directions, as described above.

[0072] If multiple loudspeaker transducers are required, for example, to cover all frequency ranges encompassed by the frequency range of the audio sub-signal, then multiple loudspeaker transducers can point in the same direction. If this frequency range is wide and the loudspeaker transducers have narrower operating frequency ranges, then multiple different loudspeaker transducers can be required for each direction.

[0073] Furthermore, if the directivity of the loudspeaker transducer is too narrow, it may be desirable to provide multiple such loudspeaker transducers that are only slightly deflected to cover the specific angular intervals of the audio sub-signals in question.

[0074] As mentioned, a much larger number of directions can be used.

[0075] The electrical sub-signals will be fed to the loudspeaker transducers. A controller or portion thereof for generating the electrical sub-signals can be provided within the loudspeaker, eliminating the need for them to be transmitted to the loudspeaker. Alternatively, the loudspeaker may include an input for receiving these signals. Clearly, this input should be configured to receive such signals and process the received signal(s)(s) as needed to obtain a signal for each loudspeaker transducer. This processing can derive electrical sub-signals from a general or combined signal received from the loudspeaker input.

[0076] The frequency range discussed is at least 100-8000Hz, but can be narrower.

[0077] The controller is configured to generate multiple audio sub-signals from the audio signal. This process is further described above.

[0078] Note that the number of audio sub-signals does not need to correspond to the number of electrical sub-signals.

[0079] As mentioned above, the same or another controller can generate electrical sub-signals from the audio signal, and generate electrical sub-signals in a manner in which the portion of the audio sub-signal in each electrical sub-signal changes over time.

[0080] In one embodiment, the input is configured to receive a stereo signal. The controller can then be configured to generate multiple audio sub-signals for each channel of the stereo audio signal. The audio sub-signals corresponding to the same frequency range can then be fed to predetermined amplifier transducers, and also fed over time, such that two signals are not fed to the same amplifier transducer with too high a portion (included in the same electrical sub-signal).

[0081] In another embodiment, the input is configured to receive a mono signal. The controller can then be configured to generate a second signal from the audio signal that is at least substantially out of phase with the mono signal, and to generate multiple audio sub-signals for each of the mono audio signal and the second signal. The audio sub-signals corresponding to the same frequency range can then be fed to predetermined amplifier transducers, and also fed over time, such that the two signals are not fed to the same amplifier transducer with too high a portion (included in the same electrical sub-signal).

[0082] In one embodiment, the controller is also configured to derive a low-frequency portion of the audio signal whose frequency is below a first threshold frequency, which can be 100Hz, 200Hz, 300Hz, 400Hz, or any frequency in between, and to include the low-frequency portion at least substantially uniformly across all electrical sub-signals. Alternatively, the loudspeaker may include a separate amplifier transducer that feeds this low-frequency signal.

[0083] In one embodiment, the controller is also configured to extract a high-frequency portion of the audio signal whose frequency is above a second threshold frequency, which can be 4000Hz, 5000Hz, 6000Hz, 7000Hz, or 8000Hz or any frequency in between, and to include the high-frequency portion at least substantially uniformly across all electrical sub-signals. Alternatively, the loudspeaker may include a separate amplifier transducer that feeds this high-frequency signal.

[0084] In one embodiment, the controller is further configured to select frequency ranges for one or more of the audio sub-signals such that the combined energy, such as combined loudness, in each audio sub-signal is within 10% of a predetermined energy / loudness value. As mentioned above, it may be preferred that the energy, loudness, or intensity is the same in each audio sub-signal. To achieve this, the frequency range of each audio sub-signal can be adapted. The predetermined energy value may be, for example, the average energy or loudness value of all audio sub-signals in the channel or all audio sub-signals, or a percentage of the energy / loudness of the audio signal, such as across the entire frequency range of the audio sub-signals.

[0085] In one embodiment, the controller is further configured to generate an electrical sub-signal for one or more electrical sub-signals, such that a portion of the audio sub-band represented in the electrical sub-band increases or decreases by at least 5% per second. In this way, the portion of the audio sub-signal in the electrical sub-signal changes considerably. Attached Figure Description

[0086] Unless otherwise stated, the accompanying drawings illustrate aspects of the innovations described herein. Referring to the drawings, where similar reference numerals in several views and this specification denote similar parts, several embodiments of the currently disclosed principles are shown by way of example rather than limitation.

[0087] Figure 1 An example of an audio device is illustrated.

[0088] Figure 2 The illustration shows a sound sphere corresponding to a representative listening environment.

[0089] Figure 3 The illustration shows another possible sound sphere corresponding to another representative listening environment.

[0090] Figure 4 The illustration shows another possible sound sphere corresponding to another representative listening environment.

[0091] Figure 5 The diagram illustrates the frequency range used for spatial sound source localization.

[0092] Figure 6 The diagram illustrates the sound distribution on the amplifier transducer.

[0093] Figure 7a The diagram illustrates another sound distribution on the amplifier transducer.

[0094] Figure 7b The diagram illustrates another sound distribution on the amplifier transducer.

[0095] Figure 8 The diagram illustrates the three-dimensional directionality factor.

[0096] Figure 9 The diagram illustrates the audio processing environment.

[0097] Figure 10 The diagram illustrates another audio processing environment. Detailed Implementation

[0098] The following describes various innovative principles related to systems for providing sound spheres with smooth, changing, or constant three-dimensional spatial transitions. For example, some aspects of the disclosed principles involve audio devices configured to project a desired sound sphere or an approximation thereof throughout the listening environment.

[0099] The embodiments of such systems described in the context of method actions are merely specific examples of the intended systems, chosen as convenient illustrative examples of the disclosed principles. One or more of the disclosed principles can be incorporated into a variety of other audio systems to achieve any of the various corresponding system features.

[0100] Therefore, systems with properties different from the specific examples discussed herein can implement one or more of the currently disclosed innovative principles and can be used in applications not described in detail herein. Consequently, such alternative embodiments also fall within the scope of this disclosure.

[0101] In some implementations, the innovations disclosed herein generally relate to systems and associated techniques for utilizing multiple beams to provide a three-dimensional sound sphere, with the beams combined to provide smoothly varying sound localization information. For example, some disclosed audio systems can project sub-bands of sound within a frequency band onto a loudspeaker transducer with subtly altered or constant phase relationships and independent amplitudes. Thus, the audio system can render added or acquired spatial information onto any input audio throughout the listening environment.

[0102] As an example only, an audio device may have an array of amplifier transducers, each constituting an independent full-range transducer. The audio device includes a processor and a memory containing instructions that, when executed by the processor, cause the audio device to render a three-dimensional waveform into a 360-degree sphere, in the form of a weighted combination of individual virtual shape components, as coordinated pairs of shape components, or otherwise, by slowly moving the audio signal along the amplifier transducers through translation processing. For each amplifier transducer, the audio device can filter the received audio signal according to a specified process. When performing dynamic sound spheres, the audio device preserves the original sound across the combined sphere components as it adds the combined sphere components in acoustic space. Therefore, for the listener, the resulting sound retains the frequency envelope of the original sound but adds or gains a dynamic or constant three-dimensional audio spatialization.

[0103] This disclosure allows the three-dimensional audio rendering to be combined with the sum of signals above and below two specified thresholds, wherein audio signals outside the thresholds do not retain information about sound localization discernible by a cognitive listening device. These two ranges are each added to form two mono audio signals, which can be simultaneously sent to all amplifier transducers. Thus, the audio device can provide a complete three-dimensional spatialization recognizable by a cognitive listening device, along with independent control of all amplifier transducers in both the low-frequency and high-frequency ranges.

[0104] This disclosure allows the management of a single mono signal input on an audio device across multiple independent spherical components, either equal to the number of amplifier transducers in the device or different from the number of virtual spherical components. Each spherical component can be a subset of a frequency range, and all components can be uniformly distributed along that range as a balanced sum of components. These components can then be translated independently across all amplifier transducers on the geometric plane, or modified as polarity-reversed pairs at relative points on the geometric plane, or otherwise, and they can be positioned at any point between adjacent planes. Used in a paired stereo configuration with two devices, this system will provide a separate three-dimensional spatialization on each mono audio channel and render the left and right channels to the two audio devices separately, resulting in a three-dimensional stereo audio rendering system. Stereo pairs can also be translated individually, and no correlation is observed at relative points.

[0105] This disclosure allows management of a stereo signal on an audio system in multiple independent iterations, the number of iterations being equal to half the number of amplifier transducers in the unit. Each pair is a subset of the frequency range of the stereo signal and can be located at relative points on a geometric entity, or at any point between adjacent planes of the entity. The stereo pairs are translated equally, so a single audio device will provide a satisfactory rendering of the input stereo signal, thus avoiding the need for two devices to render the full information of the original stereo signal, while still obtaining the described three-dimensional audio cues. The result is a point-source, three-dimensional stereo audio rendering system.

[0106] Instructions stored in the processor's memory can generate adaptive frequency band divisions, allowing for the observation of equal loudness across bands if desired. This avoids sudden changes in direction caused by energy / loudness variations within very localized frequency ranges.

[0107] I. Overview

[0108] Now for reference Figure 1 and 2The audio device or speaker 10 can be placed in the room 20. The audio device 10 renders a three-dimensional sound sphere 30, in which the listener's optimal listening area coincides with the sphere 30.

[0109] Figure 3 and 4 Other exemplary representations of the positioning of device 10 are shown. Audio device 10 may correspond to the position of one or more reflective boundaries (e.g., walls 22a, 22b) relative to device 10 and the possible positions 26a, 26a of the listener coinciding with the sound spheres 30a, 30b. The rendered three-dimensional sound spheres 30a, 30b are enhanced as the waveform folds backward from the walls.

[0110] As will be explained more fully below, a three-dimensional sound sphere can be constructed by combining spherical components. The three-dimensional sound sphere depends on variations in amplitude, phase, and time along different audio frequencies or bands. A method can be designed to manage such dependencies, and the disclosed audio devices can apply these methods to acoustic or digital signals containing audio content to render them as three-dimensional sound spheres.

[0111] Section II, through reference Figure 1 The device described herein describes the principles associated with such an audio device. Section III describes the principles associated with the desired three-dimensional sound sphere, and Section IV describes the principles associated with decomposing audio content into a combination of virtual and real spherical components and reassembling them in acoustic space. Section V discloses the principles of directionality associated with the three dimensions of the audio device and their variation with frequency. Section VI describes the principles associated with an audio processor adapted to render an approximation of the desired three-dimensional sound sphere based on an input audio signal at input 51 containing audio content. Section VII describes the principles associated with a computing environment suitable for implementing the disclosed processing methods. This will include examples of machine-readable media containing instructions that, when executed, cause processor 50 of, for example, a computing environment, to perform one or more of the disclosed methods. Such instructions may be embedded in software, firmware, or hardware. Furthermore, the disclosed methods and techniques may be executed in various forms of signal processors (again, in software, firmware, or hardware).

[0112] II. Audio equipment

[0113] Figure 1 An audio device 10 is shown, comprising a loudspeaker housing 12 in which a loudspeaker array is integrated, the loudspeaker array comprising a plurality of individual loudspeaker transducers or loudspeaker transducers S1, S2, ..., S6.

[0114] Generally, a loudspeaker array can have any number of individual loudspeaker transducers, although the array shown has six loudspeaker transducers. Figure 1 The number of loudspeaker transducers depicted is for illustrative purposes. Other arrays may have more or fewer than six transducers and may have more or fewer than three axes with transducer pairs, and each axis may have only one transducer. For example, embodiments of arrays for audio devices may have 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or more loudspeaker transducers.

[0115] exist Figure 1 In the middle, the box 12 has a generally cubic shape, which defines the central axis z of the opposite angle 16 of the cubic box.

[0116] The loudspeaker transducers S1, S2, ..., S6 in the loudspeaker array shown are uniformly distributed on the plane of the cube, at a constant or substantially constant position relative to the center of the axis, and at a uniform radial distance, polarity, and azimuth angle relative to the center of the axis. Figure 1 In the middle, the loudspeaker transducers are spherically spaced about 90 degrees apart.

[0117] Other arrangements for the loudspeaker transducers are possible. For example, the loudspeaker transducers in the array can be uniformly or non-uniformly distributed within the loudspeaker housing 10. Similarly, the loudspeaker transducers S1, S2, ..., S6 can be positioned at various selected spherical locations measured from the axis center, rather than as... Figure 1 The constant distance positions are shown. For example, each loudspeaker transducer can be distributed from two or more axis points.

[0118] Each transducer S1, S2, ..., S6 can be an electric or other type of amplifier transducer, specifically designed for sound output in a particular frequency band, such as a woofer, tweeter, midrange speaker, or full-range speaker. Audio device 10 can be combined with a seventh amplifier transducer S0 to supplement the output from the array. For example, the supplementary amplifier transducer S0 can be configured to radiate selected frequencies, such as the low-end frequencies of a subwoofer. The supplementary amplifier transducer S0 can be integrated into audio device 10, or it can be housed in a separate enclosure. Additionally or alternatively, the S0 amplifier transducer can be used for high-frequency output.

[0119] Although the loudspeaker enclosure 10 is shown as a cube, other embodiments of the loudspeaker enclosure 10 have another shape. For example, some loudspeaker enclosures may be arranged as, for example, a general prismatic structure, a tetrahedral structure, a spherical structure, an elliptical structure, a ring structure, or any other desired three-dimensional shape.

[0120] III. Three-dimensional sound sphere

[0121] Refer again Figure 2 The audio device 10 can be placed in the center of the room. In this case, as described above, the three-dimensional sound spheres are evenly distributed around the audio device 10.

[0122] By projecting sound energy into a three-dimensional sphere, the user's listening experience can be enhanced compared to a two-dimensional audio system. Furthermore, in contrast to existing technologies with one-dimensional and two-dimensional sound fields, the three-dimensional listening cues provided in this disclosure are spatial and therefore immersive, similar to sound cues in the physical world.

[0123] Furthermore, the listening space disclosed herein provides an infinite number of listening positions around the device 10, because the added spatial audio cues do not operate based on an ideal listening position, as long as the entire listening field or sphere contains a uniform or nearly uniform balance of the salient features of the original sound input.

[0124] Figure 3 Depicting the situation with Figure 2 The audio devices 10 are shown in different locations. Figure 2 In the middle, the sound field 30 has a circular shape and directs very little or no sound energy to the wall 22. Although Figure 3 The three-dimensional sound sphere shown is Figure 2 The difference shown is not the same as that of wall 22 and Figure 3 Compared to the possible listening positions where the partially folded sound sphere 30 shown now overlaps. Figure 3 The sound sphere shown can fit well into the example location of the loudspeaker because the reflection from wall 22 is incompatible with the sound sphere 30, as the spherical components constantly shift along the loudspeaker transducer, thus avoiding a constant imposition of any particular frequency or band. Similarly, with Figure 2 Compared to the position of the audio device 10 shown, Figure 4 An audio device 10 is shown in another location in the room, along with a three-dimensional sound sphere 30 coinciding with the listening position (also folded accordingly by the wall position 22), and the room layout. In this particular arrangement, with Figure 3 The same situation is occurring with respect to the projection of the sound sphere 30 by means of the moving sphere component, resulting in no constant enforcement of any particular frequency or band.

[0125] In some embodiments of the audio device, the three-dimensional sound field can be modified when the audio device 10 is extremely or very obviously close to the wall 22. For example, by using polar coordinates to represent the three-dimensional sound sphere 30, where the z-axis of the audio device 10 is located at the origin, the user can modify the sound sphere 30 from a sphere to an asymmetrical three-axis ellipsoid shape by means of "drawing," such as on a touchscreen, with the amplitude of the amplifier transducer scaled relative to the z-axis of the audio device 10.

[0126] In other embodiments, the user can select from a plurality of three-dimensional asymmetric triaxial elliptic bodies stored or remotely stored by the audio device 10. If remotely stored, the audio device 10 can load the selected triaxial asymmetric elliptic body via a communication connection. And in a further embodiment, the user can “draw” the desired triaxial asymmetric elliptic body outline or existing room boundary on a smartphone or tablet as described above, and the audio device 10 can receive a representation of the desired asymmetric triaxial elliptic body or room boundary directly or indirectly from the user’s device via a communication connection. Other forms of user input besides a touchscreen can be used, as described more fully below in conjunction with a computer environment.

[0127] IV. Modal decomposition and reassembly of a three-dimensional sound sphere

[0128] Figure 5 The frequency range between 40 (localized at 100 Hz) and 45 (localized at 3 kHz) used by listeners for spatial sound source localization in three-dimensional hearing is shown as a subset of the total frequency range of the listener's hearing. Clues for sound source localization include temporal and level differences between the two ears, spectral information, timing analysis, correlation analysis, and pattern matching. This disclosure uses this knowledge of the auditory system to add or acquire spatial information to the input sound by breaking down the frequency range between 40 and 45 into multiple frequency bands (arrows) and processing these bands. The number of frequency bands can be half the number of amplifier transducers, and can be more or less than the number of transducers.

[0129] This is only an example, and not all possible embodiments, in Figure 6In this process, high-pass filter 50, band-pass filters 51, 52, and 53, and low-pass filter 54 separate the audio stream into five sub-streams or audio sub-signals. The high-pass filters remove signal components above 4kHz, and the low-pass filters remove signal components below 100Hz. The audio streams from filters 50 and 54 are outside the three-dimensional hearing range and are sent equally to all amplifier transducers S1, S2, ..., S6—or to amplifier transducer S0, depending on different methods. A copy of the signal from each frequency band of filters 51, 52, and 53 can be modified by applying a certain degree of phase shift or by polarity reversal, and then the modified signal is sent to different points, such as relative points 180 degrees from the original signal of audio device 10, as a sum of the individual signals to achieve the signal for amplifier transducers S1-S6. The resulting audio output is a mono sound with independent spatial cues added to three pairs of connected sphere components for a mono three-dimensional sound sphere. In this variation of the example, the audio streams from filters 51, 52, and 53 are sent to amplifier transducers S1, S2, ..., S6, respectively, and moved in a random or semi-random but coordinated manner. This also provides spatial cues for the mono three-dimensional sound sphere, but with significantly different properties from the previous example.

[0130] Figure 7a This represents the same scenario, but with a stereo signal input. It is presented only as an example, not as all possible embodiments. Figure 7a In this configuration, high-pass filter 60, band-pass filters 61, 62, and 63, and low-pass filter 64 split the audio into five audio streams. The audio streams from filters 60 and 64, located outside the three-dimensional hearing range, are sent equally to all amplifier transducers S1, S2, ..., S6 as a mono signal summed for the low-pass audio and high-pass audio before transmission, since they provide little or no spatial information, or as two separate audio streams for the left and / or right channels of the low-pass and high-pass audio. The audio streams from filters 61, 62, and 63, located within the three-dimensional hearing range, are sent individually, but now in pairs to amplifier transducers [S1, S2], [S3, S4], [S5, S6], or any axial point between transducers. The resulting audio output is stereo sound, with added or acquired spatial cues to provide a point source, stereo, three-dimensional sound field.

[0131] Figure 7b This represents a scenario where the stereo signal input is treated as a single mono channel. This is merely an example, not all possible implementations. Figure 7bIn this configuration, high-pass filter 70, band-pass filters 71A, 71B, 72A, 72B, 73A, 73B, and low-pass filter 74 split the audio into eight audio streams. The audio streams from filters 70 and 74, located outside the three-dimensional hearing range, are sent equally to all amplifier transducers S1, S2, ..., S6 as a mono signal summed before transmission for the low-pass audio and for the high-pass audio (because they provide little or no spatial information), or as two separate audio streams for the left and / or right channels of the low-pass and high-pass audio. The audio streams from filters 71A, 71B, 72A, 72B, 73A, 73B, located within the three-dimensional hearing range, are sent individually to amplifier transducers [S1, S2, S3, S4, S5, S6] or any axis point between transducers. The resulting audio output consists of multiple unidirectional sounds with added or acquired spatial cues, providing point sources and multiple unidirectional three-dimensional sound fields. Therefore, compared to... Figure 7a In contrast, no correlation is required between the angles of the directions of the corresponding audio sub-signals in the output (which involve the same sub-band).

[0132] V. Directional Considerations

[0133] Figure 8 This describes various aspects of the directivity factor of the sound device 10. A directivity factor ranging from 1 to ∞ is an indication of the ability of a loudspeaker transducer (or any other sound emitter) to confine applied energy to a spherical cross-section. Audio devices exhibit varying degrees of directivity across the entire audible frequency range (e.g., approximately 20 Hz to approximately 20 kHz), generally exhibiting a lower directivity factor as the frequency approaches 20 Hz and increasing with increasing frequency. Considering that the loudspeaker transducers are uniformly or nearly uniformly distributed on an even-numbered side geometry, the directivity factor of the disclosed audio device 10 is 1 or close to 1 across the entire frequency range. The individual loudspeaker transducer directivity factor of the disclosed audio device 10 is 2 or close to 2 at low frequencies and varies across the entire frequency range, but it tends to a higher value as the frequency increases. When the directivity factor is 8, each transducer will have a spherical portion that, combined with the six transducers on the aforementioned cubic enclosure, forms a complete sphere for the audio device 10. Since the directional energy used for a single loudspeaker transducer determines a given listening window as a range of angular positions chosen with a constant radius from the origin, the user's listening experience deteriorates if the user's position relative to the loudspeaker changes. This disclosure, with its much lower directionality factor compared to existing techniques in a two-dimensional sound field, offers an infinite or far greater number of desired listening positions.

[0134] To achieve the desired sonic sphere or smoothly varying spherical components (or patterns) across all frequencies, these spherical components can be equalized so that each component consistently provides a corresponding sound field with the desired frequency response. In other words, filters can be designed to provide the desired frequency response across the spherical components. Furthermore, the equalized spherical components can then be combined to render a sonic sphere with a smooth transition of spherical components across the audible frequency range and / or selected frequency bands.

[0135] VI. Audio Processor

[0136] Figure 9 A block diagram of an audio rendering processor for playing back audio content (e.g., musical works, movie soundtracks) on an audio device 10 is shown.

[0137] The audio rendering processor 50 can be a dedicated processor, such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a collection of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). In some cases, the audio rendering processor can be implemented using a combination of machine-executable instructions that, when executed by the processor, cause the audio device to process one or more input channels, as described. The rendering processor 50 is used to receive input channels of a segment of audio program content from the input audio source 51.

[0138] Input audio source 51 can provide digital or analog input. The input audio source or input 51 may include a processor programmed to run a media player application and may include a decoder that generates digital audio input to the rendering processor. For this purpose, the decoder may be able to decode encoded audio signals encoded using any suitable audio codec, such as Advanced Audio Codec (AAC), MPEG Audio Layer II, MPEG Audio Layer III, and Free Lossless Audio Codec (FLAC). Alternatively, the input audio source may include a codec that converts, for example, analog or optical audio signals from a line input into a digital form for the audio rendering processor 50. Alternatively, there may be more than one input audio channel, such as a two-channel input, i.e., the left and right channels of a stereo recording of a musical work, or there may be more than two input audio channels, such as the entire original audio soundtrack in a 5.1 surround format, for example, a film reel or movie. Other audio format examples are 7.1 and 9.1 surround formats.

[0139] The array 58 of the loudspeaker transducers can render a desired sound sphere (or an approximation thereof) based on a combination of spherical component segments 52a...52N applied to the audio content by the audio rendering processor 50. Figure 9The rendering processor 50 can conceptually be divided between the spherical component domain and the loudspeaker transducer domain. In the component domain, segmented processing 53a...53N for each constituent spherical component 52a...53N can be applied to the audio content corresponding to the desired spherical component in the manner described above. Equalizers 54a...54N can provide equalization for each corresponding spherical component 52a...52N to adjust for changes in directionality factors caused by the specific audio device 10 and any spherical adjustments towards the desired asymmetrical ellipsoidal spherical contour, as mentioned above.

[0140] In the amplifier transducer domain, a sphere domain matrix can be applied to various sphere domain signals to provide the signal to be reproduced by each corresponding amplifier transducer in array 58. Generally, the matrix is ​​an M x N matrix, where N is the number of amplifier transducers, and M = (2 x N) + (2 x O), where O represents the number of virtual sphere components. Equalizers 56a...56N can provide equalization for each corresponding sphere component 57a...57N to adjust for changes in the directionality factor caused by a particular audio device 10 and any sphere adjustment toward the desired ellipsoidal spherical profile, as mentioned above.

[0141] It should be understood that the audio rendering processor 50 is capable of performing other signal processing operations to render the input audio signal in a desired manner for playback by the transducer array 58. In another embodiment, to determine how to modify the loudspeaker transducer signal, the audio rendering processor may use adaptive filtering to determine constant or varying boundary frequencies. Figure 10 A block diagram of an audio rendering processor of an audio device 10 for rendering synthesized sounds (e.g., a numeric keypad, a digital audio workstation (DAW)) or electric and / or acoustic musical instruments is shown.

[0142] VII. Computing Environment

[0143] Figure 10A generalized example of a suitable computing environment 100 is illustrated, which may include the operation of a controller 50, wherein methods, embodiments, processes, and techniques related to, for example, the programmatic generation of sound spheres are described. The computing environment 100 is not intended to impose any limitation on the scope or functionality of the techniques disclosed herein, as each technique can be implemented in different general-purpose or special-purpose computing environments. For example, each disclosed technique can be implemented using other computer system configurations, including wearable and handheld devices, mobile communication devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, embedded platforms, network computers, minicomputers, mainframe computers, smartphones, tablet computers, data centers, etc. Each disclosed technique can also be practiced in a distributed computing environment where tasks are performed by remote processing devices that are connected via communication links or network links, or integrated into digital or analog musical instruments. In a distributed computing environment, program modules can reside both locally and in remote memory storage devices.

[0144] The computing environment 100 includes at least one central processing unit 110 and memory 120. Figure 10 In this context, the most basic configuration 130 is included within the dashed lines. The central processing unit 110 executes computer-executable instructions and can be a real or virtual processor. In a multiprocessor system, multiple processing units execute computer-executable instructions to increase processing power, thus multiple processors can operate simultaneously. Memory 120 can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of both. Memory 120 stores software 180a, which, when executed by the processor, can, for example, implement one or more of the innovative techniques described herein.

[0145] The computing environment may have additional features. For example, computing environment 100 includes storage device 140, one or more input devices 150, one or more output devices 160, and one or more communication connections 170. Interconnection mechanisms (not shown), such as buses, controllers, or networks, interconnect the components of computing environment 100. Typically, operating system software (not shown) provides an operating environment for other software executing in computing environment 100 and coordinates the activities of the components of computing environment 100.

[0146] Storage device 140 may be removable or non-removable and may include machine-readable media of a selected form, including magnetic disks, magnetic tapes or cassettes, non-volatile solid-state storage, CD-ROMs, CD-RWs, DVDs, magnetic tapes, optical data storage devices and carrier waves, or any other machine-readable media that can be used to store information and can be accessed within computing environment 100. Storage device 140 stores instructions for software 180b that can implement the techniques described herein.

[0147] Storage device 140 can also be distributed across a network, enabling software instructions to be stored and executed in a distributed manner. In other embodiments, some of these operations may be performed by specific hardware components containing hard-wired logic. These operations can alternatively be performed by any combination of programmed data processing components and fixed hard-wired circuit components.

[0148] One or more input devices 150 may be touch input devices, such as a keyboard, keypad, mouse, pen, touchscreen, touchpad or trackball, voice input device, scanning device, or another device that provides input to computing environment 100. For audio, one or more input devices 150 may include a microphone or other transducer (e.g., a sound card or similar device that accepts audio input in analog or digital form), or a computer-readable medium reader that provides audio samples to computing environment 100.

[0149] One or more output devices 160 may be a monitor, printer, speaker transducer, DVD burner, or another device that provides output from computing environment 100.

[0150] One or more communication connections 170 enable communication with another computing entity via a communication medium (e.g., a network connection). The communication medium transmits information such as computer-executable instructions, compressed graphics information, processed signal information (including processed audio signals), or other data in modulated signals.

[0151] Therefore, the disclosed computing environment is suitable for performing the orientation estimation and audio rendering processes disclosed herein.

[0152] Machine-readable media are any available media that can be accessed within computing environment 100. By way of example and not limitation, for computing environment 100, machine-readable media include memory 120, storage device 140, communication media (not shown), and any combination of the above. Tangible machine-readable (or computer-readable) media do not include transient signals.

[0153] As explained above, some of the disclosed principles can be implemented in a tangible, non-transitory machine-readable medium (e.g., microelectronic memory) on which instructions are stored, programming one or more data processing components (collectively referred to herein as "processors") to perform the digital signal processing operations described above, including estimation, adaptation, computing, calculating, measurement, adjustment (by audio processor 50), sensing, measuring, filtering, addition, subtraction, inversion, comparison, and decision-making. In other embodiments, some of these operations (machine-processed) can be performed by specific electronic hardware components containing hard-wired logic (e.g., dedicated digital filter blocks). Those operations can alternatively be performed by any combination of programmed data processing components and fixed hard-wired circuit components.

[0154] The audio device 10 may include a loudspeaker enclosure 12 configured to generate sound. The audio device 10 may also include a processor and a non-transitory machine-readable medium (memory) storing instructions therein, which, when executed by the processor, automatically perform three-dimensional sphere construction and supporting processing, as described herein.

[0155] The examples described above generally relate to apparatus, methods, and related systems for rendering audio, and more specifically to providing desired three-dimensional spherical patterns. However, embodiments other than those described in detail above are contemplated based on the principles disclosed herein and any accompanying changes to the configuration of the various apparatuses described herein.

[0156] Orientation and other relevant references (e.g., up, down, top, bottom, left, right, back, forward, etc.) may be used to facilitate the discussion of the figures and principles herein, but are not intended to be limiting. For example, certain terms such as “up,” “down,” “upper,” “lower,” “horizontal,” “vertical,” “left,” “right,” etc., may be used. Where applicable, the use of such terms provides some clarity in dealing with relationships, particularly with respect to the illustrated embodiments. However, such terms are not intended to imply absolute relationships, positions, and / or orientations. For example, with respect to an object, simply flipping the object over can make the “upper” surface become the “lower” surface. However, it remains the same surface, and the object remains unchanged. As used herein, “and / or” means “and” or “or,” as well as “and” and “or.” Moreover, for all purposes, all patent and non-patent literature cited herein is incorporated herein by reference in its entirety.

[0157] The principles described above in conjunction with any particular example can be combined with the principles described in conjunction with another example described herein. Therefore, this detailed description should not be construed as limiting, and upon review of this disclosure, those skilled in the art will recognize the various signal processing and audio presentation techniques applicable to the various conceptual designs described herein.

[0158] Furthermore, those skilled in the art will recognize that the exemplary embodiments disclosed herein can be adapted to various configurations and / or uses without departing from the disclosed principles. Applying the principles disclosed herein, it is possible to provide a variety of systems suitable for providing a desired three-dimensional spherical sound field. For example, modules identified as part of a given computing engine in the above description or figures may be divided, distributed across one or more modules, or omitted entirely, differently from those described herein. Similarly, such modules may be implemented as part of different computing engines without departing from some of the disclosed principles.

[0159] The prior description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the disclosed innovations. Various modifications to those embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of this disclosure. Therefore, the claimed invention is not intended to be limited to the embodiments shown herein, but rather to be consistent with the full scope of the language of the claims, wherein elements referred to in the singular (such as by using the article “a” or “an”) do not mean “one and only one”, but rather “one or more” unless specifically stated otherwise. All structural and functional equivalents of features and methodological behavior known or to be known thereafter by those skilled in the art throughout the various embodiments described herein are intended to be covered by the features described and claimed herein. Moreover, nothing disclosed herein is intended to be exclusive to the public, whether such disclosure is expressly stated in the claims. Claim statements should not be interpreted unless expressly stated using the phrases “component for…” or “step for…”.

[0160] Therefore, given that the disclosed principles can be applied to many possible embodiments, we reserve the right to claim any and all combinations of the features and techniques described herein as understood by those skilled in the art, including, for example, all within the scope of the technology.

Claims

1. A method for outputting sound based on an audio signal, the method comprising: - Receive audio signals, - Generate multiple audio sub-signals from the audio signal. Each audio sub-signal represents an audio signal within a frequency range of 100-8000Hz. The frequency range of one sub-signal may not be completely included in the frequency range of another sub-signal. - Provides loudspeakers comprising multiple sound output amplifier transducers, each capable of outputting sound in a range of at least 100-8000Hz, with the amplifier transducers positioned within a room or venue. - Generate an electrical sub-signal for each loudspeaker transducer, each electrical sub-signal comprising a predetermined portion of each audio sub-signal, and - Feed the electrical signals to the loudspeaker transducer. The generation of electrical signals includes: The predetermined portion of the audio sub-signal in each electrical sub-signal is changed over time, and Provides electrical sub-signals with the same sound energy, loudness, or intensity.

2. The method of claim 1, wherein the step of receiving the audio signal includes receiving a stereo signal, and wherein the step of generating audio sub-signals includes generating a plurality of audio sub-signals for each channel of the stereo audio signal.

3. The method of claim 1, wherein the step of receiving the audio signal includes receiving a mono signal and generating a second signal from the audio signal that is inversely phase to the mono signal, and wherein the step of generating audio sub-signals includes generating a plurality of audio sub-signals for each of the mono audio signal and the second signal.

4. The method according to any one of claims 1-3, further comprising the step of: The low-frequency portion of the audio signal with a frequency below a first threshold frequency is extracted and this low-frequency portion is uniformly included in all electrical sub-signals.

5. The method according to any one of claims 1-3, further comprising the step of: The high-frequency portion of the audio signal with a frequency higher than the second threshold frequency is extracted and this high-frequency portion is uniformly included in all electrical sub-signals.

6. The method according to any one of claims 1-3, wherein the step of generating an audio sub-signal comprises selecting a frequency range for one or more audio sub-signals in the audio sub-signal, such that the energy / loudness in each audio sub-signal is within 10% of a predetermined energy / loudness value.

7. The method according to any one of claims 1-3, wherein the step of generating an electrical sub-signal comprises, for one or more electrical sub-signals, generating the electrical sub-signal such that a portion of an audio sub-band represented in the electrical sub-band increases or decreases by at least 5% per second.

8. A system for outputting sound based on an audio signal, the system comprising: - Input terminal, used to receive audio signals. - A loudspeaker, comprising multiple sound output amplifier transducers, each capable of outputting sound in the range of at least 100-8000Hz, the amplifier transducers being positioned within a room or venue. - The controller is configured as follows: - Generate multiple audio sub-signals from the audio signal. Each audio sub-signal represents an audio signal within a frequency range of 100-8000Hz. The frequency range of one sub-signal may not be completely included in the frequency range of another sub-signal. - Generate an electrical sub-signal for each loudspeaker transducer, each electrical sub-signal comprising a predetermined portion of each audio sub-signal, and - A component used to feed electrical signals to the loudspeaker transducer. The controller is configured to generate each of the electrical sub-signals such that: A predetermined portion of the audio sub-signal in each electrical sub-signal changes over time, and The sound energy, loudness, or intensity of electronic signals are the same.

9. The system of claim 8, wherein the input is configured to receive a stereo signal, and wherein the controller is configured to generate a plurality of audio sub-signals for each channel of the stereo audio signal.

10. The system of claim 8, wherein the input is configured to receive a mono signal, and wherein the controller is configured to generate a second signal from the audio signal that is inversely phase to the mono signal, and to generate a plurality of audio sub-signals for each of the mono audio signal and the second signal.

11. The system according to any one of claims 8-10, wherein the controller is further configured to derive a low-frequency portion of the audio signal whose frequency is below a first threshold frequency, and to uniformly include the low-frequency portion in all electrical sub-signals.

12. The system according to any one of claims 8-10, wherein the controller is further configured to derive a high-frequency portion of the audio signal whose frequency is higher than a second threshold frequency, and to uniformly include the high-frequency portion in all electrical sub-signals.

13. The system according to any one of claims 8-10, wherein the controller is further configured to select a frequency range for one or more audio sub-signals such that the energy / loudness in each audio sub-signal is within 10% of a predetermined energy / loudness value.

14. The system according to any one of claims 8-10, wherein the controller is further configured to generate, for one or more electrical sub-signals, such that a portion of the audio sub-band represented in the electrical sub-band increases or decreases by at least 5% per second.

Citation Information

Patent Citations

  • Method and apparatus for three-dimensional acoustic field encoding and optimal reconstruction

    CN102326417A

  • Transmission-agnostic presentation-based program loudness

    CN107112023A