Apparatus, method, and computer product for selecting spatial sound sources
By selecting a subset of spatial sound sources and applying frequency-dependent level adaptation and equalization processing, the problem of unbalanced sound source levels in spatial audio scenarios is solved, achieving sound source clarity and spatial distribution optimization, and improving the perception effect of sound scenarios.
Patent Information
- Application Number
- CN202180017351.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2021-02-17
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2041-02-17
AI Technical Summary
Existing technologies struggle to effectively process and control the frequency-related levels of multiple sound sources in a spatial audio scene, resulting in an uneven spatial distribution of the sound scene and blurred sound source perception.
By selecting a subset of spatial sound sources, applying frequency-related level adaptation and equalization processing, adjusting the frequency component levels of the sound sources, redistributing sound energy to compensate for the spatial distribution of the sound sources, and adjusting the position and diffusion degree of the sound sources through beamforming algorithms.
It improves spatial audio scenes, clarifies sound source perception, avoids spatial separation or overlap of sound sources, and enhances the spatial distribution and perception quality of sound scenes.
Smart Images

Figure CN115152250B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the invention relate to spatial audio. BACKGROUND
[0002] Spatial audio describes the capture / processing / rendering of audio which includes spatial sound sources 20 at specific locations 22 in a sound scene 10. SUMMARY
[0003] According to different but not necessarily all embodiments, there is provided an apparatus comprising means for:
[0004] applying equalization to a subset of the plurality of spatial sound sources included in the sound scene to modify the sound scene,
[0005] wherein the spatial sound sources are associated with respective locations in the sound scene,
[0006] wherein the equalization comprises frequency-dependent level adaptation, and
[0007] wherein the subset comprises a plurality of the spatial sound sources but does not comprise all of the plurality of spatial sound sources.
[0008] In some but not necessarily all examples, the means for applying equalization to the subset of the plurality of spatial sound sources is configured to apply a common equalization to the subset of the spatial sound sources.
[0009] In some but not necessarily all examples, the subset of the plurality of spatial sound sources is selected from the plurality of spatial sound sources in dependence on user input.
[0010] In some but not necessarily all examples, the user input comprises:
[0011] an indication of a spatial sound source or a location;
[0012] a frequency indication; and optionally,
[0013] an indication of emphasis or de-emphasis.
[0014] In some but not necessarily all examples, the user input directly or indirectly indicates a spatial sound source having a first characteristic, and wherein the apparatus comprises means for selecting the subset of the plurality of spatial sound sources to have the first characteristic.
[0015] In some but not necessarily all examples, the first characteristic is that the spatial sound source has a frequency-specific volume greater than a threshold value.
[0016] In some but not necessarily all examples, the apparatus comprises means for spatially redistributing sound energy to spatially compensate for the equalization of the subset of the spatial sound sources.
[0017] In some, but not necessarily all examples, the apparatus comprises means for performing the following: adapting one or more properties of one or more of the spatial sound sources not comprised in the subset of the plurality of spatial sound sources.
[0018] adapting one or more properties of one or more of the spatial sound sources not comprised in the subset of the plurality of spatial sound sources.
[0019] In some, but not necessarily all examples, the apparatus comprises means for adapting one or more of the spatial sound sources in the subset of the plurality of spatial sound sources to make them more diffuse.
[0020] In some, but not necessarily all examples, the apparatus comprises means for performing the following: adapting one or more of the spatial sound sources not in the subset of the plurality of spatial sound sources to make one or more of the sound sources less diffuse, and / or changing a position of one or more of the spatial sound sources.
[0021] In some, but not necessarily all examples, the apparatus comprises means for representing the spatial sound sources as respective spatially located time-frequency tiles, and means for performing the following: preventing a time-frequency tile from becoming spatially separated or significantly different from other contemporaneous time-frequency tiles of the same spatial sound source; and / or preventing a significantly differently located time-frequency tile from spatially overlapping other significantly differently located contemporaneous time-frequency tiles of other spatial sound sources.
[0022] In some, but not necessarily all examples, the apparatus comprises means for capturing and / or processing and / or rendering a sound scene comprising a plurality of spatial sound sources, wherein the spatial sound sources are associated with respective positions in the sound scene.
[0023] In some, but not necessarily all examples, the apparatus is configured as a headphone, a controller for a loudspeaker, or a spatial sound capture device.
[0024] According to various, but not necessarily all, embodiments, there is provided a method comprising:
[0025] applying an equalization to a subset of the plurality of spatial sound sources comprised in the sound scene to modify the sound scene,
[0026] wherein the spatial sound sources are associated with respective positions in the sound scene,
[0027] wherein the equalization comprises a frequency dependent level adaptation, and
[0028] wherein the subset comprises a plurality of the spatial sound sources, but does not comprise all of the plurality of spatial sound sources.
[0029] According to various, but not necessarily all, embodiments there is provided a computer program which, when run on a computer, causes:
[0030] applying equalization to a subset of the plurality of spatial sound sources included in the sound scene to modify the sound scene,
[0031] wherein the spatial sound sources are associated with respective positions in the sound scene,
[0032] wherein the equalization comprises frequency-dependent level adaptation, and
[0033] wherein the subset comprises the plurality of spatial sound sources, but does not comprise all of the plurality of spatial sound sources.
[0034] According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing:
[0035] applying equalization to at least one of the plurality of spatial sound sources included in the sound scene to modify the sound scene,
[0036] wherein the spatial sound sources are associated with respective positions in the sound scene;
[0037] wherein the equalization comprises a frequency-dependent level filter,
[0038] spatially redistributing sound energy to spatially compensate for the equalization of the at least one of the plurality of spatial sound sources.
[0039] In some, but not necessarily all, examples the means for applying equalization to at least one of the plurality of spatial sound sources is configured to apply equalization to all of the plurality of spatial sound sources. In some, but not necessarily all, examples the means for applying equalization to at least one of the plurality of spatial sound sources is configured to apply equalization to only one of the plurality of spatial sound sources. In some, but not necessarily all, examples the means for applying equalization to at least one of the plurality of spatial sound sources is configured to apply equalization to a subset of the plurality of spatial sound sources, wherein the subset comprises the plurality of spatial sound sources, but does not comprise all of the plurality of spatial sound sources.
[0040] According to various, but not necessarily all, embodiments there is provided a method comprising:
[0041] applying equalization to at least one of the plurality of spatial sound sources included in the sound scene to modify the sound scene,
[0042] wherein the spatial sound sources are associated with respective positions in the sound scene;
[0043] wherein the equalization comprises a frequency dependent level filter,
[0044] re-distributing the sound energy spatially to spatially compensate for at least one of the plurality of spatial sound sources.
[0045] According to various, but not necessarily all, embodiments there is provided a computer program which, when run on a computer, causes:
[0046] applying equalization to at least one of the plurality of spatial sound sources included in the sound scene to modify the sound scene,
[0047] wherein the spatial sound sources are associated with respective positions in the sound scene;
[0048] wherein the equalization comprises a frequency dependent level filter,
[0049] re-distributing the sound energy spatially to spatially compensate for at least one of the plurality of spatial sound sources.
[0050] According to various, but not necessarily all, embodiments there is provided an example as claimed in the appended claims. BRIEF DESCRIPTION OF DRAWINGS
[0051] Some example embodiments will now be described with reference to the drawings, in which:
[0052] Figure 1A and Figure 1B Example embodiments of the subject matter described herein;
[0053] Figure 2 Another example embodiment of the subject matter described herein;
[0054] Figure 3A and Figure 3B Another example embodiment of the subject matter described herein;
[0055] Figure 4A , Figure 4B , Figure 4C Further example embodiments of the subject matter described herein;
[0056] Figure 5A , Figure 5B , Figure 5C Another example embodiment of the subject matter described herein;
[0057] Figure 6 Another example embodiment of the subject matter described herein;
[0058] Figure 7 Another example embodiment of the subject matter described herein;
[0059] Figure 8A 、 Figure 8B 、 Figure 8C Further example embodiments of the subject matter described herein are shown. DETAILED DESCRIPTION
[0060] The accompanying drawings illustrate a controlled equalization for spatial audio.
[0061] Spatial audio describes the capture / processing / rendering of audio where the audio content comprises a spatial sound source 20 at a particular position 22 in a sound scene 10. The position can be defined, for example, using a two-dimensional position (e.g. (x, y)) or a three-dimensional position (e.g. (x, y, z)), or can be defined, for example, using a one-dimensional bearing (e.g. an azimuthal bearing ) or a two-dimensional bearing (e.g. an azimuthal and polar bearing ).
[0062] The spatial sound source 20 has a level (volume). Applying an equalization to the spatial sound source 20 comprises a frequency-dependent level adaptation. The spatial sound source 20 has different levels for different frequencies, and the equalization adapts one or more of these levels. The equalization adjusts the balance between the levels of the frequency components of the spatial sound source 20. The equalization may, for example, be discrete and change the level for each of a plurality of fixed (or variable) frequency bands, or may, for example, define a center frequency and a bandwidth.
[0063] Figure 1A and Figure 1B An example of a sound scene 10 comprising a plurality of spatial sound sources 20 n is illustrated. The spatial sound sources 20 n are associated with respective positions 22 n in the sound scene 10. For example, each of the spatial sound sources 20 n is located at a position 22 n within the sound scene 10.
[0064] For explanatory purposes, the spatial sound sources 20 n are illustrated visually. However, it will be appreciated that they are not visible; although they can be associated with visible objects (as illustrated).
[0065] For explanatory purposes, the set 30 m of frequency-dependent levels 32 n is illustrated visually. However, it will be appreciated that they are not visible in the sound scene 10; although they can be displayed in a user input interface, for example, in association with representations of visible objects (as illustrated).
[0066] The set 30 m of frequency-dependent levels 32 nare associated with respective spatial sound sources 20 n For example, for each spatial sound source 20 n there is a set 30 m of frequency-dependent levels 32 n In this example, but not necessarily in all examples, the frequency-dependent levels 32 m relate to the same frequency range f n for each spatial sound source 20 m In this example, but not necessarily in all examples, the frequency range f m is continuous. The frequency-dependent levels 32 m may be different levels for different spatial sound sources 20 n For example, in Figure 1A the frequency-dependent levels 32 a for the frequency range f a are higher for the spatial sound sources 201, 202 compared to the spatial sound source 203.
[0067] If the equalization is applied to the spatial sound sources 201, 202 (but not to the spatial sound source 203), a modified sound scene 10 is obtained, e.g. as illustrated in Figure 1B In Figure 1B the frequency-dependent levels 32 a for the frequency range f a are adapted (in this example increased) for the spatial sound sources 201, 202, and the frequency-dependent levels 32 a for the frequency range f a are not adapted (in this example not increased) or not equally adapted (in this example increased) for the spatial sound source 203.
[0068] Thus, Figure 1A and Figure 1B illustrate a method, which is further illustrated in Figure 2 .
[0069] In Figure 2 the method 300 comprises, at block 302, applying an equalization to a subset 24 {spatial sound sources 201, 202} of the plurality of spatial sound sources 20 included in the sound scene 10 to modify the sound scene 10. The spatial sound sources 20 n are associated with respective positions 22 n in the sound scene 10. For example, each spatial sound source 20 is associated with a position 22 in the sound scene 10. The equalization comprises a frequency-dependent level adaptation. The subset 24 comprises a plurality of spatial sound sources 20, but not all of the plurality of spatial sound sources 20. Figure 1BThe subset 24 in the example includes spatial sound sources 201 and 202, but excludes spatial sound source 203.
[0070] Applying equalization to a subset 24 of multiple spatial sound sources 20 may, for example, include: adjusting the frequency components f of the subset 24 of the multiple spatial sound sources 20. m Frequency correlation level 32 m The balance between them.
[0071] In some, but not necessarily all, examples, applying equalization to a subset 24 of multiple spatial sound sources 20 includes applying a common (shared) equalization to a subset 24 of spatial sound sources 20. For example, for spatial sound source 201 (L a1 The frequency range f a Frequency correlation level 32 a and for spatial sound source 202 (L a2 The frequency range f a Frequency correlation level 32 a It can be adapted using the same absolute value X. Therefore, L a1 (After equilibrium) = L a1 (Before equilibrium) +X and L a2 (After equilibrium) = L a2 (Before equalization) +X. The level is expressed in decibels above. The adjustment can be positive (increase or emphasize) or negative (decrease or deemphasize).
[0072] In some, but not necessarily all, examples, a subset 24 of multiple spatial sound sources 20 is selected from multiple spatial sound sources 20, depending on user input 200. Figure 3A and Figure 3B The illustration shows an example of a user entering 200.
[0073] In this example, user input interface 210 displays frequency-dependent levels 32 for spatial sound source 201. m The set 301 is represented. Frequency-related level 32. m The set of 30 n The representation can be associated with the corresponding spatial sound source 20 n Related. For example, in some examples, 20 can be targeted at each spatial sound source. n Display frequency-related level 32 m The set of 30 n The expression .
[0074] In this example, the user input interface 210 includes a touch-sensitive display. The user selects a frequency-specific correlation level 32 from a specific set 301 of frequency-specific correlation levels 32 associated with a specific spatial sound source 201. a .
[0075] Then, the user sets the frequency-related level to 32. a Drag (201) to the desired fit value A. After applying equalization to a subset 24 of multiple spatial sound sources 20, the desired fit value A will be the frequency-dependent level 32 for the spatial sound sources 201. a The value of .
[0076] In some examples, the user selects the location 221 of the spatial sound source 201 in sound scene 10 or the displayed representation of the spatial sound source 201 in sound scene 10 via the user input interface 210 of the touch-sensitive display. In response, the user input interface 210 displays a frequency-related level 32 associated with the spatial sound source 201. m Set 301. Then, the user can select set 30. j Frequency correlation level 32 i The representation is then dragged to the desired fit value.
[0077] In these examples, user input 200 includes an indication of spatial sound source 20 or location 22, and includes a frequency indication (a user-selected representation of frequency-dependent level 32). In this example, the user input also includes an indication of emphasis (increasing level) or de-emphasis (decreasing level), and additionally provides control over the amount of emphasis (level increase) or de-emphasis (level decrease). However, in other examples, the emphasis or de-emphasis indication may have pre-programmed default values and meanings.
[0078] In these examples, user input directly (by selecting a representation of spatial sound source 201) or indirectly (by selecting a location 221 in the sound scene 10 associated with spatial sound source 201) indicates a spatial sound source 201 having a first characteristic. A subset 24 of multiple spatial sound sources 20 is selected that share a common first characteristic. The subset includes Figure 1B Spatial sound sources 201 and 202 are included. Other unselected spatial sound sources 20 among the multiple unselected spatial sound sources 20 for the first subset 24 do not possess the first characteristic. The other unselected spatial sound sources are... Figure 1B Spatial sound source 203.
[0079] exist Figure 1B In the example, the first characteristic is for the frequency range f a Frequency correlation level 32 a The spatial sound sources 201 and 202 in subset 24 have a level greater than the threshold T for a specific frequency range f. a Frequency correlation level 32 a The spatial sound source 203 (not in subset 24) has a frequency range f that is less than the threshold T. a Frequency correlation level 32a Thus, a first characteristic is that the spatial sound sources 20 have a frequency-specific volume that is greater than a threshold T.
[0080] The first characteristic can for example depend on the user input.
[0081] For example, a user indication of a frequency (frequency-dependent level 32 i ) can for example specify a first parameter of the first characteristic, the frequency f i ) can for example specify a first parameter of the first characteristic, the frequency f i .
[0082] The first characteristic can depend on a user selection of spatial sound sources 20.
[0083] For example, if the user input 200 additionally provides an indication to emphasize (increase level) or de-emphasize (decrease level), and additionally provides a control to emphasize amount (level increase) or de-emphasize amount (level decrease), this information can specify a second parameter of the first characteristic, the level threshold T. The level threshold T can for example be set to a value similar to the level L before user adaptation or some percentage of that level. For example, when the level is increased by equalization, the threshold T = L (before equalization) * p, where p is less than 1 (e.g. 0.7 or 0.8). The threshold needs to be exceeded in the positive sense, i.e. the level needs to be greater than the threshold T for the other spatial sound sources 20 in the subset 24. For example, when the level is decreased by equalization, the threshold T = L (before equalization) * p, where p is greater than 1 (e.g. 1.2 or 1.3) when the level is decreased by equalization. The threshold needs to be exceeded in the negative sense, i.e. the level needs to be less than the threshold T for the other spatial sound sources 20 in the subset 24.
[0084] The subset 24 of the plurality of spatial sound sources 20 can be selected automatically or semi-automatically after the user input 200.
[0085] The equalization applied to the subset 24 of spatial sound sources 20 changes the spatial distribution of sound energy in the sound scene 10. In some, but not necessarily all, examples, the method 300 comprises: Figure 2 re-distributing the sound energy spatially at block 304 to at least partially spatially compensate for the equalization of the subset 24 of spatial sound sources 20.
[0086] The equalization applied to a target spatial sound source 20 in the subset 24 has an impact on the perceived location of that target spatial sound source 20. At block 304, the method 300 can comprise one or more steps to improve the impact.
[0087] For example, block 304 can include adapting one or more characteristics of the spatial sound sources 20. In some examples, this includes adapting one or more spatial sound sources 20 in the subset 24 of the plurality of spatial sound sources to make them more diffuse. Figure 4A The effect of making the spatial sound sources 20 more diffuse is illustrated. The spatial sound sources are less localized and spread over a larger area of the sound scene 10. In the limit, it can be fully diffuse, i.e. ambient sound.
[0088] Additionally or alternatively, block 304 can include adapting one or more characteristics of one or more spatial sound sources 20 in the plurality of spatial sound sources 20 that are not included in the subset 24 of the plurality of spatial sound sources 20. In some examples, block 304 includes adapting one or more spatial sound sources in the plurality of spatial sound sources 20 that are not in the subset 24 of the plurality of spatial sound sources to make the one or more spatial sound sources less diffuse. Figure 4B The effect of making the spatial sound sources 20 less diffuse is illustrated. The spatial sound sources are more localized and spread over a smaller area of the sound scene 10. In some examples, block 304 additionally or alternatively includes changing the position 22 of one or more spatial sound sources 20. Figure 4C The effect of changing the position of the spatial sound sources 20 is illustrated.
[0089] Figure 5A Each of the spatially localized time-frequency tiles 70 i illustrates a representation of each spatial sound source 20 as a spatially localized time- frequency tile 70. The audio signal representing the spatial sound source 20 is time- segmented, then frequency segmented to form the time-frequency tiles.
[0090] Figure 5A , Figure 5B and Figure 5C Each of the figures in Figure 5A , Figure 5B and Figure 5C illustrates multiple contemporaneous (same time period) time-frequency tiles 70 for the same spatial sound source. In each of the figures in
[0091] Figure 5A illustrates the spatial distribution of the time-frequency tiles 70 at time tl before equalization. Figure 5B illustrates the spatial distribution of the time-frequency tiles 70 at a time after equalization (before improvement), e.g. after block 302 but before block 304. Figure 5C illustrates the spatial distribution of the time-frequency tiles 70 at a time after equalization and improvement (e.g. after block 304).
[0092] Figure 5BThe effect of equalization is illustrated. As a result of equalization, there is a spatial separation of the time-frequency tiles 70. The spatial separation can cause a single spatial sound source 201 to be perceived as two distinct spatial sound sources 20. The spatial separation can cause a spatial sound source 201 to become spatially ambiguous with another spatial sound source 203 that was distinct in space before equalization Figure 5A ). The spatial separation can cause a spatial sound source 201 to become spatially ambiguous because it overlaps or is close to another spatial sound source 203.
[0093] The effect of improvement can be understood by comparing Figure 5B and Figure 5C The effect of improvement can be understood by comparing
[0094] The spatial separation of the time-frequency tiles 70 is improved by diffusing some or all of the time-frequency tiles 70 associated with a spatial sound source 201. This prevents a single spatial sound source 201 from being perceived as two (or more) distinct spatial sound sources 20.
[0095] A spatial sound source 201 remains distinct from another spatial sound source 203 (distinct in space before equalization) by diffusing the other spatial sound source 203 less (more limited in spatial extent), e.g., by diffusing one, some, or all of the time-frequency tiles 70 associated with the other spatial sound source 203 less (more limited in spatial extent).
[0096] A spatial sound source 201 remains distinct from another spatial sound source 203 (distinct in space before equalization) by repositioning the other spatial sound source 203, e.g., by moving one, some, or all of the time-frequency tiles 70 associated with the other spatial sound source 203.
[0097] Thus, the above-described methods can prevent time-frequency tiles 70 from becoming spatially separated or distinct from other contemporaneous time-frequency tiles of the same spatial sound source 20; and / or prevent spatially overlapping of distinctively positioned time-frequency tiles 70 of a spatial sound source 201 with other contemporaneous time-frequency tiles 50 of one or more other spatial sound sources 203.
[0098] A beamforming algorithm (e.g., vector-based amplitude panning) can be used to estimate how a time-frequency tile representation of a spatial sound source 20 is repositioned as a level associated with the time-frequency tile is increased. A time-frequency tile can be modeled as a weighted linear combination of audio signals at audio transducers. Increasing the level decreases the distance from the audio transducers to the spatial sound source (audio intensity scales inversely as the square of the distance). The effect of these changes in distance on the spatial sound source location can be computed.
[0099] Figure 6 An example of a controller 102 is illustrated. Implementations of the controller 102 can be as controller circuitry. The controller 102 can be implemented solely in hardware, have certain aspects included in software with firmware separately, or can be a combination of hardware and software (including firmware).
[0100] As Figure 6 illustrated in
[0101] The processor 110 is configured to read from and write to memory 120. The processor 110 can also include an output interface and an input interface, via which data and / or commands are output by the processor 110 and input to the processor 110, respectively.
[0102] The memory 120 stores a computer program 122 comprising computer program instructions (computer program code) that, when loaded into the processor 110, control the operation of the apparatus 100. The computer program instructions of the computer program 122 provide the logic and routines Figure 2 illustrated in
[0103] The apparatus 100 therefore comprises:
[0104] at least one processor 110; and
[0105] at least one memory 120 including computer program code
[0106] The at least one memory 120 and the computer program code are configured to, with the at least one processor 110, cause the apparatus 100 at least to perform:
[0107] applying equalization to a subset 24 {spatial sound sources 201, 202} of the plurality of spatial sound sources 20 included in the sound scene 10 to modify the sound scene 10, wherein each spatial sound source 20 is associated with a position 22 in the sound scene 10, wherein the equalization comprises a frequency-dependent level adaptation, and wherein the subset 24 comprises a plurality of spatial sound sources 20 but not all of the plurality of spatial sound sources 20.
[0108] As Figure 7As illustrated in the middle, the computer program 122 can arrive at the apparatus 100 via any suitable delivery mechanism 124. The delivery mechanisms 124 can include, for example, a machine readable medium, a computer readable medium, a non-transitory computer readable storage medium, a computer program product, a memory device, a record medium such as a compact disc read only memory (CD-ROM) or digital versatile disc (DVD), or a solid state memory, an article of manufacture that comprises or tangibly embodies the computer program 122. The delivery mechanisms can be, for example, a signal that can be a propagated signal, for example, a propagated signal generated by a compositional apparatus, modulated onto a carrier wave using a variety of techniques known to those skilled in the art, including but not limited to, amplitude modulation, frequency modulation, and phase modulation. The propagated signal can be in a variety of forms, including, but not limited to, single sideband, amplitude modulated, spread spectrum, orthogonal frequency division multiplexing (OFDM), and others. The propagated signal can be modulated using a variety of technologies including, but not limited to, those that utilize a plurality of phase values or those that utilize a plurality of amplitudes. The computer program 122 can be transmitted using a delivery mechanism 124 over a communication link. The communication link can be a network, including, but not limited to a telephone network, a local area network, a wide area network, the Internet, or an Intranet, among others. The computer program 122 can be transmitted using the delivery mechanisms 124, for example, a signal that can be a propagated signal, for example, a propagated signal generated by a compositional apparatus, modulated onto a carrier wave using a variety of techniques known to those skilled in the art, including but not limited to, amplitude modulation, frequency modulation, and phase modulation, among others. The propagated signal can be in a variety of forms, including, but not limited to, single sideband, amplitude modulated, spread spectrum, orthogonal frequency division multiplexing (OFDM), and others. The propagated signal can be modulated using a variety of technologies including, but not limited to, those that utilize a plurality of phase values or those that utilize a plurality of amplitudes. The apparatus 100 can propagate the computer program 122, for example, as a computer data signal.
[0109] Computer program instructions for causing an apparatus to perform or for performing at least the following:
[0110] applying equalization to a subset 24 {spatial sound sources 201, 202} of the plurality of spatial sound sources 20 included in the sound scene 10 to modify the sound scene 10, wherein each spatial sound source 20 is associated with a position 22 in the sound scene 10, wherein the equalization includes a frequency dependent level adaptation, and wherein the subset 24 includes the plurality of spatial sound sources 20 but not all of the plurality of spatial sound sources 20.
[0111] The computer program instructions can be included in a computer program, a non-transitory computer-readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program.
[0112] Although the memory 120 is illustrated as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable and / or can provide persistent / semi-persistent / dynamic / cached storage.
[0113] Although the processor 110 is illustrated as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable. The processor 110 can be a single core or multicore processor.
[0114] References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc., or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as FPGAs, ASICs, signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a field programmable gate array (FPGA), or hardware description language (HDL) descriptions for a programmable logic device (PLD), programmed memory etc.
[0115] As used in this application, the term “circuitry” can refer to one or more or all of the following:
[0116] (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and
[0117] (b) combinations of hardware circuits and software, such as (as applicable):
[0118] (i) combinations of analog and / or digital hardware circuits with software / firmware and
[0119] (ii) any portions of hardware processor(s) with software (including digital signal processors); software, such as “firmware”, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and
[0120] (c) hardware circuit(s) that requires software (e.g., firmware) for operation, but it is not present when it is not needed for the operation.
[0121] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also encompasses an implementation that is a processor that is programmed to perform functions or an implementation that is a processor working together with an application-specific integrated circuit (ASIC) that is programmed to perform functions or both or other implementations in silicon or silicon-based hardware. As examples of further implementations, the term circuitry also encompasses a processor or processors working together with software and / or firmware that are programmed to perform functions, and hardware only implementations (e.g., non- programmable logic and / or analog circuitry) that perform the functions.
[0122] Figure 2The blocks in the middle illustration can represent steps in a method and / or portions of code in a computer program 122. The specific order of the blocks as illustrated is not necessarily a required order of the blocks, but the order and arrangement of the blocks can vary. Also, some blocks can be omitted.
[0123] The apparatus 100 can be configured for capturing and / or processing and / or rendering a sound scene 10.
[0124] As Figure 8A As illustrated in the middle, the apparatus 100 can be configured as a spatial sound capturing device, such as a camera phone. As Figure 8B As illustrated in the middle, the apparatus 100 can be configured as a headphone 103. As Figure 8C As illustrated in the middle, the apparatus 100 can be configured as a controller 105 for a plurality of loudspeakers 107.
[0125] In an alternative implementation, the block 302 of the method 300 comprises:
[0126] applying an equalization to at least one of the plurality of spatial sound sources 20 included in the sound scene 10 to modify the sound scene 10,
[0127] wherein each spatial sound source 20 is associated with a position 22 in the sound scene 10; wherein the equalization comprises a frequency dependent level filter; and the block 304 comprises spatially redistributing the sound energy to spatially compensate the equalization of the at least one of the plurality of spatial sound sources.
[0128] In one implementation, applying the equalization to at least one of the plurality of spatial sound sources 20 comprises applying the equalization to all of the plurality of spatial sound sources 20.
[0129] In another implementation, applying the equalization to at least one of the plurality of spatial sound sources 20 comprises applying the equalization to only one of the plurality of spatial sound sources 20.
[0130] In another implementation, applying the equalization to at least one of the plurality of spatial sound sources 20 comprises applying the equalization to a subset 24 of the plurality of spatial sound sources 20, wherein the subset 24 comprises the plurality of spatial sound sources 20 but does not comprise all of the plurality of spatial sound sources 20.
[0131] sound source
[0132] The spatial sound sources 20 can be defined, for example, using channel-based audio, such as n.m surround sound (e.g. 5.1, 7.1 or 22.2 surround sound) or binaural audio, or scene-based audio, including spatial information about the sound field and sound sources.
[0133] The audio content can encode spatial audio as audio sources. Examples include, but are not limited to, MPEG-4 and MPEG SAOC. MPEG SAOC is an example of metadata- assisted spatial audio.
[0134] The audio content can encode spatial audio as audio sources in the form of moving virtual loudspeakers.
[0135] The audio content can encode spatial audio including sound sources as an audio signal with parametric side information or metadata. The audio signal can be, for example, a first order ambisonics (FOA) or its special case B-format, a higher order ambisonics (HOA) signal or mid-side stereo. For such audio signals, a synthesis using the audio signal and the parametric metadata is used to synthesize the audio scene, thereby creating the desired spatial perception.
[0136] The encoded audio content can be speech and / or music and / or general audio.
[0137] The 3GPP IVAS (3GPP, Immersive Voice and Audio Services) currently under development is expected to support new immersive voice and audio services, such as mediated reality.
[0138] A spatial sound source can be repositioned by mixing a direct form of the sound source (direct sound attenuated and directionally filtered) with an indirect form of the sound source (e.g., directional early reflections and / or diffuse reverberation positioned).
[0139] A spatial sound source can be widened or narrowed by spreading the time- frequency tile representing the spatial sound source over a wider or narrower area. The time- frequency tile can be positioned or repositioned as described in the preceding paragraph for a spatial sound source.
[0140] A spatial sound source can be defined as a sound object. A sound object is a data structure with metadata defining a position.
[0141] Where structural features have been described, the structural features can be replaced with means for performing one or more of the functions of the structural features, whether the function or functions are explicitly or implicitly described.
[0142] The recording of data can comprise only temporary recording, or it can comprise permanent recording, or it can comprise both temporary recording and permanent recording. Temporary recording means that the data is temporarily recorded. This can for example occur during sensing or image capturing, at dynamic memory, at buffers such as circular buffers, registers, caches. Permanent recording means that the data is in the form of an addressable data structure that is retrievable from an addressable memory space, and thus can be stored and retrieved until deleted or overwritten, although long-term storage can or can not occur. The use of the term "capturing" relates to temporary recording of data. The use of the term "capturing" in relation to images relates to permanent recording of data. Where the term "capturing" is used in the specification, it can be replaced by "recording", and vice versa.
[0143] The above examples can apply to the following components:
[0144] Automotive systems; telecommunication systems; electronic systems including consumer electronics; distributed computing systems; media systems for generating or rendering media content, including audio, visual and audiovisual content and hybrids, mediated, virtual and / or augmented reality; personal systems, including personal health systems or personal fitness systems; navigation systems; user interfaces, also known as human-machine interfaces; networks including cellular, non-cellular and optical networks; ad hoc networks; the Internet; the Internet of Things; virtualized networks; and related software and services.
[0145] The term "comprising" is used herein in the inclusive sense of "including" and not the exclusive sense of "consisting of". That is, any reference to X comprising Y indicates that X can consist of only Y, or X can include more than Y. If it is intended that "comprising" be used in the exclusive sense of "consisting only of", then the context will make this clear, by referring to "comprising only of" or by using "consisting of".
[0146] In this specification, reference can be made to various examples. The description of features or functions in relation to one example indicates that those features or functions are present in that example. The use of the term "example" or "for example" or "may" or "can" in the text denotes a description of a possible example of that feature or functionality, but not necessarily a description of the only example of that feature or functionality. Thus "example", "for example", "may" or "can" denotes a description of one example in a class of examples. A property of an example can be a property of only that example, or a property of the class that includes that example, or a property of a sub-class that includes that example, that includes some but not all examples of the class. Thus, the disclosure implicitly discloses that a feature described with reference to one example can also apply to another example, unless the context explicitly indicates otherwise. Thus, the disclosure implicitly discloses distinct examples comprising a feature, unless the context explicitly indicates otherwise. Although each feature can be described in the context of a single example, the feature can be combined with other features of that example, or other examples described herein, although the context of specific examples is not intended to limit the scope of the application to the combination of features described in that specific example, but is instead intended to stretch to the full scope of the application as claimed.
[0147] While embodiments have been described above with reference to various examples, it should be understood that many modifications and variations of the examples described herein can be made.
[0148] Features described in the preceding description can be used in combinations other than the combinations explicitly described above.
[0149] Although functions have been described with reference to certain features, those functions can be performed by other features whether or not the other features are described.
[0150] Although features have been described with reference to certain embodiments, those features can exist in other embodiments whether described or not.
[0151] The terms "a" or "the" are used herein in an inclusive, not an exclusive, sense. That is, unless otherwise noted, any reference to X using "a" or "the" indicates that X can include one or more of the referenced item. If it is intended to use "a" or "the" in an exclusive sense, it will be explicitly stated as such in the context. In some cases, the use of "at least one" or "one or more" can be used to emphasize the inclusive meaning of "a" or "the." The absence of these terms should not be taken to imply that the use of "a" or "the" is in an exclusive sense.
[0152] The presence of a feature (or combination of features) in a claim has the meaning attributed thereto in patent law irrespective of whether or not the feature (or combination of features) is recited in the summary. The term "consisting of is intended to mean "hearing only that specific integer of claim limitation and nothing else. The term "consisting essentially of means that the claim limitation excludes any additional feature, step, ingredient, component or group thereof that is not specified in the claim. The term "comprising" as used herein is intended to mean that the claim can include other elements, ingredients or steps in addition to those specifically recited in the claim.
[0153] In this specification, the use of various examples of adjectives or adjective phrases to describe features of examples. Such description of features relating to examples indicates that the feature is present in some examples exactly as described, and in other examples substantially as described.
[0154] While efforts have been made to accurately rete to those features that are considered important, it should be understood that the applicant can seek protection for any patentable feature or combination of features described herein, whether or not it is expressly described in the claims.
Claims
1. An apparatus (100) for selecting spatial sound sources, comprising means for: applying equalization to a subset (24) of a plurality of spatial sound sources (20) comprised in a sound scene (10) to modify the sound scene (10), wherein spatial sound sources (20) are associated with respective positions (22) in the sound scene (10), wherein equalization comprises frequency-dependent level adaptation, wherein the subset (24) comprises a plurality of spatial sound sources (20), but not all of the plurality of spatial sound sources (20), and the subset (24) is selected from the plurality of spatial sound sources (20) depending on a user input, wherein the user input directly indicates spatial sound sources (20) having a first property by selecting a representation of the spatial sound sources (20), or indirectly indicates spatial sound sources (20) having a first property by selecting positions (22) in the sound scene (10) associated with the spatial sound sources (20), and wherein the first property is that the spatial sound sources have a frequency volume greater than a threshold, and wherein the apparatus (100) comprises means for selecting the subset (24) of the plurality of spatial sound sources (20) to have the first property.
2. The apparatus according to claim 1, wherein the means for applying equalization to the subset of the plurality of spatial sound sources is configured to apply a common equalization to the subset of spatial sound sources.
3. The apparatus according to any of the preceding claims, wherein the user input comprises: an indication of spatial sound sources or positions; a frequency indication; and optionally, an indication of emphasis or deemphasis.
4. The apparatus of claim 1 or 2, comprising: means for spatially redistributing sound energy to spatially compensate the equalization of the subset of spatial sound sources.
5. The apparatus of claim 1 or 2, comprising: means for adapting one or more properties of a target spatial sound source, and / or means for adapting one or more properties of one or more spatial sound sources of the plurality of spatial sound sources not comprised in the subset of the plurality of spatial sound sources.
6. The apparatus of claim 1 or 2, comprising: means for adapting one or more spatial sound sources of the subset of the plurality of spatial sound sources to increase their diffusion.
7. The apparatus of claim 1 or 2, comprising: means for adapting one or more spatial sound sources of the plurality of spatial sound sources not in the subset of the plurality of spatial sound sources to decrease their diffusion, and / or to change their position.
8. The apparatus according to claim 1 or 2, comprising means for representing spatial sound sources as respective spatially localized time-frequency tiles, and means for: preventing a time-frequency tile from becoming spatially separated or significantly different from other contemporaneous time-frequency tiles of the same spatial sound source; and / or preventing a significantly differently localized time-frequency tile from spatially overlapping other significantly differently localized contemporaneous time-frequency tiles of other spatial sound sources.
9. The apparatus of claim 1 or 2, comprising means for capturing and / or processing and / or rendering the sound scene comprising the plurality of spatial sound sources, wherein a spatial sound source is associated with a respective position in the sound scene.
10. The apparatus of claim 1 or 2, configured as a headphone, a controller for a loudspeaker, or a spatial sound capturing device.
11. A method (300) for selecting spatial sound sources, comprising: applying equalization to a subset (24) of a plurality of spatial sound sources (20) comprised in a sound scene (10) to modify the sound scene (10), wherein a spatial sound source (20) is associated with a respective position (22) in the sound scene (10), wherein equalization comprises frequency-dependent level adaptation, wherein the subset (24) comprises a plurality of spatial sound sources (20), but not all of the plurality of spatial sound sources (20), and the subset (24) is selected from the plurality of spatial sound sources (20) depending on a user input, wherein the user input directly indicates a spatial sound source (20) having a first property by selecting a representation of the spatial sound source (20), or indirectly indicates a spatial sound source (20) having a first property by selecting a position (22) in the sound scene (10) associated with the spatial sound source (20), and wherein the first property is that the spatial sound source has a frequency volume greater than a threshold value, and selecting the subset (24) of the plurality of spatial sound sources (20) to have the first property.
12. A computer product which, when running on a computer, causes: applying equalization to a subset (24) of a plurality of spatial sound sources comprised in a sound scene (10) to modify the sound scene (10), wherein a spatial sound source (20) is associated with a respective position (22) in the sound scene (10), wherein equalization comprises frequency-dependent level adaptation, wherein the subset (24) comprises a plurality of spatial sound sources (20), but not all of the plurality of spatial sound sources (20), and the subset (24) is selected from the plurality of spatial sound sources (20) depending on a user input, wherein the user input directly indicates a spatial sound source (20) having a first property by selecting a representation of the spatial sound source (20), or indirectly indicates a spatial sound source (20) having a first property by selecting a position (22) in the sound scene (10) associated with the spatial sound source (20), and wherein the first property is that the spatial sound source has a frequency volume greater than a threshold value, and selecting the subset (24) of the plurality of spatial sound sources (20) to have the first property.
Citation Information
Patent Citations
Spatial audio processing emphasizing sound sources close to focal distance
CN109076306A
Playback Device For Generating Sound Events
US20100223552A1