Method and apparatus for efficient audio rendering
By determining the rendering control value based on the parameters of perceptual correlation, eliminating, clustering or converting audio sources, the problem of high computing costs in high-quality immersive audio rendering is solved, and efficient rendering is achieved without reducing audio quality.
Patent Information
- Application Number
- CN202380084899.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-20
- Filing Date
- 2023-12-12
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art has high computing costs in high-quality immersive audio rendering, resulting in a degradation in output audio quality and affecting the listening experience.
Determine rendering control values based on parameters of perceptual correlation, culling, clustering, or converting audio sources to reduce rendering complexity and maintain audio output quality.
While reducing the computing load, maintain or improve the perceived quality of audio output and reduce computing and memory system workloads.
Smart Images

Figure CN120359766A_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority benefits of U.S. Provisional Application Serial No. 63 / 431,822, filed on December 12, 2022, and U.S. Provisional Application Serial No. 63 / 491,258, filed on March 20, 2023, and each of these applications is incorporated herein by reference in its entirety.
[0002] The present disclosure generally relates to a method for rendering audio sources. In particular, the present disclosure relates to controlling the rendering of audio sources to achieve efficient audio rendering.
[0003] Although some embodiments will be described herein with specific reference to this disclosure, it should be understood that the present disclosure is not limited to such fields of use and can be applied to a broader context. Background Art
[0004] Any discussion of background art throughout the disclosure should in no way be construed as an admission that such technology is well known in the art or forms part of the common general knowledge in the art.
[0005] Audio rendering, especially for high-quality immersive content, is computationally expensive. To render immersive audio within a reasonable computation time (especially on power-constrained devices), only numerical operations with very limited complexity are allowed on the processors included in these devices. As a result, the quality of the output audio typically degrades, and the listener experience is poor.
[0006] Therefore, there is a current need for methods and apparatuses that allow for efficient rendering of audio sources without degrading the perceived quality of the resulting rendered audio output. Summary of the Invention
[0007] According to a first aspect of the present disclosure, a method for rendering an audio source is provided. The method may include, for each of a plurality of audio sources, determining a rendering control value based on one or more parameters indicating the perceived relevance of the corresponding audio source. The method may further include, for each of the plurality of audio sources, comparing the rendering control value with a perception threshold to determine whether the corresponding audio source meets a perceived irrelevance criterion. The method may further include modifying the audio sources among the plurality of audio sources for which the rendering control value meets the threshold to obtain modified audio sources. And the method may include rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio sources.
[0008] In some embodiments, the one or more parameters may include the loudness of the corresponding audio source. Then, the rendering control value may be determined based on the loudness value.
[0009] In some embodiments, the one or more parameters can include the relationship between the loudness of the single audio source and the loudness of the audio output produced by rendering the plurality of audio sources other than the single audio source.
[0010] In some embodiments, the one or more parameters can include the relationship between the short-term spectral energy of the single audio source and the spectral energy of the audio output produced by rendering the plurality of audio sources.
[0011] In some embodiments, the rendered audio output can be associated with a psychoacoustic masking model.
[0012] In some embodiments, modifying the audio source for which the rendering control value meets the threshold can include excluding the audio source. Then, the modified audio source can correspond to an excluded subset of the plurality of audio sources.
[0013] In some embodiments, rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio source can include not rendering the audio sources included in the excluded subset.
[0014] In some embodiments, the one or more parameters can include the positional proximity of the audio source relative to the listener's position.
[0015] In some embodiments, the rendering control value can be determined based on the distance from the corresponding audio source position of the audio source for which the rendering control value is determined to another audio source, and wherein the rendering control value can be compared with a fraction of the distance between the listener's position and the nearest source position.
[0016] In some embodiments, the one or more parameters can include the directional proximity relative to the listener's position.
[0017] In some embodiments, the rendering control value can be determined based on at least one of the azimuth angle or the elevation angle between the corresponding audio source position of the audio source for which the rendering control value is determined and the listener's position.
[0018] In some embodiments, modifying the audio source can include clustering the audio sources for which the rendering control value meets the threshold. Then, the modified audio source can correspond to the clustered audio sources.
[0019] In some embodiments, rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio source can include replacing the clustered audio sources with a smaller number of audio sources and rendering the smaller number of audio sources as the modified audio source.
[0020] In some embodiments, the one or more parameters may include one or more of acoustic occlusion, optical occlusion, or abstraction. Then, the rendering control value may be determined based on the degree of occlusion or abstraction of the corresponding audio source.
[0021] In some embodiments, modifying the audio source may include converting the type of the audio source for which the rendering control value meets the threshold. Then, the modified audio source may correspond to the audio source after type conversion.
[0022] In some embodiments, rendering the unmodified audio source among the plurality of audio sources relative to the modified audio source may include rendering the audio source after type conversion as the modified audio source.
[0023] In some embodiments, for an audio source having a directivity pattern among the plurality of audio sources, if the rendering control value meets the threshold, the directivity calculation may not be performed.
[0024] In some embodiments, the method may further include receiving control information indicating whether to perform modification on some or all of the audio sources for which the rendering control value meets the threshold.
[0025] In some embodiments, the control information may be one of an activation parameter or a deactivation parameter, grouping data, priority sorting information, and / or visibility information.
[0026] According to a second aspect of the present disclosure, there is provided an apparatus for rendering an audio source. The apparatus may include one or more processors configured to implement a method including the following operations: for each audio source among a plurality of audio sources, determining a rendering control value based on one or more parameters indicating the perceived relevance of the corresponding audio source; for each audio source among the plurality of audio sources, comparing the rendering control value with a perceivedness threshold to determine whether the corresponding audio source meets the perceived irrelevance criterion; modifying the audio source for which the rendering control value meets the threshold among the plurality of audio sources to obtain a modified audio source; and rendering the unmodified audio source among the plurality of audio sources relative to the modified audio source.
[0027] According to a third aspect of the present disclosure, there is provided an apparatus including a processor and a memory, the memory being coupled to the processor and storing instructions for the processor, wherein the processor is adapted to execute the method described herein.
[0028] According to a fourth aspect of the present disclosure, there is provided a program, the program including instructions which, when executed by a processor, cause the processor to perform the methods described herein. A computer-readable storage medium may store the program.
[0029] It should be appreciated that apparatus (system) features and method steps may be interchanged in various ways. In particular, as will be appreciated by those skilled in the art, the details of the disclosed methods may be implemented by a corresponding apparatus (system), and vice versa. In addition, any statement made above regarding the methods is understood to equally apply to the corresponding apparatus (system), and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Example embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0031] Figure 1 An example of a method for rendering an audio source according to an embodiment of the present disclosure is illustrated.
[0032] Figure 2 An example of culling an audio source according to an embodiment of the present disclosure is schematically illustrated.
[0033] Figure 3 An example of clustering audio sources according to an embodiment of the present disclosure is schematically illustrated.
[0034] Figure 4 Another example of clustering audio sources according to an embodiment of the present disclosure is schematically illustrated.
[0035] Figure 5 An example of converting the type of an audio source according to an embodiment of the present disclosure is schematically illustrated.
[0036] Figure 6 An example of a method for processing audio according to an embodiment of the present disclosure is illustrated.
[0037] Figure 7 Another example of a method for processing audio according to an embodiment of the present disclosure is illustrated.
[0038] Figure 8 Yet another example of a method for processing audio according to an embodiment of the present disclosure is illustrated.
[0039] Figure 9 An example of using control information according to an embodiment of the present disclosure is illustrated.
[0040] Figure 10 An example of control information parameters according to an embodiment of the present disclosure is schematically illustrated.
[0041] Figure 11Illustrated is an example of an apparatus including a processor and a memory coupled to the processor according to an embodiment of the present disclosure.
[0042] Note that where feasible, like or similar reference numerals may be used in the figures, and such reference numerals may indicate like or similar functionality. The figures depict embodiments of the disclosed apparatus (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.
[0043] In addition, in figures where connecting elements (such as solid lines or dashed lines or arrows) are used to illustrate a connection, relationship, or association between two or more other schematic elements, the absence of any such connecting element does not imply that a connection, relationship, or association may not exist. In other words, some connections, relationships, or associations between elements are not shown in the figures so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, in the case where a connecting element represents the communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or more signal paths that may be required to effect the communication. Detailed Description
[0044] Overview
[0045] The methods and apparatus described herein provide technical benefits and advantages of reducing the computational and memory system workload without degrading the perceived quality of the resulting rendered audio output. The lack of degradation in perceived quality can be understood as providing a similar (or better) listener subjective experience. For example, compared to an unmodified audio output, the present disclosure provides a method of modifying the audio output to achieve such a benefit. Additionally, the present disclosure provides the additional benefit of reducing the workload. Such reduction can be achieved by modifying and / or deactivating selected audio source digital signal processing (DSP) instances and optionally associated metadata processing.
[0046] The parameters, downmix matrix, and / or related information of the methods described herein may be defined and signaled by an encoder, decoder / renderer, and / or application. For example, the encoder may operate in an "encoder assist" mode, in which the encoder may manually allow content creators to influence an otherwise automated method. This can provide the benefit of avoiding the senseless culling described below. Alternatively, such information may be provided in the decoder and / or renderer (e.g., automatically provided in a "default" mode). Alternatively, the information may be provided by an application (e.g., automatically provided by an application or system middleware in a "system assist" mode).
[0047] Method for Rendering Audio Sources
[0048] The present disclosure relates to reducing the number of rendered audio sources to eliminate complexity on a renderer without incurring psychoacoustic effects. Such audio sources can be any type of audio input source, such as channel (static) sources, (multiple) audio objects, (multiple) higher-order high-fidelity stereo HOA, first-order high-fidelity stereo FOA, and / or B-format high-fidelity stereo.
[0049] Reference Figure 1 illustrates method 100 for rendering audio.
[0050] In step S101 of the method, for each of the multiple audio sources, a rendering control value is determined based on one or more parameters indicating the perceived correlation of the corresponding audio source.
[0051] In step S102 of the method, for each of the multiple audio sources, the rendering control value is compared with a perception threshold to determine whether the corresponding audio source meets the perceived irrelevance criterion.
[0052] In step S103 the audio sources among the multiple audio sources whose rendering control values meet the threshold are modified to obtain modified audio sources.
[0053] And in step S104 the unmodified audio sources among the multiple audio sources are rendered relative to the modified audio sources.
[0054] In an embodiment, one or more parameters (for determining the rendering control value) may include the loudness of the corresponding (single) audio source. Then, the rendering control value can be determined based on the corresponding loudness value. Depending on the use case, the loudness value can be obtained from:
[0055] 1) Metadata from existing audio standards set by the MPEG Audio Group, which audio standards include MPEG-H 3D Audio (ISO_IEC_23008-3) or MPEG-D Part 4 (Dynamic Range Control). For example, the metadata may be presented for legacy content;
[0056] 2) Information of an encoder set by the MPEG Audio Group, such as an encoder compatible with the MPEG-H 3D Audio or MPEG-I audio standards. For example, the loudness information can be estimated and transmitted via the MPEG-H audio stream (MHAS) packet payload;
[0057] 3) Information provided by an MPEG-I renderer. For example, information estimated in real time using audio content and rendering gain.
[0058] The above loudness data sources can be applied alone or in combination; for example, for content without reliable loudness information or with unreliable loudness information, social virtual reality (VR) content, etc.
[0059] Generally, according to different loudness definitions and measurement methods (e.g., short-term loudness, instantaneous loudness, level gating, IBU definition, etc.), loudness can be signaled differently and processed differently.
[0060] Alternatively or additionally, in an embodiment, one or more parameters can include the relationship between the loudness of a single audio source and the loudness of an audio output generated by rendering a plurality of audio sources other than the single audio source. For example, in a non-limiting manner, this relationship can be a signal-to-noise ratio. Other measures that can be conceived, such as using spatial masking, are also possible.
[0061] Alternatively or additionally, in an embodiment, one or more parameters can include the relationship between the short-term spectral energy of a single audio source and the spectral energy of an audio output generated by rendering a plurality of audio sources. The rendered audio output can be associated with a psychoacoustic masking model.
[0062] The above parameters can be used alone or in combination to determine the corresponding rendering control values for each of the plurality of audio sources. Then, it can be said that the rendering control value indicates the perceived relevance (perceivedness) of the corresponding audio source. Then, comparing the rendering control value with a (predetermined) perceivedness threshold for each of the plurality of audio sources can indicate whether the corresponding audio source is perceived as relevant or not, that is, indicate whether the corresponding audio source meets the perceived irrelevance criterion. For example, if the rendering control value is equal to or higher than a certain threshold, the corresponding audio source can be considered to be perceived as relevant. If the rendering control value is lower than a certain threshold, the corresponding audio source can be considered to be perceived as irrelevant, that is, the corresponding audio source meets / complies with the perceived irrelevance criterion.
[0063] For example, if the listener is about to start experiencing heavy rain noise, the audio source associated with the bird chirping audio can become perceived as irrelevant. In this case, the rendering control value will meet the corresponding (perceived relevance) threshold, and the corresponding audio source will meet the perceived irrelevance criterion.
[0064] In an embodiment, the modification of an audio source whose rendering control value meets the threshold can include excluding the audio source. Then, the modified audio source can correspond to an excluded subset of the plurality of audio sources. Referring to Figure 2 the example of , the exclusion of the audio source is schematically illustrated by forming the corresponding excluded subset. Figure 2Illustrated are a plurality of audio sources 201, 202 associated with corresponding audio scenes 200. Referring again to the above example, audio source 202 may be associated with bird song, and audio source 201 may be associated with light rain. If the listener now starts experiencing heavy rain noise 201, the audio source 202 associated with bird song may become perceptually irrelevant, i.e., the rendering control value may meet the perceptibility threshold. In this case, the audio source 202 is modified because the audio source can be culled. Then, the culled audio source may correspond to a corresponding culling subset 203.
[0065] In this context, it can be said that culling refers to removing the corresponding perceptually irrelevant audio source. That is, in an embodiment, rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio sources (whose rendering control values may not meet the perceptibility threshold) may include not rendering the audio sources included in the culling subset.
[0066] When estimating perceptual relevance considering additional perceptual energy / loudness related aspects and setting corresponding rendering control conditions based on comparing the corresponding rendering control values with a perceptibility (perceptual relevance) threshold to determine whether a corresponding audio source meets the perceptually irrelevant criterion, an improvement is achieved compared to traditional methods for detecting audio source irrelevance, because traditional methods only consider the distance (from the listener to the audio object) and / or the rendering gain (of the audio object). In other words, it can be said that traditional methods apply blind culling because the actual signal and signal energy are not considered. Therefore, traditional methods cannot detect sub - categories of perceptually irrelevant audio sources. That is, sources that are inaudible but have corresponding distance and gain values higher than the traditional culling threshold. Thus, in traditional methods, audio sources that should be retained may be removed and vice versa.
[0067] In an embodiment, one or more parameters (for determining the rendering control value) may alternatively or additionally include the positional proximity (near - source position) of the audio source relative to the listener's position. Then, the rendering control value may be determined based on the distance from the corresponding audio source position of the audio source for which the rendering control value is determined to another audio source, and the rendering control value may be compared to a fraction of the distance between the listener's position and the nearest source position. For example, the distance between N audio source positions may be less than T distance % of the distance between the listener and the nearest source. T distance % refers to a variable representing a percentage value (e.g., 5%). When the listener explores a scene containing several audio sources (i.e., the original N audio sources), the distance between the listener and the nearest source can be calculated. This distance to the nearest source can be referred to as the variable Tc. Additionally, the distances between each of the remaining audio sources in the scene can also be calculated. Each of these distances is compared to Tc. If the distance between two sources is less than T of Tcdistance %, these source clusters (grouped together). Since this comparison is done over all audio sources, the clusters can add or remove more members (audio sources).
[0068] Figure 3 Illustrates an example of an abstract visualization of clustering based on location proximity. As in Figure 3 it can be seen, cluster A 301 contains audio sources 301A that are closest to each other in terms of distance. Cluster B 302 contains different sources 302B that are also closest to each other in terms of distance. In contrast, source X 303 is far from any of the audio sources 301A, 302B in both cluster A 301 and B 302, and is thus not assigned to either cluster.
[0069] Alternatively or additionally, one or more parameters can include direction proximity (near source azimuth and elevation) relative to the listener's position. Then, the rendering control value can be determined based on at least one of the azimuth or elevation between the corresponding audio source position of the audio source for which the rendering control value is determined and the listener's position. For example, the azimuth and elevation related to the listener's position for N audio sources can be less than T azimuth degrees and T elevation degrees. Figure 4 Illustrates an example of an abstract visualization of clustering based on direction proximity and / or location proximity / distance proximity. As in Figure 4 it can be seen, assume that the angular / directional differences in azimuth and elevation between any two sources (e.g., A2 and An) seen from listener L are less than T azimuth degrees and T elevation degrees. Then these sources are clustered together. T azimuth and T elevation thresholds in degrees can be, for example, 5 degrees. It can also be assumed that since the distances between these audio sources (A, A1, A2, A3, and An) are small, they are clustered based on location proximity or distance proximity.
[0070] The above parameters can be used alone or in combination to determine the respective rendering control values for each of the multiple audio sources. Then, it can be said that the rendering control values indicate the perceived relevance of the respective audio sources. Then, comparing the rendering control values with a (predetermined) perceivedness threshold for each of the multiple audio sources can indicate whether the respective audio source is perceptually relevant, i.e., indicate whether the respective audio source meets the perceptual irrelevance criterion. For example, if the rendering control value is equal to or higher than a certain threshold, the respective audio source can be considered perceptually relevant. If the rendering control value is lower than a certain threshold, the respective audio source can be considered perceptually irrelevant, i.e., the respective audio source meets / conforms to the perceptual irrelevance criterion. As an alternative or supplement to the above loudness-related parameters, the above parameters refer to the aspects related to perceptual localization used to estimate perceptual relevance / irrelevance.
[0071] In an embodiment, modifying the audio sources may include clustering the audio sources whose rendering control values meet the perceivedness threshold. Then, the modified audio sources may correspond to the clustered audio sources. In an embodiment, rendering the unmodified audio sources among the multiple audio sources with respect to the modified audio sources (whose rendering control values may not meet the perceivedness threshold) may include replacing the clustered audio sources with a smaller number of audio sources, and rendering the smaller number of audio sources as the modified audio sources. For example, N audio sources may be replaced by one or more audio sources M (M < N, preferably M = 1) including (multiple) weighted downmix signals or multichannel audio signals. That is, assuming that a total of N audio sources are clustered, these sources may be replaced by M audio sources. M may be the number of modified audio sources to be rendered, rather than N. Thus, the M audio sources still represent the same clustering.
[0072] By clustering the audio sources whose rendering control values meet the threshold, advantageously, a number of audio sources can be replaced with a smaller number of modified audio sources, thereby reducing the number of audio sources to be rendered.
[0073] Alternatively or additionally, in an embodiment, one or more parameters may include one or more of acoustic occlusion, optical occlusion, or abstraction. Then, the rendering control value can be determined based on the degree of occlusion or abstraction of the respective audio source.
[0074] In an embodiment, modifying the audio sources may include converting the type of the audio sources whose rendering control values meet the threshold. That is, modifying the audio sources may include converting the respective audio source from one type to another type. The resulting type can be rendered with higher computational efficiency. This is schematically illustrated in the Figure 5 example. Figure 5Illustrated are a plurality of audio sources 501 associated with corresponding audio scenes 500. For example, if the distances 502 between a listener L and a number of audio sources 501 are large enough, these audio sources 501 (e.g., waterfall range sounds) can be replaced by one audio object 503 (e.g., distant waterfall point source); or many audio point sources (e.g., raindrops) can be replaced by one high-order high-fidelity stereo HOA (e.g., rain noise) signal.
[0075] Then, the modified audio source can correspond to a transcoded audio source. The parameters for determining the corresponding rendering control conditions (rendering control values compared to thresholds) for transcoding can depend in particular on the positional proximity and / or directional proximity and / or degree of acoustic / optical occlusion / abstraction as described above. For example, if a number of audio sources are occluded but still acoustically important, these sources can simply be replaced by a single point-source audio or FOA signal.
[0076] Then, rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio source (whose rendering control values may not satisfy the perception threshold) can include rendering the transcoded audio source as the modified audio source.
[0077] In addition to the above, in an embodiment, on an audio source having a directivity pattern among a plurality of audio sources, if the rendering control value satisfies the threshold, the directivity calculation may not be performed.
[0078] An example can be the directional motor engine sound of a car. Suppose the car is traveling on a circular track and the listener is standing at a specific position / platform, as in a racing scene. Then, when the car passes by the listener, the effect of the directional motor engine sound will be perceived. But when the car is at a distant position, the directivity effect is negligible because the listener can only hear the faint sound of the motor engine, making it perception-irrelevant and thus there is no need to perform the directivity calculation. In some embodiments, if there is a significant obstacle / occlusion (such as a hill or a pile of hard rocks) in the middle of the circular track, the source can even be further considered to be included in the culling subset.
[0079] In an embodiment, the directivity calculation can be skipped by additionally considering the so-called worst-case directivity gain estimation parameter. This worst-case estimation is easy to calculate, for example, simply by taking the maximum directivity gain within the given directivity pattern of the audio source. This estimation can be used in combination with other existing rendering stage thresholds (e.g., 60 dB culling threshold) such that if the audio source rendering gain satisfies the "culling" stage gain threshold, the directivity calculation / stage can be skipped.
[0080] In an embodiment, the method may further include receiving control information that indicates whether to perform a modification on some or all of the audio sources for which the rendering control value meets a threshold.
[0081] The following parameters may be used to control the application of the perceptual audio culling, clustering, and type conversion rendering phases. The (de)activation parameters may be included in the control information. For example, these parameters may allow disabling the modification of the voice-over speech signal even if the corresponding audio source may have a rendering control value that meets the threshold (i.e., meets the perceptual irrelevance criterion). Alternatively or additionally, the grouped data for culling / clustering / type conversion may be included in the control information. Such grouped data may include, for example, a way to define a group of audio sources within a specific region (e.g., the sound-emitting components of a car) or having a specific signal category (e.g., wind and rain environmental sounds). Alternatively or additionally, for each group, the encoder may define culling and clustering strategies to combine or remove objects that may be included in the control information. For example:
[0082] o If all the objects in the group are below a certain threshold loudness, the group may be combined into one object (loudness-based culling);
[0083] o If the azimuth / elevation of all the objects are within a certain range, the group may be combined into one object (direction-based clustering);
[0084] o If one or many objects in the group are acoustically occluded, the group may be combined into one object.
[0085] Alternatively or additionally, a prioritization of the signal perceptual relevance (and its information importance) with respect to the listener (and the listener's focus of attention) may be included in the control information. For example, when performing "cocktail party effect" reproduction, culling (instead of leveling / equalizing) may be applied to non-relevant audio sources to improve speech intelligibility. Alternatively or additionally, the visibility of the graphical representation of the (multiple) audio sources to the listener (i.e., whether the visual representation of the audio source is visible to the user) may be included in the control information. For example, the audio source may be associated with a visual representation that may be:
[0086] - Not in the video rendering viewport (behind the listener);
[0087] - Occluded by an optically opaque occluder.
[0088] Such audio sources may be processed differently depending on their nature and application scenario; for example, the thresholds for location proximity or direction proximity conditions may vary depending on the virtual object visibility.
[0089] The following aspects can be further used to avoid frequent "switching" in the culling, clustering, and type conversion rendering phases:
[0090] - Hysteresis (the dependence of the current decision on the previous decision history);
[0091] - Distance relative to the listener's position (and its change over time - position velocity);
[0092] - Direction relative to the listener's position (and its change over time - angular velocity). For example, a time threshold can be determined based on the measured velocity.
[0093] All parameters of the above aspects and / or the downmix matrix can be defined and signaled in the following ways:
[0094] - Encoder - "encoder assist" mode (performed manually by the content creator) to allow the content creator to influence automatic culling in the renderer (e.g., to avoid meaningless culling);
[0095] - Renderer - "default" mode (automatic);
[0096] - Application - "system assist" mode (automatic) of the application or system middleware.
[0097] It is noted that parameters for all rendering phases can be defined for each (virtual) object. All these rendering phases can be implemented as part of the MPEG-I rendering process or as an MPEG-I audio "external tool". A transition from the "culling" state, "clustering" state, and "type conversion" state to the original state can be initiated (when the corresponding conditions are met), and the (multiple) sound source energy is low. That is, in a further embodiment, the above method can be implemented by a corresponding two-step process to determine the 'correct' frame for applying the application complexity reduction measure. These two steps can include: in the first step, determining the complexity reduction measure to be applied (i.e., the modification of the audio source as described); and in the second step, determining the time point or time period for applying the corresponding measure.
[0098] The above method can be implemented by a corresponding apparatus including one or more processors. Alternatively or additionally, the above method can be implemented in the form of a corresponding program including instructions that, when executed by a processor, cause the processor to execute the method. The program can be stored on a computer-readable storage medium.
[0099] While the above aspects of controlling the rendering of multiple audio sources can be implemented separately or in combination within a single method, the above aspects of complexity reduction can alternatively be implemented by individual methods of processing audio as described below. However, these individual methods can also be implemented separately or in combination.
[0100] Elimination of audio sources
[0101] The method according to the present disclosure involves selecting (i.e., "eliminating") such audio sources.
[0102] The present disclosure relates to a method for eliminating audio sources. The elimination can be performed during the core audio decoding phase or during the audio rendering phase.
[0103] The "elimination" method involves removing perceptually irrelevant audio sources. For example, when a listener starts experiencing heavy rain noise, the elimination process will remove the audio source associated with the bird chirping audio from the rendering pipeline.
[0104] Generally, the following aspects have been considered to detect whether an audio source is irrelevant:
[0105] - Distance (from the listener to the audio object)
[0106] - Rendering gain (of the audio object)
[0107] Problem:
[0108] However, these existing detection techniques have various problems. For example, the prior art cannot detect sub - categories of perceptually irrelevant audio sources. That is, the sub - category contains sources that are inaudible but have corresponding distance (gain) values above / below the elimination threshold. The threshold can be based on distance and gain. For example, if the source distance is greater than the distance threshold D, the source is eliminated. Alternatively, if the source gain value is below the gain threshold G, the source is eliminated. Problematically, under the prior art, some audio sources may end up being classified as "not eliminated", but they may be perceptually irrelevant or inaudible.
[0109] Technical solution:
[0110] The present disclosure relates to performing the elimination of audio sources based on additional information (such as information regarding aspects related to perceptual energy / loudness). This information allows the estimation of perceptual relevance and improves the setting of the "elimination" application conditions.
[0111] One aspect of the present disclosure considers relevant information during elimination. The relevant information is the loudness of a single audio source. The loudness information (i.e., value) can be obtained from:
[0112] 1) Metadata from existing audio standards set by the MPEG Audio group, which includes MPEG - H 3D audio (ISO_IEC_23008 - 3) or MPEG - D Part 4 (dynamic range control). For example, the metadata can be presented for traditional content.
[0113] 2) Information of an encoder set by the MPEG Audio Group, such as an encoder compatible with the MPEG-H 3D Audio or MPEG-I Audio standard. For example, loudness information can be estimated and transmitted via the MPEG-H Audio Stream (MHAS) packet payload.
[0114] 3) Information provided by an MPEG-I renderer. For example, information estimated in real time using audio content and rendering gain.
[0115] Each of these sources can be considered individually or in combination. For example, combinations of different loudness data sources (1), (2), (3) can be used for certain types of applications (e.g., for content with no reliable or unreliable loudness information, social VR audio content).
[0116] Another aspect of the present disclosure contemplates information regarding the relationship between the loudness of a single audio source and the loudness of the overall rendered audio output. Some embodiments can exclude this source, for example, in the case of an SNR-based relationship. Depending on different loudness definitions and measurement methods (e.g., short-term loudness, instantaneous loudness, EBU R128 definition), the loudness information can be signaled differently and processed differently.
[0117] Another aspect of the present disclosure contemplates information regarding the relationship between the short-term spectral energy of a single audio source and the spectral energy of the final rendered audio output associated with a psychoacoustic masking model.
[0118] Figure 6 An exemplary method of excluding an audio source according to the present disclosure is illustrated.
[0119] The method includes a first step S601 . At step S601, for a plurality of audio sources, one or more values and / or relationship data related to perceptual correlation information can be obtained. The audio sources can be received and / or predetermined. The (multiple) values indicate the loudness of each audio source. In one example, the relationship data can indicate the relationship between the loudness of a single audio source and the loudness of the overall rendered audio output (excluding that single source). The single audio source will be one of the plurality of audio sources. In another example, the relationship data will relate to the relationship between the short-term spectral energy of a single audio source and the spectral energy of the final rendered audio output associated with a psychoacoustic masking model. The single audio source will likewise be one of the plurality of audio sources.
[0120] At step At S602, each value / data from S601 will be verified to see if it meets the corresponding predefined condition threshold(s). An example would be the "SNR" value, where a single audio source is considered the "useful signal" and the remaining audio sources are considered "noise". Then the SNR value is compared with the threshold. The threshold can be preset or determined dynamically. An audio source with a value that meets the corresponding condition threshold is considered perceptually irrelevant. Then a subset of the perceptually irrelevant sources from the multiple sources from S601 is selected.
[0121] In step At S603 , audio culling is performed on the perceptually irrelevant audio sources from S602 that meet the condition threshold. Step S603 outputs information related to the status (culled or not culled) of the audio sources. For example, the status of the audio source (which can be "culled" or "not culled") can be represented by a boolean variable indicating "true" or "false". For example, a boolean variable named "isCulled" is specified and initialized to "false". After processing, it can be set to "true" by the assignment operator "isCulled = true". Then, this variable is passed down the processing chain to exclude the rendering of this audio source. In subsequent steps (not shown), this culling information is provided to the audio renderer. The audio renderer uses the culling information in combination with the audio sources to determine which audio sources will be rendered.
[0122] Clustering of audio sources
[0123] This disclosure further relates to performing clustering of audio sources. The "clustering" phase should replace several audio sources with a smaller number of modified sources. Clustering refers to grouping based on certain conditions (e.g., the distance proximity of audio sources such as sources that are close to each other). Clustering can be performed during the core audio decoding phase or during the audio rendering phase.
[0124] Problem:
[0125] The current state-of-the-art solutions provided by the standardized audio group of MPEG do not support reducing the computational and memory system workload by clustering audio sources without degrading the perceived quality of the resulting rendered audio output. However, such a clustering feature is an important feature that needs to be supported in future state-of-the-art solutions for standardizing MPEG audio features.
[0126] Technical solution:
[0127] This disclosure relates to clustering of audio sources, which is based on additional perceptually localization-related aspects for estimating perceptual relevance and setting the "clustering" application conditions.
[0128] Another aspect of the present disclosure is to determine clustering based on position proximity (i.e., near-source position) relative to the listener's position. For example, if the distance between N audio source positions is less than T times the distance between the listener and the nearest source distance %, then these audio sources are replaced by one or more audio sources M, where M < N, preferably M = 1. The M resulting audio sources should contain the (multiple) weighted downmix signals or multichannel audio signals.
[0129] T distance % refers to a variable representing a percentage value (e.g., 5%). When the listener explores a scene containing several audio sources (i.e., the original N audio sources), the distance between the listener and the nearest source can be calculated. This distance to the nearest source can be referred to as the variable Tc. Additionally, the distances between each of the remaining audio sources in the scene can also be calculated. Each of these distances is compared with Tc. If the distance between two sources is less than Tdistance % of Tc, then these sources are clustered (grouped together). Since this comparison is done on all audio sources, this clustering can add more members (audio sources). Assuming a total of N audio sources are clustered, then these sources are replaced by M audio sources. M is the number of audio sources to be rendered, not N. Thus, the M audio sources still represent the same clustering.
[0130] Figure 3 illustrates an example of an abstract visualization of clustering based on position proximity. As can be seen in Figure 3 , cluster A 301 contains audio sources 301A that are closest to each other in terms of distance. Cluster B 302 contains different sources 302B that are also closest to each other in terms of distance. In contrast, source X 303 is far from any of the audio sources 301A, 302B in both cluster A 301 and B 302, and thus is not assigned to either cluster.
[0131] Another aspect of the present disclosure is to determine clustering based on direction proximity (i.e., near-source azimuth and elevation) relative to the listener's position. It should be noted that the listener's perceptual localization ability has a finer resolution in azimuth than in elevation. For example, if the listener-position-related azimuth and elevation of N audio sources are less than T azimuth degrees and T elevation degrees, respectively, then these audio sources are replaced by one or more audio sources M (M < N, preferably M = 1) that contain the (multiple) weighted downmix signals or multichannel audio signals. T azimuth and T elevationare thresholds in degrees for the azimuth and elevation angles respectively representing direction proximity, e.g., 5 degrees. When a listener explores a scene containing several audio sources (i.e., the original N audio sources), the angular / directional difference between any two sources seen from the listener's pose can be calculated (based on the azimuth and elevation angles). If the angular / directional difference between two sources is less than T azimuth and T elevation , then these sources are clustered (grouped together). Since this comparison is done over all audio sources, this clustering can add or remove more members (audio sources).
[0132] Figure 4 illustrates an example of an abstract visualization of clustering based on direction proximity and / or position proximity / distance proximity. As can be seen in Figure 4 , assuming that the angular / directional differences in terms of azimuth and elevation angles between any two sources (e.g., A2 and An) seen from the listener L are less than T azimuth degrees and T elevation degrees respectively. Then these sources are clustered together. It can also be assumed that since these audio sources (A, A1, A2, A3, and An) are very close to each other in distance, they are clustered based on position proximity or distance proximity.
[0133] Figure 7 illustrates an exemplary method for clustering audio sources according to the present disclosure.
[0134] The method includes a first step S701 . At step S701, for a plurality of audio sources, one or more perceptual relevance information values for each audio source can be determined. The perceptual relevance information can be either i) position proximity and / or ii) direction proximity.
[0135] At step At S702 , each value from S701 is verified to determine whether they meet the corresponding predefined condition thresholds. The audio sources having values that meet the corresponding condition thresholds are added to the cluster.
[0136] At step S703 , all sources within the cluster (e.g., all N audio sources) are replaced by a lower number of audio sources (e.g., M, where M < N) for rendering. The replacement can involve a weighted downmix operation on the entire set or a subset of the N audio sources. It can also be obtained by simply culling some sources.
[0137] The output of step S703 can be a set of M audio sources to be rendered. These M audio sources replace the original N audio sources without a perceived quality degradation during rendering.
[0138] Type conversion of audio sources
[0139] This disclosure relates to a "type conversion" stage of converting one or more audio sources from one type to another. The expected resulting type will be rendered with higher computational efficiency. The "type conversion" can be performed during the core audio decoding stage or during the audio rendering stage.
[0140] For example, the "type conversion" method will determine whether the distance between the listener and each of several audio sources is large enough (e.g., by comparing the distance between the listener and the audio source with a threshold). If the distance is large enough, these audio sources (e.g., waterfall range sounds) can be replaced by audio objects / sources (e.g., distant waterfall point sources). Alternatively, these audio point sources (e.g., raindrops) can be replaced by a HOA (e.g., rain noise) signal.
[0141] Problem
[0142] The current art provided by the standardized audio group of MPEG does not support reducing the computational and memory system workload by type-converting audio sources without degrading the perceived quality of the resulting rendered audio output. However, type conversion is an important feature that needs to be supported in future standardized art for standardized MPEG audio features.
[0143] Technical solution:
[0144] This disclosure relates to type-converting audio sources based on one or more of the following conditions:
[0145] - Location proximity and orientation proximity (for more details on the measurement of these parameters, see the clustering discussion above)
[0146] - Degree of acoustic / optical occlusion / abstraction
[0147] The degree of acoustic / optical occlusion / abstraction can be measured by considering several aspects, such as the acoustic properties (e.g., transmission coefficient, absorption coefficient, reflection coefficient) of the occluder that blocks the direct line of sight from the listener and the distance between the audio source and the listener. For example, if several audio sources (multi-channel audio sources) are occluded but still acoustically important, these sources can simply be replaced by a single point source audio or FOA. If the source is blocked by an occluding material such that the listener cannot see the source, the source is considered occluded. The occluding material is described by its acoustic properties (such as transmission coefficient, reflection coefficient, and absorption coefficient). In this example, the occluded multi-channel audio source can be replaced by a mono audio source obtained from the dominant channel of the multi-channel audio or from a weighted mixture of the audio sources. Alternatively, the occluded multi-channel audio source can be replaced by a stereo audio source.
[0148] Figure 8 Illustrated is an exemplary method for converting the type of an audio source according to the present disclosure.
[0149] At step S801 each of a plurality of audio sources may be evaluated to obtain one or more values for the corresponding audio source, wherein the (plural) values indicate perceptual relevance information. For example, the (plural) values may indicate: (i) positional proximity; (ii) directional proximity; and / or (iii) degree of acoustic / optical occlusion / abstraction.
[0150] At step S802 for each audio source, each of the (plural) values is evaluated to determine whether it meets the (plural) corresponding predefined condition thresholds. For example, the threshold is 5 degrees. If the obtained value is below 5 degrees, it is considered to be in close proximity in terms of the directional proximity criterion. The obtained sources that meet the threshold are classified as type-converted sources.
[0151] Then it can be said that "the obtained value meets the condition threshold".
[0152] At step S803 the type-converted sources from step S802 are converted to another type. The audio type may be multi-channel audio (e.g., stereo, 5.0, 7.0, 11.0), mono audio, high-fidelity stereo (FOA, HOA).
[0153] Control of the culling, clustering, and type conversion phases
[0154] The culling, clustering, and type conversion methods described below may be performed individually or in various combinations.
[0155] Each of these methods may be implemented in the decoding and / or rendering phases of an audio decoder. The audio decoder and / or decoding method may be compatible with the standards set by the Audio Group of MPEG organized by ISO / IEC, such as the MPEG-I immersive audio standard. For example, each of these phases may be implemented individually or in combination as part of an MPEG-I rendering process, or as an external tool for the MPEG-I immersive audio standard.
[0156] Each method may be (de) / activated by control information. The control information may be provided as a condition. The control information may be in the form of parameters to be processed by the system and / or device. (Plural) parameters may be defined for each (virtual) object.
[0157] The application of the method of the present invention (i.e., culling, clustering, and type conversion) may result in a change in the state of the audio source. For example, the audio source state changes from "not culled" to "culled" or vice versa. If such a change occurs, it is desirable that the actual change in rendering should occur when the corresponding audio source has a lower loudness. This is to avoid sudden changes in loud signals. In this case, the control information can be a loudness or energy threshold that sets the threshold for when to trigger such a change. Additionally, the control information can include the maximum change allowed within a certain time period to avoid, for example, highly repetitive changes within a short time period. An example is that a voice signal is only allowed to change once within 1 minute. Alternatively, the control information can be set by the listener or an application on the decoder / renderer side.
[0158] Examples of control information parameters that can be used to control the application of the perceptual audio culling, clustering, and type conversion rendering phases include activation parameters or deactivation parameters. For example, such a parameter can indicate when to disable a certain (some) audio source (e.g., signal to disable a voice-over voice signal). An example of activation can be applied to certain audio sources in a scene, such as a car audio source. Generally, in a scene, the car sound can be active (e.g., the engine starts, the car moves) or inactive (e.g., the engine is off, the car is parked). In this case, activation means that the method of the present invention (i.e., culling, clustering, type conversion) can be applied to the car audio source. For example, this does not apply to a voice-over voice signal where it is expected that the voice signal is always rendered, so deactivation control is performed.
[0159] Another example of a control information parameter is grouped data. For example, the grouped data can be a set of audio source IDs that belong together, such as the parts of a car that make sounds (tires, engine, horn). If the direction difference between these audio sources with respect to the listener's position is less than 5 degrees (direction proximity), then these audio sources can be grouped together. The grouped data can be used to control each of culling / clustering / type conversion. For example, the grouped data can provide a way to define a group of audio sources within a specific region (e.g., the sound-producing parts of a car) or having a specific signal category (e.g., wind and rain environmental sounds).
[0160] When grouped data is available, the encoder, decoder, and / or renderer can define culling and clustering strategies to combine or remove audio sources (e.g., objects). For example, if all the audio sources in a group are below a certain threshold loudness, then the group is combined into one object (culling based on loudness). Alternatively, if the azimuth / elevation of all the objects in a group are each within a certain corresponding range, then the group can be combined into one object (clustering based on direction). Alternatively, if one or many objects / sources in a group are acoustically occluded, then the group is combined into one object / source.
[0161] Another example of a control information parameter is related to prioritization. For example, the control information can provide a prioritization regarding the signal perception relevance (and its information importance) with respect to the listener (and the listener's attention focus). For example, when performing "cocktail party effect" reproduction, the prioritization parameter can control when to apply culling (instead of leveling / equalizing) to non-relevant audio sources to improve speech intelligibility.
[0162] Another example of a control information parameter is the visibility of the graphical representation of the audio source to the listener (i.e., whether the visual representation of the audio source is visible to the user). For example, the audio source can be associated with a visual representation that: (i) is not in the video rendering viewport (behind the listener), and (ii) is occluded by an optically opaque occluder. Such audio sources can be processed differently depending on their nature and application scenario; for example, the thresholds for position proximity or direction proximity conditions can vary according to virtual object visibility. The control information parameter will define one or more sets of predefined condition thresholds for object visibility. The following aspects can be used to avoid frequent "switching" in the culling, clustering, and type conversion rendering phases:
[0163] - Hysteresis (the dependence of the current decision on the history of previous decisions)
[0164] - Distance relative to the listener's position (and its change over time - position velocity)
[0165] - Direction relative to the listener's position (and its change over time - angular velocity)
[0166] Figure 9 Illustrated is an exemplary use of control information for controlling one or more of the culling, clustering, and / or type conversion methods of the present invention. For example, at 900, listener pose (i.e., position) information and a plurality of audio sources P can be received. At 901A, control information A can be received. At 901B, control information B can be received. At 901C, control information C can be received. In various embodiments, 901A, 901B, and 901C can each be performed individually or in combination. Control information A illustrates an example when the information controls both the culling method and the clustering method. Control information B and C are used to control only the clustering method and only the type conversion method, respectively.
[0167] Figure 10Illustrated is control information parameters that will define one or more sets of predefined condition thresholds for object / source visibility. On the one hand, source A is clearly visible from the listener's position (not blocked by an occluder), and thus a set of condition thresholds X is applied to perform audio culling, clustering, and type conversion of the present invention. On the other hand, source B is blocked by an occluder, and thus another set of condition thresholds Y is applied to perform audio culling, clustering, and type conversion of the present invention. The control information here is whether the visual representation of the audio source is visible to the user.
[0168] An apparatus for implementing the method according to the present disclosure
[0169] Finally, the present disclosure similarly relates to apparatuses (e.g., computer-implemented apparatuses) for performing the methods and techniques described in the present disclosure. Figure 11 An example of such an apparatus 1100 is shown. In particular, apparatus 1100 includes a processor 1110 and a memory 1120 coupled to the processor 1110. The memory 1120 may store instructions for the processor 1110. Depending on the use case and / or implementation, among other things, the processor 1110 may also receive suitable input data 1130. Depending on the use case and / or implementation, the processor 1110 may be adapted to execute the methods / techniques described in the present disclosure and generate corresponding output data 1140.
[0170] Explanation
[0171] A computing device implementing the techniques described above may have the following example architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the example architecture includes one or more processors (e.g., dual-core processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components may exchange communications and data via one or more communication channels (e.g., a bus), which may utilize various hardware and software to facilitate the transfer of data and control signals between the components.
[0172] The term "computer-readable medium" refers to a medium that participates in providing instructions to a processor for execution, including but not limited to non-volatile media (e.g., optical disk or magnetic disk), volatile media (e.g., memory), and transmission media. Transmission media includes but is not limited to coaxial cable, copper wire, and optical fiber.
[0173] The computer-readable medium may further include an operating system (e.g., An operating system), a network communication module, an audio interface manager, an audio processing manager, and a real-time content distributor. The operating system can be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system performs basic tasks, including but not limited to: identifying inputs from network interfaces and / or devices and providing outputs to network interfaces and / or devices; recording and managing files and directories on a computer-readable medium (e.g., a memory or a storage device); controlling peripheral devices; and managing traffic on one or more communication channels. The network communication module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.).
[0174] The architecture can be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device with one or more processors. The software can include multiple software components or can be a single body of code.
[0175] The described features can be advantageously implemented in one or more computer programs that can be executed on a programmable system including at least one programmable processor, which is coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device and to transmit data and instructions to the data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform an activity or bring about a result. The computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.
[0176] For example, suitable processors for executing an instruction program include both general and special purpose microprocessors, as well as any one or more processors or cores of any kind of computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data files, or be operatively coupled to communicate therewith; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, which include by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, an ASIC (Application Specific Integrated Circuit).
[0177] To provide for interaction with a user, the features can be implemented on a computer having a display device, such as a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor or a retinal display device, for displaying information to the user. The computer can have a touch surface input device (e.g., a touch screen), or a keyboard and a pointing device such as a mouse or a trackball, by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user.
[0178] The features can be implemented in a computer system that includes a backend component, such as a data server, or includes a middleware component, such as an application server or an Internet server, or includes a frontend component, such as a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include, for example, LANs, WANs, and the computers and networks that form the Internet.
[0179] A computing system can include clients and servers. The clients and servers are typically remote from each other and typically interact over a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., in order to display data to a user interacting with the client device and receive user input from the user). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.
[0180] A system of one or more computers can be configured to perform particular operations by virtue of software, firmware, hardware, or a combination thereof installed in the system that, in operation, causes the system to perform the actions. One or more computer programs can be configured to perform particular operations by virtue of including instructions that, when executed by a data processing apparatus, cause the apparatus to perform those operations.
[0181] Although one or more embodiments have been described by way of example and in terms of specific embodiments, it should be understood that one or more embodiments are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements that would be apparent to those skilled in the art. Accordingly, the scope of the appended claims should be accorded the broadest interpretation so as to cover all such modifications and similar arrangements.
[0182] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0183] Unless otherwise specifically stated, it will be apparent from the following discussion that, throughout the disclosure, use of terms such as "processing," "computing," "calculating," "determining," "analyzing," etc., to refer to actions and / or processes of a computer or computing system, or similar electronic computing device, indicate manipulation and / or transformation of data represented as physical (such as electronic) quantities into other data similarly represented as physical quantities.
[0184] References throughout this disclosure to "an example embodiment," "some example embodiments," or "example embodiments" mean that a particular feature, structure, or characteristic described in connection with the example embodiments is included in at least one example embodiment of the present disclosure. Thus, appearances of the phrases "in an example embodiment," "in some example embodiments," or "in example embodiments" throughout this disclosure are not necessarily all referring to the same example embodiment. Moreover, in one or more example embodiments, the particular features, structures, or characteristics may be combined in any suitable manner, which will be apparent to those of ordinary skill in the art in light of the present disclosure.
[0185] As used herein, unless otherwise specified, the use of ordinal adjectives “first,” “second,” “third,” etc. to describe a common object merely indicates different instances of like objects and is not intended to imply that the objects so described must be in a given order in terms of time, space, rank, or any other manner.
[0186] Likewise, it should be understood that the terminology and terms used herein are for the purpose of description and should not be regarded as restrictive. The use of “including,” “comprising,” or “having” and their variants is intended to cover the items listed thereafter and their equivalents as well as additional items. Unless otherwise specified or limited, the terms “mounted,” “connected,” “supported,” and “coupled” and their variants are used broadly and cover direct and indirect mounting, connection, support, and coupling.
[0187] In the following claims and the description of the present disclosure herein, any one of the terms “comprising,” “comprised of,” or “which comprises” is an open term, which means including at least the subsequent element / feature, but not excluding other elements / features. Thus, when the term “comprising” is used in a claim, the term should not be construed as being limited to the devices or elements or steps listed thereafter. For example, the scope of the expression “a device including A and B” should not be limited to a device including only elements A and B. As used herein, any one of the terms “including” or “which includes” or “that includes” is also an open term, which also means including at least the element / feature after the term, but not excluding other elements / features. Thus, “including” is synonymous with “comprising” and means “comprising.”
[0188] It should be understood that in the above description of the exemplary embodiments of the present disclosure, various features of the present disclosure are sometimes combined in a single exemplary embodiment, figure, or its description in order to simplify the present disclosure and to assist in understanding one or more of the inventive aspects. However, the method of the present disclosure should not be construed as reflecting an intention that the claims require more features than are expressly recited in each claim. On the contrary, as reflected in the following claims, the inventive aspects lie in less than all the features of a single previously disclosed exemplary embodiment. Accordingly, the claims following the description are hereby expressly incorporated into this description, where each claim stands alone as a separate exemplary embodiment of the present disclosure.
[0189] In addition, while some example embodiments described herein include some features included in other example embodiments and not other features included in other example embodiments, as will be understood by those skilled in the art, combinations of features of different example embodiments are intended to be within the scope of the present disclosure and form different example embodiments. For example, in the appended claims, any of the example embodiments in the claimed example embodiments can be used in any combination.
[0190] In the description provided herein, numerous specific details are set forth. However, it should be understood that example embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0191] Accordingly, although the modes that are considered to be the best mode of the present disclosure have been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of the present disclosure, and all such changes and modifications that fall within the scope of the present disclosure are intended to be claimed. For example, any of the formulas given above merely represent procedures that can be used. Functions can be added or removed from the block diagrams, and operations can be interchanged among the functional blocks. Steps can be added or removed from the methods described within the scope of the present disclosure.
[0192] The enumerated example embodiments
[0193] Aspects and implementations of the present disclosure can also be understood from the following enumerated example embodiments (EEEs), which are not claims.
[0194] EEE1. A method for processing multiple audio sources, the method comprising:
[0195] For each audio source of the multiple audio sources, determining a corresponding rejection value;
[0196] For each audio source of the multiple audio sources, comparing the corresponding rejection value of each source of the multiple sources with a threshold to determine whether the source is a rejected source;
[0197] Outputting information identifying a rejected subset of the multiple audio sources, wherein the rejected subset includes the rejected source(s) identified by the comparison.
[0198] EEE2. The method according to EEE1, wherein the rejection value is based on at least one of: a loudness value, a relationship between the loudness of an individual source and the overall loudness of the overall rendered audio, and a relationship between the short-term spectral energy of an individual audio source and the overall spectral energy of the final rendered audio.
[0199] EEE3. The method as described in EEE1 further includes rendering the plurality of audio sources, wherein the rendering only renders the audio sources that are not part of the culled subset of the plurality of audio sources.
[0200] EEE4. A method for processing a plurality of audio sources, the method comprising:
[0201] For each audio source of the plurality of audio sources, determining a corresponding clustering value;
[0202] For each audio source of the plurality of audio sources, comparing the corresponding clustering value of each source of the plurality of sources with a threshold to determine whether the source is a clustered source;
[0203] Outputting information identifying the clustered subset of the plurality of audio sources based on the comparison.
[0204] EEE5. The method as described in EEE4, wherein the clustering value indicates positional proximity.
[0205] EEE6. The method as described in EEE5, wherein the threshold is a fraction of the distance between the listener and the nearest source.
[0206] EEE7. The method as described in EEE6, wherein the comparison compares the distance between the audio source and the listener with the threshold.
[0207] EEE8. The method as described in EEE4, wherein the clustering value indicates directional proximity.
[0208] EEE9. The method as described in EEE8, wherein the threshold is at least one of the azimuth angle or the elevation angle between the source and the listener's position.
[0209] EEE10. The method as described in EEE4 further includes a set of clusters, and rendering the set of clusters by an audio renderer.
[0210] EEE11. A method for processing a plurality of audio sources, the method comprising:
[0211] For each audio source of the plurality of audio sources, determining a corresponding type conversion value;
[0212] For each audio source of the plurality of audio sources, comparing the corresponding type conversion value of each source of the plurality of sources with a threshold to determine whether the source is a type conversion source;
[0213] Outputting information identifying the new source type that replaces the original source type.
[0214] EEE12. A method for processing control information, wherein the control information determines whether to execute one or more of the methods described in EEE 1, 4, and / or 11.
[0215] EEE13. The method according to EEE12, wherein the control information is one of the following: activation parameter or deactivation parameter, packet data, prioritization information, and / or visibility information.
[0216] EEE14. The method according to any one of EEE 1 to 13, wherein the method is performed according to a standard set by the MPEG Audio Group of ISO / IEC.
[0217] EEE15. The method according to EEE14, wherein the standard is the MPEG-I immersive audio standard.
Claims
1. A method of rendering an audio source, the method comprising: For each of a plurality of audio sources, determining a rendering control value based on one or more parameters indicative of the perceived relevance of the corresponding audio source; For each of the plurality of audio sources, comparing the rendering control value with a perception threshold to determine whether the corresponding audio source meets a perceived irrelevance criterion; Modifying the audio sources among the plurality of audio sources for which the rendering control value meets the threshold to obtain modified audio sources; And Rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio sources.
2. The method according to claim 1, wherein The one or more parameters include the loudness of the corresponding audio source; and wherein the rendering control value is determined based on the loudness value.
3. The method according to claim 1 or 2, wherein The one or more parameters include the relationship between the loudness of a single audio source and the loudness of an audio output generated by rendering the plurality of audio sources other than the single audio source.
4. The method according to any one of claims 1 to 3, wherein, The one or more parameters include the relationship between the short-term spectral energy of a single audio source and the spectral energy of an audio output generated by rendering the plurality of audio sources.
5. The method according to claim 4, wherein, The rendered audio output is associated with a psychoacoustic masking model.
6. The method according to any one of claims 1 to 5, wherein Modifying the audio sources for which the rendering control value meets the threshold includes removing the audio sources; and wherein the modified audio sources correspond to a removed subset among the plurality of audio sources.
7. The method according to claim 6, wherein Rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio sources includes not rendering the audio sources included in the removed subset.
8. The method according to any one of claims 1 to 7, wherein The one or more parameters include the positional proximity of the audio source relative to the listener's position.
9. The method according to claim 8, wherein, The rendering control value is determined based on the distance from the corresponding audio source position of the audio source for which the rendering control value is determined to another audio source, and wherein the rendering control value is compared with a fraction of the distance between the listener's position and the nearest source position.
10. The method according to any one of claims 1 to 9, wherein, The one or more parameters include the directional proximity relative to the listener's position.
11. The method according to claim 10, wherein, The rendering control value is determined based on at least one of the azimuth angle or elevation angle between the corresponding audio source position of the audio source for which the rendering control value is determined and the listener's position.
12. The method according to any one of claims 8 to 11, wherein Modifying the audio sources includes clustering the audio sources for which the rendering control value meets the threshold; and wherein the modified audio sources correspond to clustered audio sources.
13. The method according to claim 12, wherein, Rendering the unmodified audio sources among the plurality of audio sources relative to the modified audio sources includes replacing the clustered audio sources with a smaller number of audio sources and rendering the smaller number of audio sources as the modified audio sources.
14. The method according to any one of claims 1 to 13, wherein, The one or more parameters include one or more of acoustic occlusion, optical occlusion, or abstraction; and wherein the rendering control value is determined based on the degree of occlusion or abstraction of the corresponding audio source.
15. The method according to any one of claims 8 to 14, wherein, Modifying the audio sources includes converting the type of the audio sources for which the rendering control value meets the threshold; and wherein the modified audio sources correspond to audio sources of which the type has been converted.
16. The method according to claim 15, wherein, Rendering the unmodified audio source among the plurality of audio sources relative to the modified audio source includes rendering the type-converted audio source as the modified audio source.
17. The method according to any one of claims 8 to 16, wherein, On an audio source having a directivity pattern among the plurality of audio sources, if the rendering control value meets the threshold, the directivity calculation is not performed.
18. The method according to any one of claims 1 to 17, wherein The method further includes receiving control information indicating whether to perform modification on some or all of the audio sources for which the rendering control value meets the threshold.
19. The method according to claim 18, wherein, The control information is one of the following: activation parameter or deactivation parameter, grouping data, priority sorting information, and / or visibility information.
20. An apparatus for rendering an audio source, the apparatus including one or more processors configured to implement a method including the following operations: For each audio source among the plurality of audio sources, determine a rendering control value based on one or more parameters indicating the perceived relevance of the corresponding audio source; For each audio source among the plurality of audio sources, compare the rendering control value with a perception threshold to determine whether the corresponding audio source meets the perceived irrelevance criterion; Modify the audio sources among the plurality of audio sources for which the rendering control value meets the threshold to obtain modified audio sources; And Render the unmodified audio sources among the plurality of audio sources relative to the modified audio sources.
21. An apparatus includes a processor and a memory, the memory being coupled to the processor and storing instructions for the processor, wherein, The processor is adapted to execute the method according to any one of claims 1 to 19.
22. A program including instructions that, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 19.
23. A computer-readable storage medium storing the program according to claim 22.