Method, device and system for presentation, coding and decoding of discrete directivity information
By uniformly distributing unit vectors on a 3D sphere and calculating directional gains at the decoder, the method addresses the inefficiencies of existing discrete directional data representations, achieving efficient 6DoF audio rendering with reduced computational complexity and bit rate.
Patent Information
- Application Number
- JP2025134532
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-07-02
- Filing Date
- 2025-08-12
- Publication Date
- 2025-12-16
AI Technical Summary
Existing methods for representing and encoding discrete directional data of sound sources are suboptimal for 6DoF rendering, leading to unnecessary interpolation and large bitstream sizes due to redundancies and irrelevance in directional data.
A method for determining a number of unit vectors on a 3D sphere based on desired representation accuracy, using a predetermined arrangement algorithm to distribute these vectors uniformly, and calculating directional gains at the decoder, reducing the need for interpolation and bitstream size.
Provides efficient 6DoF audio rendering with reduced computational complexity and bit rate without degrading psychoacoustic perception by approximating directional information uniformly on a 3D sphere.
Smart Images

Figure 2025183213000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Provisional Application No. 62 / 869,622 (Reference No. D19038USP1), filed July 2, 2019, and European Application No. 19183862.2 (Reference No. D19038EP), filed July 2, 2019, the disclosures of both applications being incorporated herein by reference in their entirety.
[0002] The present disclosure relates to providing methods and apparatus for processing and encoding audio content that includes discrete directivity information (directional data) for at least one sound source. In particular, the present disclosure relates to representing, encoding, and decoding the discrete directivity information. [Background technology]
[0003] Real-world sound sources, whether natural or artificial (e.g., speakers, musical instruments, voices, mechanical devices), radiate sound non-isotropically. Characterizing a sound source's complex radiation pattern (or "directivity") can be crucial for proper rendering, especially in interactive environments such as video games and virtual / augmented reality applications. In these environments, users typically interact with directional audio objects by walking around them, thus changing the user's auditory perspective of the generated sound. Users may also be able to grasp and dynamically rotate virtual objects, again requiring the rendering of the corresponding sound source's radiation pattern in different directions. In addition to more realistic rendering of the direct propagation effect from the sound source to the listener, radiation characteristics also affect reverberant sound, as they play a major role in higher-order acoustic coupling between the sound source and its environment (e.g., the virtual environment in a video game). This, in turn, affects other spatial cues, such as perceived distance.
[0004] The radiation pattern of the sound source, or its parametric representation, needs to be transmitted as metadata to the six degrees of freedom (6DoF) audio renderer. The radiation pattern can be represented, for example, by a spherical harmonic decomposition or discrete vector data.
[0005] However, direct application of traditional discrete directional representations has proven to be suboptimal for 6DoF rendering. Summary of the Invention [Problem to be solved by the invention]
[0006] Therefore, there is a need for methods and apparatus for improved representation and / or improved coding of discrete directional data (directional information) of directional sound sources. [Means for solving the problem]
[0007] One aspect of the present disclosure relates to a method for processing audio content that includes directional information for at least one sound source. The method may be performed in an encoder in the case of encoding. Alternatively, the method may be performed in a decoder prior to rendering. The sound source may be, for example, a directional sound source and / or may be related to an audio object. The directional information may be discrete directional information. Furthermore, the directional information may be related to an audio object. The directional information may be part of metadata about the object. The directional information may include a first set of first directional unit vectors representing directional directions and associated first directional gains. The first directional unit vectors may be non-uniformly distributed on the surface of the 3D sphere. A unit vector refers to a vector of unit length. The method may include determining, as a count, a number of unit vectors to place on the surface of the 3D sphere based on a desired representation accuracy. The determining step may also refer to determining, based on the desired representation accuracy, a number of unit vectors to be generated for placement on the surface of the 3D sphere. The determined number of unit vectors may be defined as the cardinality of a set of unit vectors. The desired representation accuracy may be, for example, a desired angular accuracy or a desired directional accuracy. Furthermore, the desired representation accuracy may correspond to a desired angular resolution (e.g., in degrees). The method may further include generating a second set of second directional unit vectors by using a predetermined arrangement algorithm to distribute the determined number of unit vectors on the surface of a 3D sphere. The predetermined arrangement algorithm may be an algorithm for approximately uniformly distributing the unit vectors on the surface of the 3D sphere. The predetermined arrangement algorithm may be scalable depending on the number of unit vectors to be arranged / generated (i.e., the number may be a control parameter of the predetermined arrangement algorithm). The method may further include determining, for each second directional unit vector, an associated second directional gain based on the first directional gain of one or more first directional unit vectors of the group of first directional unit vectors that are closest to the respective second directional unit vector. The group of first directional unit vectors may be a suitable subgroup or subset of the first set of first directional unit vectors.
[0008] In this configuration, the proposed method provides a representation of the discrete directional information (i.e., the determined number and second directional gain) that can be rendered at the decoder without the need for interpolation to provide a "uniform response" to object-to-listener orientation changes. Furthermore, the representation of the discrete directional information can be coded at a low bit rate, since the perceptually relevant directional unit vectors are not stored in the representation but can be calculated at the decoder. Finally, the proposed method reduces the computational complexity during rendering.
[0009] In some embodiments, the number of unit vectors may be determined such that, when the unit vectors are distributed over the surface of a 3D sphere by a predetermined placement algorithm, they approximate the direction indicated by the first directional unit vector of the first set with at most a desired representation accuracy.
[0010] In some embodiments, the number of unit vectors may be determined such that, for each first directional unit vector in the first set, there is at least one unit vector whose directional difference with respect to the respective first directional unit vector is less than a desired representation accuracy when the unit vectors are distributed on the surface of the 3D sphere by a predetermined placement algorithm. The directional difference may be, for example, an angular distance. The directional difference may be defined with respect to an appropriate directional difference norm.
[0011] In some embodiments, determining the number of unit vectors may include using a pre-established functional relationship between the representational precision and the number of corresponding unit vectors that are distributed on the surface of the 3D sphere by a predetermined placement algorithm and that approximate at most the directions indicated by the first directional unit vectors of the first set with the respective representational precision.
[0012] In some embodiments, determining an associated second directional gain for a given second directional unit vector may include setting the second directional gain to a first directional gain associated with a first directional unit vector that is closest (for purposes of this disclosure, closeness is defined by an appropriate distance criterion) to the given second directional unit vector. Alternatively, this determination may include, for example, stereo projection or triangulation.
[0013] In some embodiments, the predetermined placement algorithm may include superimposing a spiral path on the surface of the 3D sphere, the spiral path extending from a first point on the 3D sphere to a second point on the 3D sphere opposite the first point, and placing unit vectors successively along the spiral path, wherein a spacing between the spiral path and / or an offset between each two adjacent unit vectors along the spiral path may be determined based on the number of unit vectors.
[0014] In some embodiments, determining the number of unit vectors may further include mapping (rounding) the number of unit vectors to one of a predetermined number. The predetermined number may be transmitted by a bitstream parameter. For example, the bitstream parameter may be a 2-bit parameter such as a directivity_precision parameter. In the case of encoding, the method may include encoding the determined number into a value of the bitstream parameter.
[0015] In some embodiments, the desired rendering accuracy may be determined based on a model of the perceptual directional sensitivity thresholds of a human listener (eg, a human reference listener).
[0016] In some embodiments, the cardinality of the second directional unit vectors of the second set may be less than the cardinality of the first directional unit vectors of the first set, which may imply that the desired representational precision is less than the representational precision provided by the first directional unit vectors of the first set.
[0017] In some embodiments, the first and second directional unit vectors may be expressed in spherical or Cartesian coordinate systems. For example, the first directional unit vector may be uniformly distributed in the azimuth-elevation plane, which implies a non-uniform (spherical) distribution on the surface of a 3D sphere. The second directional unit vector may be non-uniformly distributed in the azimuth-elevation plane in a manner that is (quasi-)uniformly distributed on the surface of the 3D sphere.
[0018] In some embodiments, the directional information represented by the first set of first directional unit vectors and associated first directional gains may be stored in the Spatial Oriented Format for Acoustics (SOFA) format. The SOFA format includes formats standardized by the Audio Engineering Society (see, for example, AES69-2015). Additionally or alternatively, the directional information represented by the second set of first directional unit vectors and associated second directional gains may be stored in the SOFA format.
[0019] In some embodiments, the method may be a method for encoding audio content and may further include encoding the determined number of unit vectors together with the second directivity gain into a bitstream. The method may further include outputting the bitstream. This assumes that at least part of the proposed method is performed on the encoder side.
[0020] Another aspect of the present disclosure is an audio control system that includes directional information about at least one sound source. The present invention relates to a method for decoding audio content. The directional information may include a number (e.g., a count) indicating the number of unit vectors approximately uniformly distributed on the surface of a 3D sphere, and, for each such unit vector, an associated directional gain. The unit vectors may be assumed to be distributed on the surface of the 3D sphere by a predetermined constellation algorithm. Here, the predetermined constellation algorithm may be an algorithm for approximately uniformly spherically distributing the unit vectors on the surface of the 3D sphere. The method may include receiving a bitstream including audio content. The method may further include extracting the number and the directional gain from the bitstream. The method may further include determining (e.g., generating) a set of directional unit vectors by using the predetermined constellation algorithm to distribute the number of unit vectors on the surface of the 3D sphere. In this sense, the number of unit vectors may function as a control parameter of the predetermined constellation algorithm. The method may further include associating each directional unit vector with its directional gain. This aspect assumes that the proposed method is distributed between the encoder and decoder sides.
[0021] In some embodiments, the method may further include determining, for a given target directional unit vector pointing in a direction from the sound source to the listener position, a target directional gain for the target directional unit vector based on associated directional gains of one or more directional unit vectors of a group of directional unit vectors that are closest to the target directional unit vector, The group of directional unit vectors may be a suitable subgroup or subset of the set of directional unit vectors.
[0022] In some embodiments, determining the target directional gain for the target directional unit vector may include setting the target directional gain to the directional gain associated with the directional unit vector that is closest to the target directional unit vector.
[0023] Another aspect of the present disclosure relates to a method for decoding audio content including directional information for at least one sound source. The directional information may include a first set of first directional unit vectors representing directional directions and associated first directional gains. The method may include receiving a bitstream including the audio content. The method may further include extracting the first set of directional unit vectors and associated first directional gains from the bitstream. The method may further include determining a number of unit vectors to arrange on the surface of a 3D sphere based on a desired representation accuracy as a count. The method may further include generating a second set of second directional unit vectors by using a predetermined arrangement algorithm to distribute the determined number of unit vectors on the surface of the 3D sphere. Here, the predetermined arrangement algorithm may be an algorithm for approximately uniformly distributing the unit vectors on the surface of the 3D sphere. The method may further include, for a second directional unit vector, determining an associated second directional gain based on the first directional gains of one or more first directional unit vectors of the group of first directional unit vectors that are closest to the respective second directional unit vector. The method may further include, for a given target directional unit vector pointing in a direction from the sound source to the listener position, determining a target directional gain for the target directional unit vector based on the associated second directional gains of one or more second directional unit vectors of the group of second directional unit vectors that are closest to the target directional unit vector. The group of second directional unit vectors may be a suitable subgroup or suitable subset of the second set of second directional unit vectors. This aspect assumes that all of the proposed method is performed on the decoder side.
[0024] In some embodiments, the target pointing to the target pointing unit vector Determining the directional gain may include setting the target directional gain to a second directional gain associated with a second directional unit vector that is closest to the target directional unit vector.
[0025] In some embodiments, the method may further include extracting from the bitstream an indication of whether a second set of directional unit vectors should be generated. This indication may be a one-bit flag, e.g., a directivity_type parameter. If the indication indicates that a second set of directional unit vectors should be generated, the method may further include determining the number of unit vectors and generating second directional unit vectors of the second set. Otherwise, the number of unit vectors and the (second) directional gain may be extracted from the bitstream.
[0026] Another aspect of the present disclosure relates to an apparatus for processing audio content including directional information for at least one sound source. The directional information may include a first set of first directional unit vectors representing directional directions and associated first directional gains. The apparatus may include a processor configured to perform the steps of the method of the first aspect and any of its embodiments.
[0027] Another aspect of the present disclosure is an apparatus for decoding audio content including directional information for at least one sound source. The directional information may include a count indicating the number (e.g., count number) of unit vectors approximately uniformly distributed on the surface of a 3D sphere, and, for each such unit vector, an associated directional gain. The unit vectors may be assumed to be distributed on the surface of the 3D sphere by a predetermined placement algorithm. Here, the predetermined placement algorithm may be an algorithm for approximately uniformly spherically distributing the unit vectors on the surface of the 3D sphere. The apparatus may include a processor configured to perform the steps of the method of the second aspect and any of its embodiments.
[0028] Another aspect of the present disclosure relates to an apparatus for decoding audio content including directional information for at least one sound source. The directional information may include a first set of first directional unit vectors representing directional directions and associated first directional gains. The apparatus may include a processor configured to perform the steps of the method of the third aspect and any of its embodiments.
[0029] Another aspect of the present disclosure relates to a computer program comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of the above first to third aspects and any of the embodiments thereof.
[0030] Another aspect of the present disclosure relates to a computer readable medium storing the computer program of the preceding aspect.
[0031] Another aspect of the present disclosure relates to an audio decoder including a processor connected to a memory storing instructions for the processor, the processor being configured to perform a method according to each one of the above aspects or embodiments.
[0032] Another aspect of the present disclosure relates to an audio encoder including a processor connected to a memory storing instructions for the processor, the processor being configured to perform a method according to each one of the above aspects or embodiments.
[0033] Further aspects of the present disclosure relate to corresponding computer programs and computer-readable storage media.
[0034] It will be understood that method steps and apparatus features are in various ways interchangeable. In particular, those skilled in the art will understand that details of a disclosed method can be implemented as an apparatus configured to perform some or all of the method steps, and vice versa. In particular, it will be understood that each term referred to in the context of a method equally applies to the corresponding apparatus, and vice versa. [Brief explanation of the drawings]
[0035] Examples of embodiments of the present disclosure will now be described with reference to the accompanying drawings, in which like or similar elements are designated with like reference numerals, and in which:
[0036] [Figure 1A] FIG. 1A is a diagram that schematically illustrates an example representation of directional information including discrete directional unit vectors and associated directional gains. [Figure 1B] FIG. 1B is a diagram illustrating a schematic example of a representation of directional information including discrete directional unit vectors and associated directional gains. [Figure 1C] FIG. 1C is a diagram that schematically illustrates an example of a representation of directional information including discrete directional unit vectors and associated directional gains. [Figure 2] FIG. 2 is a diagram illustrating a schematic example of a directional unit vector and its associated directional gain. [Figure 3] FIG. 3 is a diagram illustrating a schematic example of an arrangement of directional unit vectors on the surface of a 3D sphere according to a desired representation accuracy. [Figure 4] FIG. 4 is a diagram illustrating another example of the arrangement of directional unit vectors on the surface of a 3D sphere according to a desired representation accuracy. [Figure 5] FIG. 5 is a graph that schematically illustrates the relationship between the number of unit vectors and the resulting representation accuracy, given a given placement algorithm for placing unit vectors on the surface of a 3D sphere. [Figure 6] FIG. 6 is a graph illustrating a modeled relationship between the number of unit vectors and the resulting representation accuracy, given a given placement algorithm for placing unit vectors on the surface of a 3D sphere. [Figure 7A] FIG. 7A is a diagram that schematically illustrates an example representation of directional information including discrete directional unit vectors and associated directional gains, according to an embodiment of the present disclosure. [Figure 7B]FIG. 7B is a diagram that schematically illustrates an example representation of directional information including discrete directional unit vectors and associated directional gains, according to an embodiment of the present disclosure. [Figure 7C] FIG. 7C is a diagram that schematically illustrates an example representation of directional information including discrete directional unit vectors and associated directional gains, according to an embodiment of the present disclosure. [Figure 8A] FIG. 8A is a diagram illustrating a conventional representation of discrete directional information for different representation accuracy. [Figure 8B] FIG. 8B is a diagram illustrating a schematic representation of discrete directional information for different representation accuracy according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating, in flow chart form, a method for processing or encoding audio content including directional information for at least one sound source, according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating, in flow chart form, a schematic example of a method for decoding audio content including directional information for at least one sound source, according to an embodiment of the present disclosure. [Figure 11] FIG. 11 is a diagram illustrating, in flow chart form, another example of a method for decoding audio content including directional information for at least one sound source, according to an embodiment of the present disclosure. [Figure 12] FIG. 12 is a diagram that schematically illustrates an apparatus for processing or encoding audio content that includes directional information about at least one sound source, according to an embodiment of the present disclosure. [Figure 13] FIG. 13 is a diagram that schematically illustrates an apparatus for decoding audio content that includes directional information for at least one sound source, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0037] As noted above, in this disclosure, identical or similar elements are indicated by identical or similar reference numerals, and repeated description is omitted for the sake of brevity.
[0038] Audio formats that include directional data (directional information) about sound sources can be used for 6DoF rendering of audio content. In some of these audio formats, the directional data is discrete directional data stored as a set of discrete vectors (e.g., azimuth, elevation) and magnitude (e.g., gain) (e.g., in the SOFA format). However, as described above, directly applying such conventional discrete directional representations for 6DoF rendering has proven to be suboptimal. In particular, for conventional discrete directional representations, the vector directions are typically significantly non-equidistantly spaced in 3D space, necessitating interpolation between the vector directions during rendering (e.g., 6DoF rendering). Furthermore, the directional data contains redundancies and irrelevance, resulting in a large bitstream size for encoding the representation.
[0039] An example of a conventional representation of discrete directional information of a sound source is illustrated schematically in FIGS. 1A, 1B, and 1C. The conventional representation includes a plurality of discrete directional unit vectors 10 and associated directional gains 15. FIG. 1A illustrates a 3D view of the directional unit vectors 10 arranged on the surface of a 3D sphere. In this example, these directional unit vectors 10 are uniformly (i.e., equidistantly) arranged in the azimuth-elevation plane, resulting in a non-uniform spherical arrangement on the surface of the 3D sphere. This can be seen in FIG. 1B, which illustrates a top-down view of the 3D sphere on which the directional unit vectors 10 are arranged. Finally, FIG. 1C illustrates the directional gains 15 for the directional unit vectors 10, thereby indicating the radiation pattern (or "directivity") of the sound source.
[0040] Although the representation of discrete directional information can be improved because the direction can be calculated at the decoder side (e.g., via a formula, table, or other pre-computed lookup information), such conventional representations may involve unnecessarily fine directional sampling from a psychoacoustic point of view.
[0041] This disclosure provides a set of M discrete acoustic source directional gains. Assume an initial (e.g., conventional) representation of discrete directional information about a sound source (acoustic source) containing JPEG2025183213000002.jpg54. JPEG2025183213000003.jpg54 is a non-uniformly distributed directional unit vector It is defined on JPEG2025183213000004.jpg512. Here, each directional unit vector JPEG2025183213000005.jpg53 is the directivity gain associated with itself. JPEG2025183213000006.jpg516. A directional unit vector is a directional vector of unit length. Directional unit vector JPEG2025183213000007.jpg53, 210, and associated directional gain JPEG2025183213000008.jpg54 is shown in FIG. 2. In this example, the directional unit vector JPEG2025183213000009.jpg53 is placed on the surface 230 of a 3D sphere, which is a unit sphere. A set of directional unit vectors JPEG2025183213000010.jpg53 may be referred to as a first directional unit vector of the first set for the purposes of this disclosure. JPEG2025183213000011.jpg54 may be referred to as a first directional gain associated with each of the first directional vectors.
[0042] As mentioned above, the directional unit vector The non-uniform distribution of JPEG2025183213000012.jpg53 requires the decoder to adjust the directional gain to achieve a uniform response across object-to-listener orientation changes. JPEG2025183213000013.jpg54 requires interpolation.
[0043] To address this issue, the present disclosure provides a method for reconstructing the original data in a manner that produces equivalent (e.g., subjectively non-distinguishable) 6DoF audio rendering output. Optimized directional representation approximating JPEG2025183213000014.jpg53 The purpose is to provide JPEG2025183213000015.jpg53. Here, the directional unit vector JPEG2025183213000016.jpg53 and / or directional unit vectors JPEG2025183213000017.jpg53 may be represented, for example, in a spherical coordinate system or a Cartesian coordinate system.
[0044] Optimized Representation JPEG2025183213000018.jpg53 is a directional vector It is defined on the quasi-uniform distribution of JPEG2025183213000019.jpg53. This results in a smaller bitstream size Bs. That is, JPEG2025183213000020.jpg523 and / or allows for a computationally efficient decoding process. For the purposes of this disclosure, quasi-uniform means uniform up to a given (e.g., desired) representational accuracy.
[0045] To do so, this disclosure assumes that the object-to-listener orientation is arbitrary with respect to a uniform probability distribution, and that the object-to-listener orientation representation accuracy (i.e., the desired representation accuracy) is known, e.g., defined based on the subjective directional sensitivity threshold of a human listener (e.g., a human reference listener).
[0046] The present disclosure provides at least the following technical advantages: The first technical advantage relates to the advantage of parameterizing directional information using a uniform directional representation in 3D space (not in the azimuth-elevation plane). The second ... This benefits from discarding the directional information contained in JPEG2025183213000021.jpg53.
[0047] A uniform directional representation is not easy because JPEG2025183213000022.jpg53-direction uniform distribution problems (e.g., equidistantly spaced N points on the surface of a 3D unit sphere) are generally impossible to solve exactly for any N>4, and numerical approximations for generating (quasi-)equidistantly distributed points on a 3D unit sphere are often very complex (e.g., iterative, stochastic, computationally intensive).
[0048] Original data Reducing irrelevance and redundancy in JPEG2025183213000023.jpg53 is also not easy, since it is highly related to the definition of orientation representation accuracy based on psychoacoustic considerations.
[0049] Based on at least these technical advantages, the present disclosure proposes an efficient approximation method for uniform directional representation that avoids interpolation of directional gains at the decoder side and allows for significant bitrate reduction without degrading the psychoacoustic directional perception obtained from the 6DoF rendered output.
[0050] 9 illustrates, in flowchart form, an example method 900 for processing (or encoding) audio content including (discrete) directional information for at least one sound source (e.g., audio object) according to an embodiment of the present disclosure. Assume that the directional information relates to directional information G as defined above, i.e., includes a first set of first directional unit vectors representing directional directions and associated first directional gains. The directional information G may be included in the audio content as part of metadata for the sound source (e.g., audio object).
[0051] As an initial step (not shown in the flowchart), method 900 may obtain audio content. Directivity information represented by a first set of first directivity vectors and associated first directivity gains may be stored in SOFA format.
[0052] In step S910, the number N of unit vectors for placement on the surface of the 3D sphere is determined (e.g., calculated) as a count based on a desired representation accuracy D. This may involve determining (e.g., calculating) the number N of (quasi-)equidistantly distributed direction or (directional) unit vectors (e.g., based on a given orientation representation accuracy D). Here, quasi-equidistantly distributed is understood to mean equidistantly distributed with at most representation accuracy D. The representation accuracy D may correspond, for example, to angular accuracy or directional accuracy. In this sense, representation accuracy may correspond to angular resolution. In some implementations, the desired representation accuracy may be determined based on a model of the perceptual directional threshold of a human listener (e.g., a human reference listener).
[0053] In particular, the output of this step is a single integer, namely, the number N of directional unit vectors. The actual generation of directional unit vectors is performed in step S920, described below. In other words, step S910 determines the cardinality of the set of directional unit vectors to be generated. The number N of unit vectors is determined by the following equation: JPEG2025183213000024.jpg53 unit vectors may be determined to approximate the direction indicated by the first directional vector of the first set with up to a desired representation accuracy D. Thus, for a given alignment The placement algorithm may be an algorithm for approximately uniformly distributing the unit vectors spherically (e.g., at most representational accuracy) on the surface of a 3D sphere. Examples of such placement algorithms are described below. In other words, the number N of unit vectors may be determined such that, when the unit vectors are distributed on the surface of the 3D sphere by the predetermined placement algorithm, for each first directional unit vector in the first set, there is at least one unit vector among the unit vectors whose directional difference with respect to the respective first directional unit vector is less than the desired representational accuracy D. The number N may function as a scalar (i.e., control parameter) for the predetermined placement algorithm. That is, the predetermined placement algorithm may be suitable for distributing any number of unit vectors on the surface of a 3D sphere.
[0054] In the above, the direction difference may be, for example, an angular distance (e.g., an angle). The direction difference may be defined with respect to an appropriate direction difference norm (e.g., a direction difference norm depending on the scalar product of the directional unit vectors involved).
[0055] In step S920, a second set of second directional unit vectors is generated by using a predetermined placement algorithm to distribute the determined number N of unit vectors over the surface of a 3D sphere. As described above, the predetermined placement algorithm is an algorithm for approximately uniformly distributing the unit vectors over the surface of the 3D sphere. The second directional unit vectors are the directional unit vectors defined above. JPEG2025183213000025.jpg511. This step therefore uses a predetermined placement algorithm controlled by the scaler N to generate the directional vector JPEG2025183213000026.jpg511. Preferably, the cardinality of the second directional unit vector of the second set is less than the cardinality of the first directional unit vector of the first set. This assumes that the desired representation accuracy D is less than the representation accuracy provided by the first directional unit vector of the first set.
[0056] In step S930, for a second directional unit vector, an associated second directional gain is determined (e.g., calculated) based on the first directional gain. For example, this determination may be based on the first directional gain of one or more first directional unit vectors of a group of first directional unit vectors that are closest to the second directional unit vector. For example, this determination may include stereo projection or triangulation. In a particularly simple implementation, the second directional gain for a given second directional unit vector is set to the first directional gain associated with the first directional unit vector that is closest to the given second directional vector (i.e., has the shortest directional distance to the given second directional vector). Generally, this step can be performed as follows: JPEG2025183213000027.jpg53 of the original data G defined above Directivity approximation defined on JPEG2025183213000028.jpg53 JPEG2025183213000029.jpg53. The directivity information represented by the second set of second directivity vectors and associated second directivity gains may exist (e.g., may be stored) in a SOFA format.
[0057] If the method 900 is an encoding method, the method 900 further includes steps S940 and S950, which will be described below, in which case the method 900 may be performed in an encoder.
[0058] In step S940, the determined number N of unit vectors is coded into the bitstream together with the second directivity gain. The encoding may involve encoding a bitstream including JPEG2025183213000030.jpg53 and a number N. The directivity information represented by the second set of second directivity vectors and associated second directivity gains may exist (e.g., may be stored) in a SOFA format.
[0059] In step S950, the bitstream is output, for example, the bitstream may be output for transmission to a decoder or for storage on a suitable storage medium.
[0060] FIG. 10 illustrates in flowchart form an example method 1000 for decoding audio content including (discrete) directivity information for at least one sound source (e.g., audio object) according to an embodiment of the present disclosure. Method 1000 may be performed in a decoder. The audio content may be encoded into a bitstream, for example, by steps S910-S950 of method 900 described below. Thus, the directivity information may include a number N indicating the number of unit vectors approximately uniformly distributed on the surface of a 3D sphere, and (a representation of) an associated directivity gain for each such unit vector. The associated directivity gain may be the second directivity gain (data The unit vectors can be used for a given placement algorithm (e.g., audio content placement). The vectors may be assumed to be distributed over the surface of a 3D sphere by a predetermined distribution algorithm (same as the one used to process / encode the content), where the predetermined distribution algorithm is an algorithm for approximately uniform spherical distribution of unit vectors over the surface of a 3D sphere.
[0061] In step S1010, a bitstream containing audio content is received.
[0062] In step S1020, the number N and the directional gain are extracted from the bit stream (e.g., by a demultiplexer). Decode the bitstream containing JPEG2025183213000032.jpg53 and the number N to obtain the data JPEG2025183213000033.jpg53 and the number N.
[0063] In step S1030, a set of directional unit vectors is determined (e.g., generated) by using a predetermined arrangement algorithm to distribute N unit vectors on the surface of a 3D sphere. This step may proceed in the same manner as step S920 above. Each directional unit vector determined in this step has an associated directional gain from the directional gains extracted from the bitstream in step S1020. Assuming that the same predetermined arrangement algorithm is used in processing / encoding the audio content and decoding the audio content, the directional unit vectors generated in step S1030 are determined in the same order as the second directional unit vectors generated in step S920. Then, encoding the second directional gains as an ordered set in the bitstream in step S940 allows unambiguous assignment of a directional gain to each of the generated directional unit vectors in step S1030.
[0064] In step S1040, for a given target directional unit vector pointing in a direction from the sound source to the listener's position, a target directional gain for the target directional unit vector is determined (e.g., calculated) based on the directional gain associated with the directional unit vector. For example, the target directional gain may be determined (e.g., calculated) based on the directional gain associated with one or more directional unit vectors in a group of directional unit vectors that are closest to the target directional unit vector.
[0065] For example, this determination may involve stereo projection or triangulation. In a particularly simple implementation, the target directional gain for a target directional unit vector is set to the directional gain associated with the directional unit vector closest to the target directional vector (i.e., the directional unit vector with the shortest directional distance to the target directional vector). In general, this process is performed using a directional unit vector with the shortest directional distance to the target directional vector. Defined above JPEG2025183213000034.jpg53 It may involve using JPEG2025183213000035.jpg53.
[0066] Alternatively, the steps outlined above can be distributed differently between the encoder and decoder. For example, in situations where the encoder cannot perform the operations of method 900 listed above (e.g., when the accuracy of the proposed approximation (representation accuracy) can be defined only on the decoder side), the necessary steps can be performed only on the decoder side. This does not result in a smaller bitstream size, but still has the advantage of saving computational complexity on the decoder side for rendering.
[0067] 11 illustrates in flowchart form a corresponding example of a method 1100 for decoding audio content including (discrete) directional information for at least one sound source (e.g., audio object) according to an embodiment of the present disclosure. Assume that the directional information relates to directional information G as defined above, i.e., includes a first set of first directional unit vectors representing directional directions and associated first directional gains. In this sense, contrary to method 1000, method 1100 receives as input audio content whose directional information has not yet been optimized by a method according to the present disclosure. The directional information G may be included in the audio content as part of metadata about the sound source (e.g., audio object).
[0068] In step S1110, a bitstream containing audio content is received. Alternatively, the audio content may be obtained by any other possible means, depending on the application.
[0069] In step S1120, a first set of directional unit vectors and associated first directional gains are extracted from the bitstream (or obtained by any other possible means, depending on the application). In one example, the directional vectors and associated first directional gains may be demultiplexed from the bitstream.
[0070] In step S1130, the number of vectors for placement on the surface of the 3D sphere is determined as a count based on the desired representation accuracy. This step may proceed in the same manner as step S910 above.
[0071] In step S1140, a second set of second oriented vectors is generated by using a predetermined placement algorithm to distribute the determined number of unit vectors over the surface of a 3D sphere. A uniform unit vector is generated. The predetermined placement algorithm is an algorithm for approximately uniformly distributing the unit vectors on the surface of a 3D sphere. This step may proceed in the same manner as step S920 above.
[0072] In step S1150, for each second directional unit vector, an associated second directional gain is determined based on the first directional gain. For example, the associated second directional gain may be determined for each second directional unit vector based on the first directional gain of one or more directional unit vectors in the group of first directional unit vectors that are closest to the respective second directional unit vector. Thus, this step may proceed in the same manner as step S930 above.
[0073] In step S1160, for a given target directional unit vector pointing in a direction from the sound source to the listener's position, a target directional gain for the target directional unit vector is determined based on the second directional gain. For example, the target directional gain may be determined for the target directional unit vector based on the associated second directional gains of one or more second directional unit vectors in a group of second directional unit vectors that are closest to the target directional unit vector. This step may proceed in the same manner as step S1040 above.
[0074] In a particularly simple implementation, the target directional gain for a target directional unit vector is set to the second directional gain associated with the second directional unit vector that is closest to the target directional vector (i.e., has the shortest directional distance to the target directional vector).
[0075] Since there may be flexibility in which steps are performed on the encoder side and the decoder side, it is further suggested to transmit to the decoder which steps the decoder should perform (or, in other words, which format the directional data has). This can be easily done with one bit of information, for example, using the bitstream syntax for directional representation transmission shown in Table 1 below. Possible bitstream variable definitions for directional representation transmission are shown in Table 2 below.
[0076] [Table 1]
[0077] [Table 2]
[0078] In accordance with the above, a method for decoding audio content according to an embodiment of the present disclosure includes: determining from the bitstream whether a second set of directional unit vectors should be generated; Further, the method may include determining the number of unit vectors and generating a second set of second directional unit vectors if (and only if) the indication indicates that a second set of directional unit vectors should be generated. The indication may be a one-bit flag, for example, the directivity_type parameter defined above.
[0079] Using the methods of the present disclosure, we can generate representations of discrete directional data that do not require interpolation during 6DoF rendering to provide a "uniform response" to object-to-listener orientation changes. Furthermore, perceptually relevant directional unit vectors Because JPEG2025183213000038.jpg53 is calculated rather than stored, a low bit rate can be achieved in transmitting the representation.
[0080] An example of a representation of the discrete directional data of a sound source that can be achieved by the method according to the present disclosure is illustrated schematically in Figures 7A, 7B, and 7C. This representation should be compared with the representations illustrated schematically in Figures 1A, 1B, and 1C. Figure 7A shows a (second) directional unit vector placed on the surface of a 3D sphere. JPEG2025183213000039.jpg53, 20. These directional unit vectors 20 are spatially uniformly distributed over the surface of a 3D sphere. This implies a non-uniform distribution in the azimuth-elevation plane. This can be seen in FIG. 7B, which is a top-down view of the 3D sphere on which the directional unit vectors 20 are located. Finally, FIG. 7C illustrates the (second) directional gain 25 for the (second) directional unit vector 20. This gives an indication of the radiation pattern (or "directivity") of the sound source. The envelope of this pattern is substantially identical to the envelope of the pattern illustrated in FIG. 1C and contains the same amount of relevant psychoacoustic information.
[0081] 8A and 8B are different JPEG2025183213000040.jpg illustrates a further example comparing a conventional representation of discrete directional data of a sound source with a representation according to an embodiment of the present disclosure for 53 directional unit vectors (and corresponding orientation representation accuracy D). 8B (bottom) illustrates a representation according to an embodiment of the present disclosure. JPEG2025183213000042.jpg53 is an example. The leftmost panel is JPEG2025183213000043.jpg511 and This is relevant for the case of JPEG2025183213000044.jpg511. The second panel from the left is JPEG2025183213000045.jpg511 and This is relevant for the case of JPEG2025183213000046.jpg511. The third panel from the left is JPEG2025183213000047.jpg512 and This concerns the case of JPEG2025183213000048.jpg511. The rightmost panel is JPEG2025183213000049.jpg512 and This is relevant in the case of JPEG2025183213000050.jpg511.
[0082] A specific implementation example of the above method steps of the method according to the embodiment of the present disclosure will now be described.
[0083] For specific implementation examples of these, JPEG2025183213000051.jpg Assume that the original set of 53 discrete acoustic source directivity measurements (estimates) G is given by the following radiation pattern format:
number
[0084] Using the above assumptions, step S920 of method 900 (or step S1140 of method 1100) may proceed as follows.
[0085] N directional vectors that approximate a uniform directional distribution in 3D space (i.e., positions on a 3D unit sphere) Any suitable numerical approximation method (placement algorithm) may be used to calculate (i.e., generate) JPEG2025183213000058.jpg511 (see, for example, D.P. Hardina, T. Michaelsab, E.B.S.aff "A Comparison of Popular Point Configurations on S 2 ” (2016) Dolomites Research Notes on Approximation: Volume 9, Pages 16-49). However, this disclosure is not intended to be limited by, but may be modified in any way by, for example, Kogan, Jonathan “A New Computationally Efficient We propose to consider a specific approximation method (placement algorithm) based on the paper "Method for Spacing n Points on a Sphere" (2017) Rose-Hulman Undergraduate Mathematics Journal: Volume 18, Issue 2, Article 5. The reason for this choice is that the method has low computational complexity and it uses a single control parameter, It depends on JPEG2025183213000059.jpg53, and there is no limit to its control parameter N ( JPEG2025183213000060.jpg59).
[0086] The following equation (e.g., solved in the encoder and decoder): Define JPEG2025183213000061.jpg53 and in the bitstream Avoid explicitly storing JPEG2025183213000062.jpg53.
number
number
number
[0087] More generally, the predetermined placement algorithm may include superimposing a spiraling path on the surface of a 3D sphere. The spiraling path extends from a first point on the sphere (e.g., one pole) to a second point on the sphere opposite the first point (e.g., the other pole). The predetermined placement algorithm may then sequentially place unit vectors along the spiraling path. The spacing of the spiraling path and the offset (e.g., step) between each two adjacent unit vectors along the spiraling path may be determined based on the number N of unit vectors.
[0088] Use the following example of a MatLab function to calculate a directional vector JPEG2025183213000068.jpg53 can be generated.
number
[0089] Using the following example MatLab script, we can calculate the vector JPEG2025183213000070.jpg53 can be represented.
number
[0090] Using the above assumptions, step S910 of method 900 (or step S1130 of method 1100) may proceed as follows.
[0091] directional vector To calculate JPEG2025183213000072.jpg53, the control parameters are calculated based on the orientation representation accuracy value D defined below. JPEG2025183213000073.jpg53 must be specified.
number
[0092] In simple terms, any (∀) direction Corresponding direction for JPEG2025183213000075.jpg53 JPEG2025183213000076.jpg54 (e.g., as defined by the method of step S920) At least one (symbol: F, upside down, left and right inverted) index that differs from JPEG2025183213000077.jpg53 by a value less than or equal to the orientation representation accuracy D JPEG2025183213000078.jpg52 exists.
[0093] This is illustrated schematically in Figure 3. In Figure 3, the directional unit vector The maximum distance 310 from the nearest directional unit vector of JPEG2025183213000079.jpg53, 20 is less than the desired representation accuracy D. This is because the surface of the 3D sphere is The periphery of JPEG2025183213000080.jpg53 is subdivided into multiple cells, and each cell is the directional unit vector of that cell. JPEG2025183213000081.jpg53, any other directional unit vector JPEG2025183213000082.jpg53, assuming it includes all directions closer than the nearest directional unit vector This can be achieved by ensuring that the directional difference in any direction on the cell boundary for JPEG2025183213000083.jpg53 is no greater than the desired representation accuracy D.
[0094] Therefore, the representation accuracy (direction representation accuracy) value D represents the worst case, which is illustrated diagrammatically in Figure 4. The sound radiation pattern G is JPEG2025183213000084.jpg54 has a non-zero value and zero for all other directions: JPEG2025183213000085.jpg518, where directional radiation pattern with orientation representation accuracy D (e.g., expressed in degrees) JPEG2025183213000086.jpg53 represents a cone 420 with radius D, 410.
[0095] In some implementations, determining the number N of unit vectors may include using a pre-established functional relationship between the representation accuracy D and the corresponding number N of unit vectors, which are distributed on the surface of the 3D sphere by a predetermined placement algorithm, to determine the first set of first directional unit vectors (e.g., The directions indicated by JPEG2025183213000087.jpg53) are approximated with a maximum representation accuracy of D.
[0096] For example, such a functional relationship can be obtained by a brute force method, for example, by repeatedly distributing different numbers N of directional unit vectors on a surface and determining the resulting representation accuracy, as illustrated in the graph of FIG. 5 (circular marker 510) for the placement algorithm described above with reference to equations (2)-(4). JPEG2025183213000088.jpg53 and JPEG2025183213000089.jpg53, which can be approximated using a linear function (continuous line 520 in FIG. 5).
number
[0097] Thus, in this example, the minimum number N of quasi-equidistantly distributed points N on a unit sphere required to achieve a desired directional representation accuracy D can be calculated by the following functional relationship N=N(D):
number
[0098] This method has an efficiency range for N<~2000, and the resulting orientation representation accuracy D is close to the subjective directional sensitivity threshold. JPEG2025183213000092.jpg56. Figure 6 illustrates this relationship 610 on a log-log scale. The dashed rectangle in this graph illustrates the efficiency range for N<~2000. The modeled relationship between the number of unit vectors N and representation accuracy D is also illustrated in Table 3 below for selected values.
[0099] [Table 3]
[0100] Step S930 of method 900 (or step S1150 of method 1100) may proceed as follows.
[0101] JPEG2025183213000094.jpg53 of the original data G (e.g., a first set of first directional unit vectors and associated first directional gains) defined above. Directional data approximation defined on JPEG2025183213000095.jpg53 Any approximation (e.g., stereo projection) method can be used to obtain JPEG2025183213000096.jpg53 (e.g., the associated second directivity gain). If this operation is performed on the encoder side (e.g., step S930 of method 900), it does not play a significant role in computational complexity.
[0102] On the other hand, directional data approximation A particularly simple procedure for determining JPEG2025183213000097.jpg53 (e.g., the second directional gain) is to calculate the gain for each directional unit vector For JPEG2025183213000098.jpg53 (e.g., the second directional unit vector), each directional unit vector The directional unit vector with the smallest directional difference for JPEG2025183213000099.jpg53 Directivity gain of JPEG2025183213000100.jpg53 (e.g., the first directional unit vector) JPEG2025183213000101.jpg53( JPEG2025183213000102.jpg53) (for example, the first directivity gain). Selecting the "nearest neighbor" of JPEG2025183213000103.jpg53 can proceed as follows:
number
[0103] Bitstream encoding (eg, in step S940 of method 900) and bitstream decoding (eg, in step S1020 of method 1000) may proceed according to the following considerations.
[0104] The generated bitstream is JPEG2025183213000105.jpg53 generation process (e.g., in step S1030 of method 1000) and the corresponding set of directional gains To control JPEG2025183213000106.jpg59, we need to include the encoding scalar value N.
[0105] Directional Data To transport JPEG2025183213000107.jpg53, there are two possible modes:
[0106] One possible mode (the first mode) is the directional gain The goal is to encode the complete set of JPEG2025183213000108.jpg527, where the bitstream is encoded in the corresponding direction, e.g., according to the order in the bitstream. N gain values assigned to JPEG2025183213000109.jpg53 would contain the complete sequence JPEG2025183213000110.jpg59.
[0107] Another possible mode (the second mode) is to store a partial subset as a bitstream, JPEG2025183213000111.jpg58, JPEG2025183213000112.jpg529, In this case, the bitstream is encoded as JPEG2025183213000113.jpg518, for example, by an explicit index in the bitstream. JPEG2025183213000114.jpg52 transmission (i.e., index in the subset The corresponding direction indicated by the transmission of JPEG2025183213000115.jpg52 1 array N assigned to JPEG2025183213000116.jpg53 subset gain values It will only contain JPEG2025183213000117.jpg58.
[0108] The bitstream size Bs for both possible modes can be estimated as follows: For the first mode, the bitstream size Bs can be estimated as follows:
number
[0109] For the second mode, the bitstream size Bs can be estimated as follows:
number
[0110] To achieve better bitstream coding efficiency for JPEG2025183213000120.jpg511, in some implementations, numerical approximation methods (e.g., curve fitting) can be used. One particular advantage of the present disclosure is that it is possible to apply 1D approximation methods (assuming that the data G is fitted along a 1D spiral path s i (This is because the azimuth-elevation plane is defined as A conventional representation of discrete directional information using uniformly distributed directional unit vectors within JPEG2025183213000121.jpg511 would require applying 2D approximation methods and considering boundary conditions.
[0111] To achieve better bitstream encoding efficiency for JPEG2025183213000122.jpg55, in some implementations, determining the number N of unit vectors may include mapping the number N of unit vectors to one of a set of predetermined numbers, for example, by rounding to the nearest number among the set of predetermined numbers. The predetermined number can then be transmitted to the decoder via a bitstream parameter (e.g., bitstream parameter directivity_precision). In this case, there may be an agreement between the encoder and decoder regarding the relationship between the value of the bitstream parameter and the corresponding number among the predetermined numbers. This agreement may be established, for example, by storing the same lookup table at the encoder and decoder.
[0112] In other words, to achieve better bitstream coding efficiency, For JPEG2025183213000123.jpg53, the optimal binary representation (e.g., It may be advisable to use preselected settings that result in a resolution of 1000 kb (JPEG2025183213000124.jpg55=2 bits) and accuracy D.
[0113] [Table 4]
[0114] An example of bitstream syntax for directivity size transmission is shown in Table 5 below.
[0115] [Table 5]
[0116] An example of a possible bitstream variable definition for directivity size transmission is shown in Table 6 below.
[0117] [Table 6]
[0118] The audio directivity modeling of the method in 6DoF rendering (eg, in step S1040 of method 1000 or step S1160 of method 1100) may proceed as follows.
[0119] For each given object-to-listener relative direction P (target direction vector), the closest direction vector The index corresponding to JPEG2025183213000128.jpg54 JPEG2025183213000129.jpg52 is determined as follows:
number
[0120] The corresponding directional gain is then applied to render the sound source at the listener's position. JPEG2025183213000131.jpg59 is applied to this object signal.
[0121] For convenience of notation and explanation, the radiation pattern of the sound source is assumed to be broadband, constant, and S 2 We have assumed that all of space is covered. However, the present disclosure is equally applicable to spectral frequency-dependent radiation patterns (e.g., by performing the proposed method band-by-band). Furthermore, the present disclosure is equally applicable to time-dependent radiation patterns and radiation patterns that include any subset of directions.
[0122] Furthermore, the concepts and methods described in this disclosure may be specified in a frequency and time varying manner, may be applied directly in the spatial or time domain, may be defined globally or in an object dependent manner, may be hard coded into the audio renderer, or may be specified via a corresponding input interface.
[0123] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on a digital signal processor or microprocessor. The components of may be implemented as hardware and / or as application specific integrated circuits. The signals appearing in the above-described methods and systems may be stored on a medium such as a random access memory or an optical storage medium. They may be transferred over a network such as a radio, satellite, wireless or wired network, e.g., the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer devices used to store and / or render audio signals.
[0124] FIG. 12 schematically illustrates an example apparatus 1200 (e.g., an encoder) for encoding audio content, according to an embodiment of the present disclosure. The apparatus 1200 may include an interface system 1210 and a control system 1220. The interface system 1210 may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces. The control system 1220 may include at least one of a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components. Thus, in some implementations, the control system 1220 may include one or more processors and one or more non-transitory storage media operably connected to the one or more processors.
[0125] According to some such examples, the control system 1220 may be configured to receive, via the interface system 120, audio content to be processed / encoded. The control system 1220 may be further configured to: determine a number of unit vectors to arrange on the surface of the 3D sphere as a count number based on a desired representation accuracy (e.g., as in step S910 above); generate a second set of second directional unit vectors by using a predetermined arrangement algorithm to distribute the determined number of unit vectors on the surface of the 3D sphere (where the predetermined arrangement algorithm is an algorithm for approximately uniformly spherically distributing the unit vectors on the surface of the 3D sphere) (e.g., as in step S920 above); determine, for each second directional unit vector, an associated second directional gain based on the first directional gain of one or more first directional unit vectors closest to the respective second directional unit vector in the group of first directional unit vectors (e.g., as in step S930 above); and encode the determined number together with the second directional gain in the bitstream (e.g., as in step S940 above). The control system 1220 may further be configured to output the bitstream (eg, as in step S950 above) via the interface system.
[0126] FIG. 13 schematically illustrates an example apparatus 1300 (e.g., a decoder) for decoding audio content according to an embodiment of the present disclosure. The apparatus 1300 may include an interface system 1310 and a control system 1320. The interface system 1310 may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces. The control system 1320 may include at least one of a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components. Thus, in some implementations, the control system 1320 may include one or more The system may include a processor and one or more non-transitory storage media operably connected to the one or more processors.
[0127] According to some such examples, control system 1320 may be configured to receive a bitstream including audio content via interface system 1310. Control system 1320 may be further configured to extract the number and directional gain from the bitstream (e.g., as in step S1010 above), generate a set of directional unit vectors by using a predetermined placement algorithm to distribute the number of unit vectors on the surface of a 3D sphere (e.g., as in step S1020 above), and, for a given target directional unit vector pointing in a direction from a sound source to a listener position, determine a target directional gain for the target directional unit vector based on associated directional gains of one or more directional unit vectors in the group of directional unit vectors that are closest to the target directional unit vector (e.g., as in step S1030 above).
[0128] Also, according to some such examples, control system 1320 may be configured to receive a bitstream including audio content via interface system 1310 (e.g., as in step S111 above). Control system 1320 may include steps such as: extracting a first set of directional vectors and associated first directional gains from the bitstream (e.g., as in step S1120 above); determining a number of unit vectors to place on the surface of a 3D sphere as a count number based on a desired representation accuracy (e.g., as in step S1130 above); generating a second set of second directional unit vectors by using a predetermined placement algorithm to distribute the determined number of unit vectors on the surface of the 3D sphere (where the predetermined placement algorithm is an algorithm for approximately uniformly distributing the unit vectors on the surface of the 3D sphere) (e.g., as in step S1140 above); The system may be further configured to: for a directional unit vector, determine an associated second directional gain based on a first directional gain of one or more first directional unit vectors of the group of first directional unit vectors that are closest to the respective second directional unit vector (e.g., as in step S1150 above); and for a given target directional unit vector pointing in a direction from the sound source to the listener position, determine a target directional gain for the target directional unit vector based on an associated second directional gain of a second directional unit vector of the group of second directional unit vectors that is closest to the target directional unit vector (e.g., as in step S1160 above).
[0129] In some examples, either or each of the above apparatuses 1200 and 1300 may be implemented within a single device. However, in some implementations, an apparatus may be implemented within more than one device. In some such implementations, the functionality of the control system may be included within more than one device. In some examples, an apparatus may be a component of another device.
[0130] Unless otherwise indicated, and as will be apparent from the following description, throughout this disclosure, references using terms such as "processing," "operating," "calculating," "determining," "analyzing," and the like will be understood to refer to the operations and / or processing of a computer or computing system or similar electronic computing device that manipulates and / or converts data represented as physical quantities, such as quantities of electrons, into other data also represented in physical quantities.
[0131] Similarly, the term "processor" may refer to any device or part of a device that processes electronic data, for example from registers and / or memory, and transforms the electronic data into other electronic data (which may, for example, be stored in registers and / or memory). A "computer" or "computing machine" or "computing platform" may include one or more processors.
[0132] In one exemplary embodiment, the methodologies described herein are executable by one or more processors that accept computer-readable (machine-readable) code, including a set of instructions that, when executed by the one or more processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify operations to be performed is included. Thus, one example is a typical processing system including one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem including main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. The processing system may also be a distributed processing system with processors connected by a network. If the processing system requires a display, such a display may be, for example, a liquid crystal display (LCD) or a cathode ray tube display (CRT). If manual data entry is required, the processing system may also include input devices, such as one or more of an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, and the like. The processing system may also include a storage system, such as a disk drive unit. In some configurations, the processing system may include a sound output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium that holds computer-readable code (e.g., software) that includes a set of instructions that, when executed by one or more processors, cause the processor to perform one or more of the methods described herein. It should be noted that, when the method includes several elements, e.g., several steps, no ordering of such elements is implied unless specifically stated. The software may reside on a hard disk, or may reside, completely or at least partially, in RAM and / or within the processor during execution by the computer system.Thus, the memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may form or be included in a computer program product.
[0133] In alternative exemplary embodiments, one or more processors may operate as standalone devices or may be connected (e.g., connected via a network) to other processors in a networked deployment, one or more processors may operate in the capacity of a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. The one or more processors may form a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify operations to be performed by the machine.
[0134] Also, it should be noted that the term "machine" is intended to include any collection of machines that individually or collectively execute a set of instructions (or multiple sets) for performing any one or more of the methodologies described herein.
[0135] Thus, one exemplary embodiment of each of the methods described herein may include a computer program holding a set of instructions, e.g., a computer program, for execution on one or more processors, e.g., one or more processors that are part of a web server configuration. The present disclosure may be implemented in the form of a computer-readable carrier medium. Accordingly, as will be appreciated by those skilled in the art, exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a dedicated device, an apparatus such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code including a set of instructions that, when executed on one or more processors, cause the processors to implement the method. Accordingly, aspects of the present disclosure may take the form of a method, an entirely hardware exemplary embodiment, an entirely software exemplary embodiment, or an exemplary embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.
[0136] The software may also be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in the illustrated embodiment, the term "carrier medium" should be interpreted to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more sets of instructions. The term "carrier medium" should also be interpreted to include any medium that can store, encode, or retain sets of instructions for execution by one or more processors, causing the one or more processors to perform any one or more of the methodologies of the present disclosure. Carrier media may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus subsystem. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. Thus, for example, the term "carrier medium" should be taken to include, but is not limited to, computer products embodied in solid-state memories, optical and magnetic media, media carrying propagated signals detectable by at least one processor or more processors and representing a set of instructions that when executed implements a method, and transmission media within a network carrying propagated signals detectable by at least one processor of one or more processors and representing a set of instructions.
[0137] It will be understood that the steps of the described methods are performed, in one example embodiment, by a suitable processor (or processors) of a processing (e.g., computer) system executing instructions (e.g., computer-readable code) stored in storage. It will also be understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented using any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.
[0138] Throughout this disclosure, references to "one exemplary embodiment," "some exemplary embodiments," or "exemplary embodiments" mean that a particular feature, structure, or characteristic described in connection with an exemplary embodiment is included in at least one exemplary embodiment of the disclosure. Thus, the appearances of the phrases "in one exemplary embodiment," "in some exemplary embodiments," or "in exemplary embodiments" in various places throughout this disclosure do not necessarily all refer to the same exemplary embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more exemplary embodiments.
[0139] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object merely indicates that different instances of a similar object are being referred to and is not intended to imply that the objects so described must be arranged in a given sequence, in time, space, order, or in any other manner.
[0140] In the following claims and in this specification, the terms "comprise" or "comprising" are both open terms meaning that the terms include at least the elements / features in question, but do not exclude others. Therefore, when used in the claims, the terms "comprise" or "comprising" should not be interpreted as limiting the means or elements or steps listed thereto. For example, the scope of the expression "device comprising A and B" should not be limited to a device consisting of only elements A and B. As used in this specification, the terms "comprise" or "comprised" are also both open terms meaning that the terms include at least the elements / features in question, but do not exclude others. Therefore, "comprise" is synonymous with and means "comprise."
[0141] In the foregoing description of exemplary embodiments of the present disclosure, it should be understood that various features of the present disclosure may be grouped together in a single exemplary embodiment, figure, or description thereof to simplify the disclosure and facilitate understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, aspects of the present invention may reside in fewer than all features of a single foregoing disclosed exemplary embodiment. Accordingly, the claims that follow this specification are expressly incorporated herein, with each claim standing on its own as a separate exemplary embodiment of the present disclosure.
[0142] Furthermore, it will be understood by those skilled in the art that some exemplary embodiments described herein may include some features included in other exemplary embodiments but not others, meaning that combinations of features of different exemplary embodiments are within the scope of the present disclosure and form different exemplary embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0143] In the description provided herein, many specific details are set forth. However, it will be understood that example embodiments of the present disclosure may be practiced without these specific details. In other instances, details of well-known methods, structures and techniques have been omitted so as not to obscure an understanding of this description.
[0144] Thus, while what is believed to be the best mode of the disclosure has been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the disclosure, and that all such variations and modifications are intended to be claimed within the scope of the disclosure. For example, any formulas given above are merely representative of procedures that may be used. Functions may be added to or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure.
[0145] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). 1. A method for processing audio content including directional information for at least one sound source, the directional information including a first set of first directional unit vectors representing directional directions and associated first directional gains, the method comprising: determining, as a count, a number of unit vectors to place on the surface of the 3D sphere, the number of unit vectors being related to a desired representation accuracy; generating a second set of second directional unit vectors by using a predetermined placement algorithm to distribute the determined number of unit vectors over the surface of the 3D sphere, wherein the predetermined placement algorithm is an algorithm for approximately uniformly distributing the unit vectors in a spherical fashion over the surface of the 3D sphere; determining, for each of the second directional unit vectors, an associated second directional gain based on a first directional gain of one or more first directional unit vectors of a group of first directional unit vectors that are closest to the respective second directional unit vector; The method includes: 2. A method as described in EEE1, wherein the number of unit vectors is determined such that when the unit vectors are distributed over the surface of the 3D sphere by the predetermined placement algorithm, they approximate the direction indicated by a first directional unit vector of the first set with at most the desired representation accuracy. 3. The method according to EEE1 or 2, wherein the number of unit vectors is determined such that, for each of the first directional unit vectors in the first set, there is at least one unit vector among the unit vectors whose directional difference with respect to the respective first directional unit vector is less than the desired representation accuracy when the unit vectors are distributed on the surface of the 3D sphere by the predetermined placement algorithm. 4. A method as described in any one of the preceding EEE, wherein determining the number of unit vectors comprises using a pre-established functional relationship between representation precision and the number of corresponding unit vectors distributed on the surface of the 3D sphere by the predetermined placement algorithm and approximating the direction indicated by a first directional unit vector of the first set at at most the respective representation precision. 5. A method according to any one of the preceding paragraphs EEE, wherein the step of determining an associated second directional gain for a given second directional unit vector comprises: setting the second directional gain to the first directional gain associated with the first directional unit vector that is closest to the given second directional unit vector. A method comprising: 6. A method according to any one of the preceding EEE, wherein the predetermined placement algorithm comprises superimposing on the surface of the 3D sphere a spiral path extending from a first point on the 3D sphere to a second point on the 3D sphere opposite the first point, and placing the unit vectors successively along the spiral path; determining a spacing of the spiral path and an offset between each two adjacent unit vectors along the spiral path based on the number of unit vectors; method. 7. A method as described in any one of the preceding EEE, wherein determining the number of unit vectors further comprises mapping the number of unit vectors to one of a predetermined number, wherein the predetermined number can be signaled by a bitstream parameter. 8. A method according to any one of the preceding EEEs, wherein the desired representational accuracy is determined based on a model of the perceptual directional sensitivity threshold of a human listener. 9. A method according to any one of the preceding claims, wherein the cardinality of the second directional unit vectors of the second set is less than the cardinality of the first directional unit vectors of the first set. 10. A method according to any one of the preceding EEE, wherein the first and second directional unit vectors are expressed in a spherical coordinate system or a Cartesian coordinate system. 11. The method of any one of the preceding EEE, wherein the first of the first set and the associated first directivity gain is stored in SOFA format; and / or the directional information represented by the second set of first directional unit vectors and associated second directional gains is stored in a SOFA format; method. 12. A method according to any one of the preceding EEE, comprising: encoding the audio content; encoding the determined number of unit vectors together with the second directivity gain into a bitstream; outputting the bitstream; The method further comprises: 13. A method of decoding audio content comprising directional information for at least one sound source, the directional information comprising a count indicating a number of unit vectors approximately uniformly distributed on a surface of a 3D sphere and, for each such unit vector, an associated directional gain, the unit vectors being assumed to be distributed on the surface of the 3D sphere by a predetermined placement algorithm, the predetermined placement algorithm being an algorithm for approximately uniformly spherically distributing the unit vectors on the surface of the 3D sphere, the method comprising: receiving a bitstream containing the audio content; extracting the number and the directional gain from the bitstream; generating a set of directional unit vectors by using the predetermined placement algorithm to distribute the number of unit vectors over the surface of the 3D sphere; The method includes: 14. A method according to any preceding EEE, comprising: for a given target directional unit vector pointing in a direction from the sound source to a listener position, determining a target directional gain for the target directional unit vector based on associated directional gains of one or more directional unit vectors of a group of directional unit vectors that are closest to the target directional unit vector. A method that further encompasses 15. A method according to any preceding claim, wherein determining the target pointing gain for the target pointing unit vector comprises: setting the target directional gain to the directional gain associated with the directional unit vector that is closest to the target directional unit vector. The method includes: 16. A method of decoding audio content including directional information for at least one sound source, the directional information including a first set of first directional unit vectors representing directional directions and associated first directional gains, the method comprising: receiving a bitstream containing the audio content; extracting the first set of directional unit vectors and the associated first directional gains from the bitstream; determining, as a count, a number of unit vectors to place on the surface of the 3D sphere, the number of unit vectors being related to a desired representation accuracy; generating a second set of second directional unit vectors by using a predetermined placement algorithm to distribute the determined number of unit vectors over the surface of the 3D sphere, wherein the predetermined placement algorithm is an algorithm for approximately uniformly distributing the unit vectors in a spherical fashion over the surface of the 3D sphere; determining, for each of the second directional unit vectors, an associated second directional gain based on the first directional gain of one or more first directional unit vectors of a group of first directional unit vectors that are closest to the respective second directional unit vector; and, for a given target directivity unit vector pointing in a direction from the sound source to a listener position, determining a target directivity gain for the target directivity unit vector based on associated second directivity gains of one or more second directivity unit vectors of a group of second directivity unit vectors that are closest to the target directivity unit vector; The method includes: 17. The method of claim 16, wherein determining the target pointing gain for the target pointing unit vector comprises: setting the target directional gain to the second directional gain associated with the second directional unit vector that is closest to the target directional unit vector. The method includes: The method described in 18.EEE16 is extracting from the bitstream an indication of whether the second set of directional unit vectors should be generated; if the indication indicates that the second set of directional unit vectors should be generated, determining the number of unit vectors and generating second directional unit vectors of the second set; The method further comprises: 19. An apparatus for processing audio content including directional information for at least one sound source, the directional information including a first set of first directional unit vectors representing directional directions and associated first directional gains, the apparatus comprising a processor configured to perform the steps of the method according to any one of EEE1-12. Device. 20. An apparatus for decoding audio content comprising directivity information for at least one sound source, the directivity information comprising a count indicating a number of unit vectors approximately uniformly distributed on a surface of a 3D sphere and, for each such unit vector, an associated directivity gain, the unit vectors being assumed to be distributed on the surface of the 3D sphere by a predetermined positioning algorithm, the predetermined positioning algorithm being an algorithm for approximately uniformly spherically distributing the unit vectors on the surface of the 3D sphere, the apparatus comprising a processor configured to perform the steps of the method according to any one of EEE13 to EEE15. Device. 21. An apparatus for decoding audio content comprising directional information for at least one sound source, the directional information comprising a first set of first directional unit vectors representing directional directions and associated first directional gains, the apparatus comprising a processor configured to perform the steps of the methods described in any one of EEE16 to EEE18. 22. A computer program comprising instructions which, when executed by a processor, cause said processor to perform a method as set forth in any one of EEE1 to EEE18. 23. A computer-readable medium storing a computer program according to EEE22.
Claims
[Claim 1] 1. A method for decoding audio content comprising directional information for at least one sound source of the audio content, comprising: receiving a bitstream containing the audio content; extracting a first parameter indicative of the number of unit vectors approximately uniformly distributed on the surface of a 3D sphere, and for each unit vector a second parameter indicative of the associated directivity gain; generating a set of directional unit vectors based on the first parameters, the second parameters, and an algorithm for approximately uniformly distributing the unit vectors on the surface of the 3D sphere; A method comprising:
Citation Information
Patent Citations
Adaptive audio content generation
JP2016526828A
Inserting audio channels into sound field descriptions
JP2017513053A
Fast and memory efficient encoding of sound objects using spherical harmonic symmetries
US20190069110A1
Apparatus and method for encoding or decoding directional audio coding parameters using different time / frequency resolutions
WO2019097017A1