Apparatus and method for encoding or decoding AR / VR metadata using a universal codebook
The apparatus and method enhance immersive audio experiences in augmented and virtual reality by efficiently encoding and decoding audio information about acoustic environments, addressing the lack of efficient methods in existing technologies.
Patent Information
- Application Number
- JP2025501395
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-12
- Filing Date
- 2023-07-12
- Publication Date
- 2025-09-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing audio encoding and decoding technologies for augmented and virtual reality lack efficient methods to provide additional audio information about real or virtual acoustic environments, such as reverberation, which is crucial for creating immersive experiences.
An apparatus and method for encoding and decoding audio signals using a universal codebook that includes entropy decoding and encoding modules to efficiently handle additional audio information, such as reflection and diffraction data, allowing for the generation of immersive audio experiences.
Enhances the immersive audio experience in augmented and virtual reality by efficiently encoding and decoding audio information about acoustic environments, including reflections and diffractions, supporting six degrees of freedom and providing realistic sound propagation.
Smart Images

Figure 2025528680000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for encoding or decoding, and more particularly to an apparatus and method for encoding or decoding augmented reality (AR) or virtual reality (VR) metadata using a universal codebook. [Background technology]
[0002] Further improvement and development of audio coding techniques is a continuing task of audio coding research, with the aim of creating a realistic audio experience for the listener, for example in augmented reality or virtual reality scenarios that take into account audio effects such as reverberation caused by reflections on objects, walls, etc., while at the same time encoding and decoding audio information with high efficiency. Summary of the Invention [Problem to be solved by the invention]
[0003] One of these new audio technologies, for example, that aims to create an improved listening experience for augmented or virtual reality is MPEG-I. MPEG-I is a new and developing standard for virtual and augmented reality applications. It aims to create AR or VR experiences that are natural, realistic, and provide an overall compelling experience for the ears as well as the eyes.
[0004] For example, MPEG-I technology could be used to enable listeners to move freely around a concert venue when listening to a concert in VR, rather than being routed to just one spot. Alternatively, MPEG-I technology could be employed in broadcasts of e-sports or sporting events, where users can move around the stadium while watching the game.
[0005] Previous solutions allow for visual or audio experiences from a single observation point, in what is known as three degrees of freedom (3DoF). In contrast, the upcoming MPEG-I standard supports a full six degrees of freedom (6DoF). In 3DoF, users can freely move their head and receive input from multiple planes. However, in 6DoF, users can move within the virtual space. They can walk around, explore all viewing angles, and even interact with the virtual world. MPEG-I technology is equally applicable to augmented reality (AR), where users act within a real world augmented by virtual elements. For example, they can place several virtual musicians in their living room and enjoy their own personal concert.
[0006] To this end, MPEG-I provides advanced techniques for creating convincing, highly immersive audio experiences, including consideration of many aspects of acoustics. One example is sound propagation within a room and around obstacles. Another is sound sources, which can be stationary or moving, with moving sources producing the Doppler effect. Sound propagation is assumed to have realistic radiation patterns and sizes. For example, MPEG-I technology aims to take into account the diffraction of sound around obstacles or room corners and provide efficient rendering of these effects.
[0007] Overall, MPEG-I aims to provide a long-term stable format for rich VR and AR content. Playback using MPEG-I will be possible on both dedicated receiving devices and everyday smartphones. MPEG-I aims to deliver VR and AR content as next-generation video services through existing distribution channels, allowing providers to offer users truly exciting and immersive experiences with entertainment, documentary, educational, or sports content.
[0008] It may be desirable to provide additional audio information, such as information about a real or virtual acoustic environment and / or their effects such as reverberation, to a decoder, e.g., as additional audio information. Providing such information in an efficient manner would be highly appreciated.
[0009] To summarize the above, it would be appreciated if improved concepts for audio encoding and decoding were provided. [Means for solving the problem]
[0010] The object of the present invention is to provide an improved concept for audio encoding and decoding. The object of the present invention is solved by the subject matter of the independent claims. Particular embodiments are provided in the dependent claims.
[0011] According to one embodiment, there is provided an apparatus for generating one or more audio output signals from one or more encoded audio signals, the apparatus comprising at least one entropy decoding module for decoding the encoded additional audio information to obtain decoded additional audio information if the encoded additional audio information is entropy coded, and further comprising a signal processor for generating the one or more audio output signals in response to the one or more encoded audio signals and in response to the decoded additional audio information.
[0012] Further, according to an embodiment, there is provided an apparatus for encoding one or more audio signals and additional audio information, the apparatus comprising: an audio signal encoder for encoding the one or more audio signals to obtain one or more encoded audio signals, and at least one entropy coding module for encoding the additional audio information using entropy coding to obtain the encoded additional audio information.
[0013] Further, according to one embodiment, there is provided an apparatus for generating one or more audio output signals from one or more encoded audio signals. The apparatus comprises an input interface for receiving the one or more encoded audio signals and for receiving additional audio information data. The apparatus further comprises a signal generator for generating the one or more audio output signals in response to the encoded audio signals and in response to second additional audio information. The signal generator is configured to use the additional audio information data and to obtain the second additional audio information using the first additional audio information if the additional audio information data indicates a redundant state. The signal generator is further configured to obtain the second additional audio information using the additional audio information data without using the first additional audio information if the additional audio information data indicates a non-redundant state.
[0014] Further, an apparatus for encoding one or more audio signals and generating additional audio information data according to an embodiment is provided. The apparatus includes an audio signal encoder for encoding one or more audio signals to obtain one or more encoded audio signals. The apparatus further includes an additional audio information generator for generating the additional audio information data, the additional audio information generator indicating a non-redundant operation mode and a redundant operation mode. The additional audio information generator is configured to generate the additional audio information data such that the additional audio information data includes second additional audio information when the additional audio information generator indicates the non-redundant operation mode. The additional audio information generator is further configured to generate the additional audio information data such that the additional audio information data does not include the second additional audio information or includes only a portion of the second additional audio information when the additional audio information generator indicates the non-redundant operation mode, so that the second additional audio information can be obtained using the additional audio information data together with the first additional audio information.
[0015] Further, according to one embodiment, there is provided a method for generating one or more audio output signals from one or more encoded audio signals, the method comprising: - decoding the encoded additional audio information to obtain decoded additional audio information, if the encoded additional audio information is entropy coded; generating one or more audio output signals in response to the one or more encoded audio signals and in response to the decoded additional audio information.
[0016] Further, according to an embodiment, a method for encoding one or more audio signals and additional audio information is provided, the method comprising: encoding one or more audio signals into one or more encoded audio signals; encoding the additional audio information using entropy coding to obtain encoded additional audio information.
[0017] Further, according to another embodiment, there is provided a method for generating one or more audio output signals from one or more encoded audio signals, the method comprising: receiving one or more encoded audio signals and receiving additional audio information data; generating one or more audio output signals in response to the encoded audio signal and in response to the second additional audio information.
[0018] The method includes using the additional audio information data and using the first additional audio information to obtain the second additional audio information if the additional audio information data indicates a redundant state, and further including using the additional audio information data and without using the first additional audio information to obtain the second additional audio information if the additional audio information data indicates a non-redundant state.
[0019] Further, according to an embodiment, a method for encoding one or more audio signals and for generating additional audio information data is provided, the method comprising: encoding one or more audio signals to obtain one or more encoded audio signals; and generating additional audio information data.
[0020] In the non-redundant mode of operation, the generation of the additional audio information data is such that the additional audio information data comprises second additional audio information, and in the redundant mode of operation, the generation of the additional audio information data is such that the additional audio information data does not comprise the second additional audio information or only comprises a part of the second additional audio information, such that the second additional audio information is obtainable using the additional audio information data together with the first additional audio information.
[0021] Furthermore, computer programs are provided, each computer program being configured to perform one of the methods described above when run on a computer or signal processor.
[0022] In the following, embodiments of the invention will be explained in more detail with reference to the drawings. [Brief explanation of the drawings]
[0023] [Figure 1] 1 illustrates an apparatus for generating one or more audio output signals from one or more encoded audio signals according to one embodiment; [Figure 2] FIG. 10 illustrates an apparatus for generating one or more audio output signals according to another embodiment, further comprising at least one non-entropy decoding module and a selector. [Figure 3]FIG. 10 illustrates an apparatus for generating one or more audio output signals according to a further embodiment, the apparatus comprising a non-entropy decoding module, a Huffman decoding module, and an arithmetic decoding module. [Figure 4] 1 illustrates an apparatus for encoding one or more audio signals and additional audio information according to one embodiment; [Figure 5] 4 shows an apparatus for encoding one or more audio signals and additional audio information according to another embodiment, comprising at least one non-entropy encoding module and a selector; [Figure 6] FIG. 10 illustrates an apparatus for generating one or more audio output signals according to a further embodiment, the apparatus comprising a non-entropy coding module, a Huffman coding module, and an arithmetic coding module. [Figure 7] FIG. 1 illustrates a system according to one embodiment. [Figure 8] FIG. 2 illustrates a particular embodiment illustrating encoding of additional audio data and decoding of the encoded additional audio data. [Figure 9] FIG. 10 illustrates an apparatus for generating one or more audio output signals from one or more encoded audio signals according to another embodiment. [Figure 10] 1 illustrates an apparatus for encoding one or more audio signals and for generating additional audio information data according to one embodiment; [Figure 11] FIG. 1 illustrates a system according to another embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0024] FIG. 1 illustrates an apparatus 100 for generating one or more audio output signals from one or more encoded audio signals according to one embodiment.
[0025] The apparatus 100 comprises at least one entropy decoding module 110 for decoding the coded additional audio information to obtain decoded additional audio information, if the coded additional audio information is entropy coded.
[0026] Furthermore, the apparatus 100 comprises a signal processor 120 for generating one or more audio output signals in response to the one or more encoded audio signals and in response to the decoded additional audio information.
[0027] FIG. 2 shows an apparatus 100 for generating one or more audio output signals according to another embodiment, in which, compared to the apparatus 100 of FIG. 1, the apparatus 100 of FIG. 2 further comprises at least one non-entropy decoding module 111 and a selector 115.
[0028] The at least one non-entropy decoding module 111 may be configured to decode the encoded additional audio information, for example, to obtain the decoded additional audio information, if the encoded additional audio information is not entropy coded.
[0029] The selector 115 may be configured to select one of the at least one entropy decoding module 110 and the at least one non-entropy decoding module 111 for decoding the encoded additional audio information, depending on, for example, whether the encoded additional audio information is entropy coded or not.
[0030] According to one embodiment, the encoded additional audio information may include, for example, augmented reality or virtual reality data.
[0031] In one embodiment, the encoded additional audio information is dependent on a real listening environment, or a virtual listening environment, or an augmented listening environment.
[0032] In a typical application scenario, the listening environment is assumed to be modeled and coded at the encoder side, and the modeling of the listening environment is assumed to be received at the decoder side.
[0033] Typical additional audio information about the listening environment can be, for example, information about a number of reflective objects from which sound waves can be reflected. Generally, reflective objects relevant for reflection are those that have an extension (significantly) greater than the wavelength of audible sound. Therefore, when considering reflections, walls or other large reflective objects are particularly important. Such reflective objects can be appropriately represented, for example, by surfaces from which sound is reflected.
[0034] In a three-dimensional environment, a surface may be characterized, for example, by three points in a three-dimensional coordinate system, and each of these three points may be defined, for example, by its x-coordinate value, its y-coordinate value, and its z-coordinate value. Thus, for each of the three points, three x-, y-, and z-values are required, thus a total of nine coordinate values are required to define the surface.
[0035] A more efficient representation of a surface would be, for example, its normal vector This can be achieved by defining the surface using TIFF2025528680000002.tif55 and by using a scalar distance value d that defines the distance from the defined origin to the surface. If TIFF2025528680000003.tif55 is defined by azimuth and elevation angles (the normal vector has length 1, so it does not need to be encoded), then the surface can be represented by three values: a scalar distance value d of the surface, and a normal vector d of the surface. It can only be defined by the azimuth and elevation angles of TIFF2025528680000004.tif55.
[0036] Typically, for efficient coding, the azimuth and elevation angles can be appropriately quantized. For example, each azimuth angle can be expressed as 2 nThe elevation angle may have one of two different azimuth angle values, for example, each elevation angle may have one of two different azimuth angle values. n-1 The elevation angle may be encoded to have one of several different elevation angle values.
[0037] As outlined above, the representation of walls plays an important role when defining a reflection-focused listening environment. This is true for indoor scenarios where interior walls play a very important role, for example with regard to early reflections. However, this also applies to outdoor scenarios where building walls represent the majority of the relevant reflective objects.
[0038] In a normal environment, it is observed that many walls stand at angles of about 90° to each other. For example, in an indoor scenario, there are many horizontal and vertical walls. Due to structural deviations, the relationship between walls is not always exactly 90°, and may be, for example, 89.8°, 89.6°, 90.3°, etc., but it has been found that there still exists a significant proportion of walls that have a relationship of about 90° and about 0° to each other.
[0039] For example, the elevation angle of a wall may be defined to be, for example, 0° if the wall is a horizontal wall, and may be defined to be, for example, 90° if the wall surface is a vertical wall. Then, in a real-world example, there will be a significant percentage of walls that have an elevation angle of about 90° (e.g., 89.8°, 89.7°, 90.2°) and a significant percentage of walls that have an elevation angle of about 0° (e.g., 0.3°, -0.2°, 0.4°).
[0040] The same observations about elevation angles often also apply to azimuth angles, since rooms often have a rectangular shape.
[0041] However, returning to the example of elevation angle, it should be noted that if the 0° value of elevation angle is defined differently than described above, other values will result that a typical wall will exhibit. For example, if a surface is defined to have an elevation angle of 0°, and is inclined 20° relative to the horizontal, many real-world walls may have, for example, an elevation angle of approximately −20° (e.g., −19.8°, −20.0°, −20.2°), and many real-world walls may have, for example, an elevation angle of approximately 70° (e.g., 69.8°, 70.0°, 70.2°). Nevertheless, a significant proportion of walls will have the same elevation angle at certain elevation angles (approximately −20° and approximately 70° in this example). The same is true for azimuth angle.
[0042] Additionally, some other walls have other specific typical elevation angles, for example, roofs are typically sloped at 45° or 35° or 30°, and certain frequencies of these values occur in real-world examples.
[0043] Furthermore, it should be noted that not all real-world rooms have a rectangular ground plane shape, but may exhibit other regular shapes, for example. Consider, for example, a room with an octagonal ground plane shape. However, it is expected that some azimuth angles, such as azimuth angles around 0°, 45°, 90°, and 135°, will occur more frequently than others.
[0044] Furthermore, in outdoor examples, walls often exhibit similar azimuth angles. For example, two parallel walls of one house will exhibit similar azimuth angles, but this may also be relevant for walls of adjacent houses, which are often built in rows with regular and similar ground shapes relative to each other. There, too, the walls of adjacent houses will exhibit similar azimuth values and therefore have similarly oriented reflective walls / surfaces.
[0045] From the above observations, it has been found that it is often particularly suitable to encode and decode additional audio information using entropy coding, which applies in particular to scenarios where the occurrence of certain values among all possible values occurs (significantly) more frequently than other values.
[0046] In particular embodiments, elevation angle values of a surface (e.g., representing a reflective object) may be encoded and decoded, for example, using entropy coding, for example, using Huffman coding, or using arithmetic coding.
[0047] Similarly, in particular embodiments, azimuth angle values of a surface (e.g., representing a reflective object) may be encoded and decoded using, for example, entropy coding, for example, using Huffman coding, or using arithmetic coding.
[0048] The above considerations also apply to other application scenarios: for example, for a given audio source position s and, say, a given listener position l, the reflection sequence may define, for example, a number of one or more surfaces identified by a number of one or more surface indices, which define the surfaces from which sound waves originating from the audio source on a particular propagation path are reflected until they reach (are heard at) the listener position.
[0049] For example, for a source at position s and a listener at position l, a reflection sequence [5,18] defines that, on a particular propagation path, a sound wave from the source at position s is first reflected off a surface with surface index 5, then reflected off a surface with surface index 18, until it finally reaches the listener at position l (audibly so that the listener can still perceive it). A second reflection sequence may, for example, be the reflection sequence [3,12]. A third reflection sequence includes only [5], indicating that, on a particular propagation path, a sound wave from source s is reflected only by surface 5, then audibly reaches the listener at position l. A fourth reflection sequence [3,7] defines that, on a particular propagation path, a sound wave from source s is first reflected off a surface with surface index 3, then reflected off a surface with surface index 7, until it finally reaches the listener audibly. All reflection sequences for a listener at position l and a source at position s together define the set of reflection sequences for the listener at position l and the source at position s.
[0050] However, there may also be other defined surfaces, for example surfaces with surface index 6, 8, 9, 10, 11, or 15, which may be located far from the listener's position l and far from the source's position s. These surfaces will occur less frequently or not at all in the set of reflection sequences for a listener at position l and a source at position s. From this observation, it has been found that it is often desirable to code the set of reflection sequences using entropy coding.
[0051] Furthermore, even when multiple sets of reflection sequences are jointly encoded for multiple different listener positions and / or multiple different source positions, it may still be desirable to employ entropy coding. For example, in a particular listening environment, a user-reachable area may be defined, and it may be assumed that the user never travels through, for example, dense bush or other inaccessible areas. In some application scenarios, sets of reflection sequences for user positions in these inaccessible areas are not provided. Walls in these areas are usually located far away from all defined possible user positions, and therefore will appear less frequently in the multiple sets of reflection sequences. This results in different occurrences of surface indexes in the multiple sets of reflection sequences, and therefore, it is proposed to entropy code these surface indexes in the reflection sets.
[0052] In one embodiment, the actual occurrence of different values of the additional audio information may e.g. be observed and, e.g., based on this observation, either entropy coding or non-entropy coding may e.g. be employed. Using non-entropy coding when the occurrences of different values occur at the same or at least approximately similar frequencies has, among other advantages, the advantage that a predetermined codeword-to-symbol relationship may e.g. be employed which does not need to be transmitted from the encoder to the decoder.
[0053] We return again to a more general example that can be applied to other applications than the one just described.
[0054] According to one embodiment, the encoded additional audio information may, for example, comprise propagation information that depends on one or more propagations of one or more sound waves along one or more propagation paths in a real listening environment or a virtual listening environment or an augmented listening environment.
[0055] In one embodiment, the propagation information may be, for example, reflection information that depends on one or more reflections at one or more reflecting objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, or the virtual listening environment, or the augmented listening environment.
[0056] According to one embodiment, the propagation information may be, for example, diffraction information that depends on one or more diffractions at one or more diffractive objects of one or more sound waves propagating along one or more propagation paths in the real listening environment or the virtual listening environment or the augmented listening environment.
[0057] According to one embodiment, the encoded additional audio information may include data for rendering early reflections. Signal processor 120 may be configured to generate one or more audio output signals in response to the data for rendering early reflections, for example.
[0058] In one embodiment, the signal processor 120 may be configured to generate a binaural signal, for example including two binaural channels, as one or more audio output signals.
[0059] According to one embodiment, the at least one entropy decoding module 110 may comprise a Huffman decoding module 116 for decoding the encoded additional audio information, for example if the encoded additional audio information is Huffman coded.
[0060] In one embodiment, the at least one entropy decoding module 110 may comprise an arithmetic decoding module 118 for decoding the encoded additional audio information, for example, if the encoded additional audio information is arithmetically coded.
[0061] FIG. 3 is a diagram illustrating an apparatus 100 for generating one or more audio output signals according to another embodiment, the apparatus 100 comprising a non-entropy decoding module 111, a Huffman decoding module 116, and an arithmetic decoding module 118.
[0062] The selector 115 may be configured to select, for example, one of the at least one non-entropy decoding module 111, the Huffman decoding module 116, and the arithmetic decoding module 118 to decode the encoded additional audio information.
[0063] According to one embodiment, the at least one non-entropy decoding module 111 may comprise, for example, a fixed length decoding module for decoding the encoded additional audio information if the encoded additional audio information is fixed length coded.
[0064] In one embodiment, the apparatus 100 may be configured to, for example, receive selection information, and the selector 115 may be configured to, for example, select one of the at least one entropy decoding module 110 and the at least one non-entropy decoding module 111 in response to the selection information.
[0065] According to one embodiment, the device 100 may be configured to receive, for example, a codebook or a coding tree on which the encoded additional audio information depends, and at least the entropy decoding module 110 may be configured to decode the encoded additional audio information, for example, using the codebook or using the coding tree.
[0066] In one embodiment, the apparatus 100 may be configured to receive, for example, an encoding of a structure of a coding tree on which the encoded additional audio information depends. At least the entropy decoding module 110 may be configured, for example, to reconstruct multiple codewords of the coding tree according to the structure of the coding tree. Furthermore, at least the entropy decoding module 110 may be configured, for example, to decode the encoded additional audio information using the codewords of the coding tree.
[0067] For example, typical coded information that may be transmitted from an encoder to a decoder may be, for example, an N-element codeword list containing all N codewords of the code, and a symbol list containing all N symbols encoded by the N codewords of the code. A codeword at position p, with 1≦p≦N, in the codeword list may be defined to encode the symbol at position p in the symbol list.
[0068] For example, the contents of the following two lists may be transmitted, each of the symbols representing a surface index that identifies a particular surface:
[0069] [Table 1] However, instead of transmitting a codeword list, according to one embodiment, a representation of a coding tree may be transmitted, e.g., from an encoder, which may be received, e.g., by a decoder, which may be configured, e.g., to construct a codeword list from the received representation of the coding tree.
[0070] For example, each internal node (e.g., except for the root node of the coding tree) may be represented, for example, by a first bit value (e.g., 0), and each leaf node of the coding tree may be represented, for example, by a second bit value (e.g., 1). Considering the codeword list above, we have:
[0071] [Table 2] Traversing the coding tree from the leftmost branch to the rightmost branch, encoding every new interior node as you traverse the coding tree with 0, and encoding every leaf node as you traverse the coding tree with 1, results in an encoding of the coding tree where the codeword above is represented as
[0072] [Table 3] The resulting coding tree representation is 01 1 01 01 01 1. At the decoder side, the representation of the coding tree can be decomposed into a list of codewords.
[0073] Codeword 1: First leaf node arrives at second node: Codeword 1 with bits 00 Codeword 2: Then another leaf node follows: Codeword 2 with bits: 01 Codeword 3: All nodes to the left of the root node are found, go to the right branch of the root node, and the first leaf to the right of the root node is in the second node: Codeword 3 with bits "10" Codeword 4: Move one node up (under the first branch 1); move the internal node (0) down into the right branch (second branch 1); move the leaf node (1) into the left branch (branch 0): Codeword 4: "110". (Leaf node under branch 1-1-0) Codeword 5: Lift one node up (under the second branch 1); Lower the internal node (0) into the right branch (third branch 1); Move the leaf node (1) into the left branch (branch 0): Codeword 5: "1110". (Leaf node under branch 1-1-1-0) Codeword 6: Move one node up and down the right branch (fourth branch 1), which is the leaf node (1): Codeword 6: "1111" (leaf node under branch 1-1-1-1).
[0074] By encoding the coding tree structure instead of the codewords, coding efficiency is improved.
[0075] In one embodiment, the apparatus 100 may further comprise a memory storing, for example, a codebook or a coding tree, and at least the entropy decoding module 110 may be configured to decode the additional audio information that was encoded using, for example, the codebook or using the coding tree.
[0076] According to one embodiment, the apparatus 100 may be configured to receive encoded additional audio information, for example, including a plurality of transmitted symbols and an offset value, and the at least one non-entropy decoding module 111 may be configured to decode the encoded additional audio information, for example, using the plurality of transmitted symbols and using the offset value.
[0077] In one embodiment, the data for rendering the early reflections may include, for example, information regarding the location of one or more walls, which may be one or more real or virtual walls, in the environment. The signal processor 120 may be configured to, for example, generate one or more audio output signals in response to the information regarding the location of the one or more walls.
[0078] According to one embodiment, the information about each of the one or more walls may include, for example, information about the azimuth angle and / or elevation angle of the wall, where the azimuth angle of the wall may, for example, be entropy coded and / or the elevation angle of the wall may, for example, be entropy coded. One or more entropy decoding modules of the at least one entropy decoding module 110 are configured to decode the entropy coded azimuth angle of the wall and / or the entropy coded elevation angle of the wall.
[0079] In one embodiment, one or more of the at least one entropy decoding modules 110 are configured to decode the entropy coded azimuth angles of the walls and / or the entropy coded elevation angles of the walls using a codebook or coding tree.
[0080] According to one embodiment, the encoded additional audio information may include, for example, voxel position information, which may include, for example, information regarding one or more positions of one or more voxels of the plurality of voxels in a three-dimensional coordinate system. The signal processor 120 may be configured, for example, to generate one or more audio output signals in response to the voxel position information.
[0081] In one embodiment, the at least one entropy decoding module 110 may be configured to decode the encoded additional audio information, e.g., the encoded additional audio information being entropy coded, e.g., A list of triangle indices, e.g. earlySurfaceFaceIdx, and The array length of the list of triangle indices, e.g., earlySurfaceFaceIdx, and the array length of earlySurfaceLengthFaceIdx, an array with azimuth angles specifying the surface normals in spherical coordinates (e.g., Hessian normal form), e.g., earlySurfaceAzi; an array with elevation angles specifying the surface normals in spherical coordinates (e.g., Hessian normal form), e.g., earlySurfaceEle; An array with distance values (e.g., in Hessian normal form), e.g., earlySurfaceDist, an array with the listener's position, e.g., an array with the listener voxel index, e.g., earlyVoxelL; an array with one or more sound source locations, e.g., an array with source voxel indices, e.g., earlyVoxelS; a removal list or removal set specifying the set of reflection sequences to be removed or the indexes of the reflection sequences in the reference reflection sequence list to be removed, e.g., a differently coded removal list or a differently coded removal set, e.g., earlyVoxelIndicesRemovedDiff; The number of reflection sequences or reflection paths, e.g., earlyVoxelNumPaths, An array specifying the reflection order, e.g., a two-dimensional array, e.g., earlyVoxelOrder, It may include at least one of a reflection sequence, such as an earlyVoxelSurf.
[0082] FIG. 4 illustrates an apparatus 200 for encoding one or more audio signals and additional audio information according to one embodiment.
[0083] The apparatus 200 comprises an audio signal encoder 210 for encoding one or more audio signals to obtain one or more encoded audio signals.
[0084] Furthermore, the apparatus 200 comprises at least one entropy coding module 220 for encoding the additional audio information using entropy coding to obtain encoded additional audio information.
[0085] 5 illustrates an apparatus 200 for encoding one or more audio signals and additional audio information according to another embodiment. Compared to the apparatus 200 of FIG. 4, the apparatus 200 of FIG. 4 further comprises at least one non-entropy coding module 221 and a selector 215.
[0086] The at least one non-entropy coding module 221 may be configured, for example, to encode the additional audio information to obtain encoded additional audio information.
[0087] The selector 215 may be configured to select one of the at least one entropy encoding module 220 and the at least one non-entropy encoding module 221 for encoding the additional audio information, for example, depending on the symbol distribution within the additional audio information to be encoded.
[0088] According to one embodiment, the encoded additional audio information may include, for example, augmented reality or virtual reality data.
[0089] In one embodiment, the encoded additional audio information is dependent on a real listening environment, or a virtual listening environment, or an augmented listening environment.
[0090] According to one embodiment, the additional audio information may include, for example, propagation information dependent on one or more propagations of one or more sound waves along one or more propagation paths in the real listening environment or the virtual listening environment or the augmented listening environment.
[0091] In one embodiment, the propagation information may be, for example, reflection information that depends on one or more reflections at one or more reflecting objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, or the virtual listening environment, or the augmented listening environment.
[0092] According to one embodiment, the propagation information may be, for example, diffraction information that depends on one or more diffractions at one or more diffractive objects of one or more sound waves propagating along one or more propagation paths in the real listening environment or the virtual listening environment or the augmented listening environment.
[0093] According to one embodiment, the encoded additional audio information may include data for rendering early reflections.
[0094] In one embodiment, the at least one entropy coding module 220 may comprise a Huffman coding module 226 for encoding the additional audio information using, for example, Huffman coding.
[0095] According to one embodiment, the at least one entropy coding module 220 may comprise, for example, an arithmetic coding module 228 for encoding the additional audio information using arithmetic coding.
[0096] FIG. 6 is a diagram illustrating an apparatus 200 for generating one or more audio output signals according to another embodiment, the apparatus 200 comprising a non-entropy coding module 221, a Huffman coding module 226 and an arithmetic coding module 228.
[0097] The selector 215 may be configured to select, for example, one of at least one non-entropy coding module 221, a Huffman coding module 226, and an arithmetic coding module 228 to encode the additional audio information.
[0098] In one embodiment, the at least one non-entropy coding module 221 may comprise, for example, a fixed length coding module for encoding the additional audio information.
[0099] According to one embodiment, the device 200 may be configured to generate selection information indicating, for example, one of the at least one entropy coding module 220 and the at least one non-entropy coding module 221 employed to encode the additional audio information.
[0100] In one embodiment, the apparatus 200 may be configured to transmit, for example, a codebook or coding tree employed to encode the additional audio information.
[0101] In one embodiment, the apparatus 200 may be configured to transmit, for example, an encoding of the structure of a coding tree on which the encoded additional audio information depends.
[0102] According to one embodiment, the apparatus 200 may further comprise a memory storing, for example, a codebook or a coding tree, and at least the entropy coding module 220 may be configured to encode the additional audio information using, for example, a codebook or using a coding tree.
[0103] In one embodiment, the at least one entropy coding module 220 may be configured to encode the additional audio information, such that the encoded additional audio information may comprise, for example, a plurality of transmitted symbols and an offset value.
[0104] According to one embodiment, the data for rendering early reflections may include, for example, information about the position of one or more walls in the environment, which may be one or more real or virtual walls.
[0105] In one embodiment, the information about each wall of the one or more walls may include, for example, information about the azimuth angle and / or elevation angle of the wall, where the azimuth angle of the wall may, for example, be entropy coded and / or the elevation angle of the wall may, for example, be entropy coded. One or more entropy coding modules of the at least one entropy coding module 220 are configured to encode the additional audio information, such that the encoded additional audio information may comprise, for example, an entropy coded azimuth angle of the wall and / or an entropy coded elevation angle of the wall.
[0106] According to one embodiment, the one or more entropy coding modules are configured to code the entropy coded azimuth angles of the walls and / or the entropy coded elevation angles of the walls using a codebook or coding tree.
[0107] In one embodiment, the encoded additional audio information may include, for example, voxel position information, which may include, for example, information regarding one or more positions of one or more voxels of the plurality of voxels within a three-dimensional coordinate system.
[0108] According to one embodiment, the at least one entropy coding module 220 may be configured to encode the additional audio information using, for example, entropy coding, the encoded additional audio information being, for example, A list of triangle indices, e.g. earlySurfaceFaceIdx, and The array length of the list of triangle indices, e.g., earlySurfaceFaceIdx, and the array length of earlySurfaceLengthFaceIdx, an array with azimuth angles specifying the surface normals in spherical coordinates (e.g., Hessian normal form), e.g., earlySurfaceAzi; an array with elevation angles specifying the surface normals in spherical coordinates (e.g., Hessian normal form), e.g., earlySurfaceEle; An array with distance values (e.g., in Hessian normal form), e.g., earlySurfaceDist, an array with the listener's position, e.g., an array with the listener voxel index, e.g., earlyVoxelL; an array with one or more sound source locations, e.g., an array with source voxel indices, e.g., earlyVoxelS; a removal list or removal set specifying the set of reflection sequences to be removed or the indexes of the reflection sequences in the reference reflection sequence list to be removed, e.g., a differently coded removal list or a differently coded removal set, e.g., earlyVoxelIndicesRemovedDiff; The number of reflection sequences or reflection paths, e.g., earlyVoxelNumPaths, An array specifying the reflection order, e.g., a two-dimensional array, e.g., earlyVoxelOrder, It may include at least one of a reflection sequence, such as an earlyVoxelSurf.
[0109] Figure 7 illustrates a system according to one embodiment, comprising the device 200 of Figure 4 for encoding one or more audio signals and additional audio information to obtain one or more encoded audio signals and encoded additional audio information, and further comprising the device 100 of Figure 1 for generating one or more audio output signals from the one or more encoded audio signals in response to the encoded additional audio information.
[0110] FIG. 8 illustrates a specific embodiment showing encoding of additional audio data and decoding of the encoded additional audio data. In FIG. 8, the additional audio data is AR data or VR data that is encoded at an encoder side to obtain encoded AR data or VR data. Metadata may also be encoded. The encoded AR data or encoded VR data is then decoded at a decoder side to obtain decoded AR data or decoded VR data. At the encoder side, a selector operates an encoder switch to select one of N different encoder modules for encoding the AR data or VR data. In FIG. 8, the selector provides information to the decoder side so that a corresponding decoding module among the N decoding modules is selected to decode the encoded AR data or encoded VR data.
[0111] Further embodiments are provided below.
[0112] According to one embodiment, a system for encoding and decoding a data sequence is provided, comprising an encoder subsystem and a decoder subsystem. The encoder subsystem may, for example, include at least two different encoding methods, an encoder selector, and an encoder switch for selecting one of the encoding methods. The encoder subsystem may, for example, transmit the selected selection, encoding parameters of the selected encoder, and data encoded by the selected encoder. The decoder subsystem may, for example, include a corresponding decoder and a decoder switch for selecting one of the decoding methods.
[0113] In one embodiment, the data series may include, for example, AR / VR data.
[0114] According to one embodiment, the data series may include, for example, metadata for rendering early reflections.
[0115] In one embodiment, at least one fixed length encoder / decoder may be used, for example, and at least one variable length encoder / decoder may be used, for example.
[0116] According to one embodiment, one of the variable length encoders / decoders is a Huffman encoder / decoder.
[0117] In one embodiment, the encoding parameters may include, for example, a codebook or a decoding tree.
[0118] According to one embodiment, the encoding parameters may include, for example, an offset value, the combination of which with the transmitted symbols results in a decoded data sequence.
[0119] FIG. 9 illustrates an apparatus 300 for generating one or more audio output signals from one or more encoded audio signals according to another embodiment.
[0120] The device 300 comprises an input interface 310 for receiving one or more encoded audio signals and for receiving additional audio information data.
[0121] Furthermore, the apparatus 300 comprises a signal generator 320 for generating one or more audio output signals in response to the encoded audio signal and in response to the second additional audio information.
[0122] The signal generator 320 is configured to use the additional audio information data when the additional audio information data indicates a redundancy condition, and to use the first additional audio information to obtain the second additional audio information.
[0123] Furthermore, the signal generator 320 is configured to obtain the second additional audio information using the additional audio information data without using the first additional audio information if the additional audio information data indicates a non-redundant state.
[0124] According to one embodiment, the input interface 310 may be configured to receive propagation information data, for example, as additional audio information data. The signal generator 320 may be configured to generate one or more audio output signals depending on the second additional audio information, for example, the second propagation information. Furthermore, the signal generator 320 may be configured to obtain the second propagation information using the propagation information data and the first additional audio information, for example, the first propagation information, when the propagation information data indicates a redundant state. Furthermore, the signal generator 320 may be configured to obtain the second propagation information using the propagation information data without using the first propagation information, for example, when the propagation information data indicates a non-redundant state.
[0125] According to one embodiment, the first propagation information and / or the second propagation information may depend on one or more propagations of one or more sound waves along one or more propagation paths in, for example, a real listening environment or a virtual listening environment or an augmented listening environment.
[0126] In one embodiment, the propagation information data can include, for example, reflection information data and / or diffraction information data. The first propagation information can include, for example, first reflection information and / or first diffraction information. Further, the second propagation information can include, for example, second reflection information and / or second diffraction information.
[0127] According to one embodiment, the input interface 310 may be configured to receive the reflection information data, for example, as propagation information data. The signal generator 320 may be configured to generate one or more audio output signals dependent on the second propagation information, for example, the second reflection information. Furthermore, the signal generator 320 may be configured to use the reflection information data and obtain the second reflection information using the first propagation information, for example, the first reflection information, when the reflection information data indicates a redundant state. Furthermore, the signal generator 320 may be configured to obtain the second reflection information using the reflection information data without using the first reflection information, for example, when the reflection information data indicates a non-redundant state.
[0128] In one embodiment, the first reflection information and / or the second reflection information may, for example, depend on one or more reflections at one or more reflective objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment.
[0129] The first and second reflection information may, for example, comprise a set of reflection sequences as described above. As already outlined, for example, for a given audio source position s and, for example, a given listener position l, the reflection sequence may, for example, define a number of one or more surfaces, identified by a number of one or more surface indices, from which sound waves originating from the audio source on a particular propagation path are reflected until they reach (are heard at) the listener position.
[0130] All these reflection sequences defined for a listener at position l and a source at position s form a set of reflection sequences.
[0131] For example, it is known that the sets of reflection sequences for adjacent listener positions are very similar. Therefore, it is proposed that the encoder encodes only those reflection sequences (e.g., in the reflection information data) that are not included in the similar set of reflection sequences (e.g., in the first reflection information) and indicates only those reflection sequences of the similar set of reflection sequences that are not valid for the current set of reflection sequences. Similarly, each decoder uses the received reduced information (e.g., the reflection information data) to obtain the current set of reflection sequences (e.g., the second reflection information) from the similar set of reflection sequences (e.g., the first reflection information).
[0132] In one embodiment, the input interface 310 may be configured to receive, for example, diffraction information data as propagation information data. The signal generator 320 may be configured to generate one or more audio output signals dependent on the second propagation information, for example, the second diffraction information. Furthermore, the signal generator 320 may be configured to obtain the second diffraction information using the diffraction information data and the first propagation information, for example, the first diffraction information, when the diffraction information data indicates a redundant state. Furthermore, the signal generator 320 may be configured to obtain the second diffraction information using the diffraction information data without using the first diffraction information, for example, when the diffraction information data indicates a non-redundant state.
[0133] According to one embodiment, the first diffraction information and / or the second diffraction information may, for example, depend on one or more diffractions at one or more diffractive objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment.
[0134] For example, the first and second diffraction information may include, for example, a set of diffraction sequences for a listener at position l and a source at position s. A set of diffraction sequences may be defined, for example, similarly to a set of reflection sequences, but for diffracting objects (e.g., objects that cause diffraction) rather than reflecting objects. In many cases, the diffracting object and the reflecting object may be, for example, the same object. When considering these objects as reflecting objects, the surfaces of these objects are considered, while when considering these objects as diffracting objects, the edges of these objects are considered for diffraction.
[0135] According to one embodiment, if the propagation information data indicates a redundant condition, the propagation information data may indicate one or more propagation sequences to be removed from the first propagation information, e.g., a first set of propagation sequences, and / or may indicate one or more propagation sequences to be added to the first set of propagation sequences to obtain the second propagation information, e.g., a second set of propagation sequences. The signal generator 320 may be configured, for example, to use the propagation information data to update the first set of propagation sequences to obtain the second set of propagation sequences.
[0136] In one embodiment, each reflection sequence of the first set of reflection sequences and the second set of reflection sequences may, for example, represent one or more groups of reflective objects or one or more groups of diffractive objects.
[0137] In one embodiment, if the propagation information data indicates a non-redundant state, the propagation information data may include, for example, a second set of propagation sequences, and the signal generator 320 may be configured, for example, to determine the second set of propagation sequences from the propagation information data.
[0138] According to one embodiment, a first set of propagation sequences may be associated with, for example, a first listener position and a first source position. A second set of propagation sequences may be associated with, for example, a second listener position and a second source position. The first listener position may be different from, for example, the second listener position, and / or the first source position may be different from, for example, the second source position.
[0139] In one embodiment, the first set of propagation sequences may be, for example, a first set of reflection sequences. The second set of propagation sequences may be, for example, a second set of reflection sequences. Each reflection sequence in the first set of reflection sequences may include, for example, information about a group of one or more reflecting objects in the reflection sequence, where sound waves emitted by an audio source at a first source location and perceptible to a listener at a first listener location are reflected on their way to a current listener location. Each reflection sequence in the second set of reflection sequences may include, for example, information about a group of one or more reflecting objects in the reflection sequence, where sound waves emitted by an audio source at a second source location and perceptible to a listener at a second listener location are reflected on their way to a current listener location.
[0140] According to one embodiment, the one or more encoded audio signals are associated with audio sources located at source locations in the second set of the reflection sequence. The signal generator 320 may be configured to generate the one or more audio output signals using the one or more encoded audio signals and using the second set of the reflection sequence, for example, such that the one or more audio output signals may include early reflections of sound waves emitted by the audio sources at the source locations in the second set of the reflection sequence.
[0141] In one embodiment, the input interface 310 can be configured to receive, for example, reflection information data as propagation information data. The signal generator 320 can be configured to obtain, for example, multiple sets of reflection sequences, each of which can be associated with, for example, a listener position and a source position. The input interface 310 can be configured to receive, for example, a notification. To determine the second set of reflection sequences, the signal generator 320 can be configured to, for example, use the notification to determine a first listener position and a first source position when the reflection information data indicates a redundant state, and select one of the multiple sets of reflection sequences as the first set of reflection sequences associated with the first listener position and the first source position.
[0142] For example, each reflection sequence of each set of reflection sequences of the multiple sets of reflection sequences may include, for example, information about one or more groups of reflecting objects in the reflection sequence, where sound waves emitted by an audio source at a source location in said set of reflection sequences and perceptible to a listener at a listener location in said set of reflection sequences are reflected on their way to the current listener location.
[0143] According to one embodiment, if the reflection information data indicates a redundant state, the notification may indicate, for example, selecting the first listener position and the first source position such that the first listener position is adjacent to the second listener position and / or such that the first source position is adjacent to the second listener position. If the reflection information data indicates a redundant state, the signal generator 320 may be configured to, for example, determine the first listener position and / or the first source position according to the notification.
[0144] In one embodiment, if the reflection information data indicates a redundant state, the notification may indicate, for example, selecting the first listener position and the first source position such that the first listener position is adjacent to the second listener position and the first source position is the same as the second listener position. The signal generator 320 is configured to determine the first listener position and the first source position according to the notification.
[0145] Alternatively, in one embodiment, if the reflection information data indicates a redundant state, the notification may indicate, for example, selecting the first listener position and the first source position such that the first listener position is identical to the second listener position and the first source position is adjacent to the second listener position. The signal generator 320 may be configured to determine the first listener position and the first source position according to the notification, for example.
[0146] According to one embodiment, a first position and a second position are adjacent in a coordinate system if, in each coordinate direction of the coordinate system, the first position is immediately before or after the second position, or is identical to the second position, and if, in at least one coordinate direction of the coordinate system, the first position and the second position are different from each other.
[0147] In one embodiment, the notification may include, for example: the reflective information data exhibiting a non-redundant state; the reflection information data indicates a first redundancy state, whereby the first listener position and the first source position are selected such that the first source position is identical to the second source position and such that the first listener position is adjacent to the second listener position, where in a first coordinate direction of the coordinate system the first listener position is immediately before the second listener position, and where in a second coordinate direction and a third coordinate direction of the coordinate system the first listener position is identical to the second listener position; the reflection information data indicates a second redundancy state, whereby the first listener position and the first source position are selected such that the first source position is identical to the second source position and such that the first listener position is adjacent to the second listener position, where the first listener position immediately precedes the second listener position in a second coordinate direction of the coordinate system, and where the first listener position is identical to the second listener position in the first coordinate direction and a third coordinate direction of the coordinate system; The reflection information data may indicate a third redundant state, whereby the first listener position and the first source position are selected such that the first source position is identical to the second source position and such that the first listener position is adjacent to the second listener position, where in a third coordinate direction of the coordinate system the first listener position is immediately before the second listener position, and where in the first coordinate direction and the second coordinate direction of the coordinate system the first listener position is identical to the second listener position.
[0148] If the notification indicates the first redundancy state or the second redundancy state or the first redundancy state, the signal generator 320 may be configured to determine, for example, the first listener position and the first source position according to the notification.
[0149] According to one embodiment, each of the first listener position, the first source position, the second listener position, and the second source position may, for example, define the position of a voxel among a plurality of voxels in a three-dimensional coordinate system.
[0150] For example, each of the listener positions and source positions for each of the multiple sets of reflection sequences may define the position of a voxel among multiple voxels in a three-dimensional coordinate system, for example.
[0151] In one embodiment, the signal generator 320 may be configured to generate a binaural signal, for example including two binaural channels, as one or more audio output signals.
[0152] FIG. 10 illustrates an apparatus 400 for encoding one or more audio signals and for generating additional audio information data according to one embodiment.
[0153] The apparatus 400 comprises an audio signal encoder 410 for encoding one or more audio signals to obtain one or more encoded audio signals.
[0154] Furthermore, the device 400 comprises an additional audio information generator 420 for generating additional audio information data, the additional audio information generator 420 exhibiting a non-redundant mode of operation and a redundant mode of operation.
[0155] The additional audio information generator 420 is configured to generate the additional audio information data such that the additional audio information data comprises second additional audio information when the additional audio information generator 420 indicates a non-redundant mode of operation.
[0156] Furthermore, the additional audio information generator 420 is configured to generate the additional audio information data when the additional audio information generator 420 indicates a non-redundant operating mode, such that the additional audio information data does not include the second additional audio information or includes only a part of the second additional audio information, so that the second additional audio information can be obtained using the additional audio information data together with the first additional audio information.
[0157] According to one embodiment, the additional audio information generator 420 may be, for example, a propagation information generator for generating propagation information data as the additional audio information data. The propagation information generator may be configured, for example, to generate the propagation information data such that, when the propagation information generator indicates a non-redundant operating mode, the propagation information data includes second additional audio information, which is the second propagation information. Furthermore, the propagation information generator may be configured, for example, to generate the propagation information data such that, when the propagation information generator indicates a non-redundant operating mode, the propagation information data does not include the second propagation information or includes only a part of the second propagation information, such that the second propagation information can be obtained using the propagation information data together with the first propagation information.
[0158] According to one embodiment, the first propagation information and / or the second propagation information may depend on one or more propagations of one or more sound waves along one or more propagation paths in, for example, a real listening environment or a virtual listening environment or an augmented listening environment.
[0159] In one embodiment, the propagation information data can include, for example, reflection information data and / or diffraction information data. The first propagation information can include, for example, first reflection information and / or first diffraction information. The second propagation information can include, for example, second reflection information and / or second diffraction information.
[0160] According to one embodiment, the propagation information generator may be, for example, a reflection information generator for generating reflection information data as propagation information data. The reflection information generator may be configured, for example, when the reflection information generator indicates a non-redundant operating mode, to generate the reflection information data such that the reflection information data includes the second reflection information as the second propagation information. Furthermore, the reflection information generator may be configured, for example, when the reflection information generator indicates a non-redundant operating mode, to generate the reflection information data such that the reflection information data does not include the second reflection information or includes only a portion of the second reflection information, such that the second reflection information can be obtained using the reflection information data together with the first propagation information, which is the first reflection information.
[0161] In one embodiment, the first reflection information and / or the second reflection information may, for example, depend on one or more reflections at one or more reflective objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment.
[0162] According to one embodiment, the propagation information generator may be, for example, a diffraction information generator for generating diffraction information data as propagation information data. The diffraction information generator may be configured, for example, when the diffraction information generator indicates a non-redundant operating mode, to generate the diffraction information data such that the diffraction information data includes the second diffraction information as the second propagation information. Furthermore, the diffraction information generator may be configured, for example, when the diffraction information generator indicates a non-redundant operating mode, to generate the diffraction information data such that the diffraction information data does not include the second diffraction information or includes only a portion of the second diffraction information, such that the second diffraction information can be obtained using the diffraction information data together with the first propagation information, which is the first diffraction information.
[0163] In one embodiment, the first diffraction information and / or the second diffraction information may, for example, depend on one or more diffractions at one or more diffractive objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment.
[0164] According to one embodiment, the propagation information generator may be configured to generate the propagation information data such that, for example, in a redundant operation mode, the propagation information data may indicate one or more propagation sequences to be removed from the first propagation information, for example, the first set of propagation sequences, and / or may indicate one or more propagation sequences to be added to the first set of propagation sequences to obtain the second propagation information, for example, the second set of propagation sequences.
[0165] In one embodiment, each propagation sequence of the first set of propagation sequences and the second set of propagation sequences may, for example, represent one or more groups of reflective objects or one or more groups of diffractive objects.
[0166] In one embodiment, the propagation information generator may be configured to generate the propagation information data, such that in a non-redundant mode of operation, the propagation information data may include, for example, a second set of propagation sequences.
[0167] According to one embodiment, a first set of propagation sequences may be associated with, for example, a first listener position and a first source position. A second set of propagation sequences may be associated with, for example, a second listener position and a second source position. The first listener position may be different from, for example, the second listener position, and / or the first source position may be different from, for example, the second source position.
[0168] In one embodiment, the first set of propagation sequences may be, for example, a first set of reflection sequences. The propagation information generator may be, for example, a reflection information generator. The second set of propagation sequences may be, for example, a second set of reflection sequences. The propagation information data may be, for example, reflection information data. Each reflection sequence in the first set of reflection sequences may include, for example, information about a group of one or more reflecting objects in the reflection sequence, where sound waves emitted by an audio source at a first source location and perceptible to a listener at a first listener location are reflected on their way to a current listener location. The reflection information generator may be configured to generate the reflection information data, for example, such that each reflection sequence in the second set of reflection sequences may include, for example, information about a group of one or more reflecting objects in the reflection sequence, where sound waves emitted by an audio source at a second source location and perceptible to a listener at a second listener location are reflected on their way to a current listener location.
[0169] According to one embodiment, one or more encoded audio signals are associated with audio sources located at source positions of a second set of the reflection sequence.
[0170] In one embodiment, the reflection information generator may be configured to generate notifications suitable for determining a first listener position and a first source position for a first set of reflection sequences, for example in a redundant mode of operation.
[0171] According to one embodiment, the reflection information generator may be configured to generate a notification, e.g., in a redundant mode of operation, such that the notification may indicate, e.g., to select a first listener position and a first source position such that the first listener position is adjacent to the second listener position and / or such that the first source position is adjacent to the second listener position.
[0172] In one embodiment, the reflection information generator may be configured to generate a notification, such that in a redundant mode of operation, the notification may indicate, for example, to select a first listener position and a first source position such that the first listener position is adjacent to the second listener position and the first source position is the same as the second listener position.
[0173] Alternatively, in one embodiment, the reflection information generator may be configured to generate a notification such that, for example, in a redundant mode of operation, the notification may indicate selecting a first listener position and a first source position such that the first listener position is identical to the second listener position and such that the first source position is adjacent to the second listener position.
[0174] According to one embodiment, a first position and a second position are adjacent in a coordinate system if, in each coordinate direction of the coordinate system, the first position is immediately before or after the second position, or is identical to the second position, and if, in at least one coordinate direction of the coordinate system, the first position and the second position are different from each other.
[0175] In one embodiment, the reflection information generator may, for example, in a redundant mode of operation, determine whether the notification is e.g. the reflective information data exhibiting a non-redundant state; the reflection information data indicates a first redundancy state, whereby the first listener position and the first source position are selected such that the first source position is identical to the second source position and such that the first listener position is adjacent to the second listener position, where in a first coordinate direction of the coordinate system the first listener position is immediately before the second listener position, and where in a second coordinate direction and a third coordinate direction of the coordinate system the first listener position is identical to the second listener position; the reflection information data indicates a second redundancy state, whereby the first listener position and the first source position are selected such that the first source position is identical to the second source position and such that the first listener position is adjacent to the second listener position, where the first listener position immediately precedes the second listener position in a second coordinate direction of the coordinate system, and where the first listener position is identical to the second listener position in the first coordinate direction and a third coordinate direction of the coordinate system; The notification can be configured to generate such that the reflection information data indicates a third redundant state, whereby the first listener position and the first source position are selected such that the first source position is identical to the second source position and such that the first listener position is adjacent to the second listener position, where in a third coordinate direction of the coordinate system the first listener position is immediately before the second listener position, and where in the first coordinate direction and the second coordinate direction of the coordinate system the first listener position is identical to the second listener position.
[0176] According to one embodiment, each of the first listener position, the first source position, the second listener position, and the second source position may, for example, define the position of a voxel among a plurality of voxels in a three-dimensional coordinate system.
[0177] Figure 11 shows a system according to another embodiment, comprising the apparatus 400 of Figure 10 for encoding one or more audio signals to obtain one or more encoded audio signals and for generating additional audio information data, and further comprising the apparatus 300 of Figure 9 for generating one or more audio output signals from the one or more encoded audio signals in response to the additional audio information data.
[0178] Further specific embodiments are provided below. More specifically, binary encoding and decoding of metadata is considered.
[0179] The current working draft of the MPEG-I 6DoF Audio standard ("RM0 First Draft Version") states that earlySurfaceDataJSON, earlySurfaceConnectedDataJSON, and earlyVoxelDataJSON are represented as follows: "A zero-terminated string in ASCII encoding, which contains a document in JSON format as a provisional data format." In this input document, we propose to use an encoding method to replace this provisional data format with a binary data format, which results in a significantly smaller bitstream size.
[0180] This core experiment is based on the first draft version of RM0. It aims to replace the JSON format of early reflection metadata with a binary encoding format. By applying certain techniques, a significant reduction in the size of the early reflection payload is achieved while introducing a small amount of quantization error.
[0181] Techniques applied to reduce payload size include: 1. Data consolidation: Variables that are no longer used by the RefSoft renderer (earlySurfaceConnectedData) are removed.
[0182] 2. Coordinate system: The unit normal vector of the reflecting surface is transmitted in spherical coordinates rather than Cartesian coordinates to reduce the number of coefficients from 3 to 2.
[0183] 3. Quantization: The coefficients defining the reflecting surfaces are quantized at high resolution (near-lossless coding).
[0184] 4. Entropy Coding: A codebook-based generalized coding scheme is used for entropy coding of the transmitted symbols. The applied method is particularly useful for data sequences with a very large number of symbols, but is also suitable for a small number of symbols.
[0185] 5. Inter-voxel redundancy reduction: The similarity of voxel data in the vicinity of a voxel is exploited to further reduce the bitstream size. A differential technique is used to encode the difference between the current voxel data set and the adjacent voxel data set.
[0186] While the runtime complexity of the renderer is unaffected by the proposed changes, the decoder is simplified as the JSON data parsing step is no longer required.
[0187] Furthermore, the proposed replacement also reduces the library dependencies of the renderer and the encoder, since generating and parsing JSON documents is no longer required.
[0188] For all "Test 1" and "Test 2" scenes, the proposed coding method provides an average reduction of 21.33% in overall bitstream size over P13. Considering only scenes reflecting mesh data, the proposed coding method provides an average reduction of 28.91% in overall bitstream size over P13.
[0189] In the following, information regarding additions / replacements is considered. The encoding method presented in this core experiment is intended as a replacement for the main part of payloadEarlyReflections(). It means that the corresponding payload handler in the reference software for packets of type PLD_EARLY_REFLECTIONS will be replaced accordingly.
[0190] Further technical information is provided below. In particular, it is proposed to remove unused variables.
[0191] The RM0 bitstream parser generates the data structures earlySurfaceData and earlySurfaceConnectedData from the bitstream variables earlySurfaceDataJSON and earlySurfaceConnectedDataJSON. This data defines the static scene geometry and the reflective surfaces of triangles that belong to the connected surface area. The motivation for dividing the set of all triangles that belong to the reflective surfaces into several groups of connected areas was to allow the renderer to check only a subset during visibility tests. However, the reference software implementation no longer makes use of this specific information. Internally, the Intel Embree library is used for fast ray tracing using a proprietary acceleration method (bounding volume hierarchical data structure).
[0192] Therefore, it is proposed to simplify these data structures by combining them into a single data structure without connected surface information:
[0193] [Table 4] In the following, quantization is taken into account.
[0194] Instead of sending the Cartesian coordinates of the unit normal vector N0, it is more efficient to send the spherical coordinates as a distance where one of the values is a constant and does not need to be sent. TIFF2025528680000009.tif542TIFF2025528680000010.tif650TIFF2025528680000011.tif417
[0195] The azimuth angle is as follows: TIFF2025528680000012.tif49 in 12-bit, elevation It is proposed to quantize TIFF2025528680000013.tif47 at 11 bits. TIFF2025528680000014.tif1247TIFF2025528680000015.tif1059The elevation angle of surface normal N0 is as follows. TIFF2025528680000016.tif1243TIFF2025528680000017.tif1071
[0196] This quantization scheme ensures that integer multiples of 5° and various dividers of 360° that are powers of 2 lie directly on the quantization grid. The resulting 4032 quantization steps for azimuth and 2017 quantization steps for elevation can be considered quasi-reversible due to the high resolution.
[0197] For the quantization of the surface distance d, we propose a resolution of 1 mm, which is the same resolution also used to transmit the scene geometry data.
[0198] The actual number of bits used to transmit these values depends on the entropy coding scheme described in the following section.
[0199] In the following, entropy coding according to a particular embodiment is considered. When the symbol distribution is not uniform, entropy coding can be used to reduce the amount of bits required to transmit the data. A widely used method for entropy coding is Huffman coding, which uses smaller codewords for more frequent symbols and longer codewords for less frequent symbols, resulting in a smaller average word size. Recently, arithmetic coding has become popular, in which the complete message text is coded at once. For example, adaptive arithmetic coding schemes are used to code directional data. This adaptive method is particularly advantageous when the symbol distribution is steadily changing over time.
[0200] In the case of early reflection metadata, no assumptions can be made about the temporal behavior of the symbol distribution (such as some symbols occurring more frequently at the beginning of a transmission and other symbols occurring more frequently at the end of a transmission). It is more reasonable to assume that the symbol distribution is fixed and can be determined during encoder initialization. Furthermore, adjusting the symbol distribution at run time and using a symbol distribution that deviates from the a priori known symbol distribution would effectively negate the theoretical benefits of adaptive arithmetic coding methods.
[0201] For this reason, it has been proposed to use classical Huffman coding for entropy coding of early reflection metadata. This requires either a predefined codebook to be used, the codebook used, or a binary decoding tree to be transmitted along with the corresponding symbol list. The latter can be efficiently generated by a recursive algorithm, traversing the decoding tree and encoding leaves, i.e., valid codewords, with "1" and branches with "0". If the current word is not a valid codeword, i.e., the algorithm is at a branch in the decoding tree, two recursions are performed: one on the left where the current word is extended with "0"s, and one on the right where the current word is extended with "1". The following pseudocode shows the decoding tree encoding algorithm:
[0202] Using a predefined codebook is actually one of three options: using a predefined codebook, or using a codebook containing a codeword list and a symbol list, or using a decoding tree and a symbol list. function traverseTreeEncode(Bitstream reference bs, List <int>reference symbol_list, List <bool>code) { if (code in codebookInverse) { bs.append(1); symbol = codebookInverse[code]; symbol_list.append(symbol); } else { bs.append(0); traverseTreeEncode(bs, symbol_list, code + 0); traverseTreeEncode(bs, symbol_list, code + 1); } }
[0203] The algorithm also generates a list of all symbols in tree traversal order. The same mechanism can be used on the decoder side to extract the decoding tree topology and valid codewords. function traverseTreeDecode(Bitstream reference bs, List <int>reference code_list, List <bool>code) { bool isLeaf = bs.readBool(); if (isLeaf) { code_list.append(code); } else { traverseTreeDecode(bs, code_list, code + 0); traverseTreeDecode(bs, code_list, code + 1); } } This results in a very efficient encoding of the decoding tree, since only a single bit is spent for each codeword and each branch.
[0204] In addition to the topology of the decoding tree, the symbol list needs to be transmitted in tree traversal order for a complete transmission of the codebook.
[0205] In some cases, transmitting a codebook in addition to the symbols can result in an even larger bitstream than simple fixed-length coding. Therefore, we introduce a new generic method for transmitting data using a codebook. Our proposed method utilizes either variable-length coding or fixed-length coding using the coding schemes described above. In the latter case, instead of the complete codebook, only the word size, i.e., the number of bits in each codeword, has to be transmitted. Optionally, a common offset of integer values of symbols may be given in the bitstream if the difference with the offset results in a smaller word size. The following function parses such a generic codebook and returns a data structure for the current codebook instance:
[0206] [Table 5]
[0207] In this implementation, the keyword "Bitarray" is used as an alias for a bit sequence of a particular length, and the keyword "append()" indicates how to extend the length of the array by one or more elements added to the end. The recursive tree traversal function is defined as follows:
[0208] [Table 6]
[0209] Since they have different symbol distributions, we propose to use individual codebooks for the following arrays: earlySurfaceLengthFaceIdx earlySurfaceFaceIdx · earlySurfaceAzi · earlySurfaceEle earlySurfaceDist earlyVoxelL (see next section) earlyVoxelS (see next section) earlyVoxelIndicesRemovedDiff (see next section) earlyVoxelNumPaths (see next section) earlyVoxelOrder (see next section) earlyVoxelSurf (see next section) In the following, inter-voxel redundancy reduction according to a particular embodiment is described.
[0210] The early reflection voxel database earlyVoxelDatabase[l][s] stores a list of reflection sequences that are potentially visible to the source in the voxel with index s and the listener in the voxel with index l. Often, this list of reflection sequences is very similar for neighboring voxels. Reducing this inter-voxel redundancy can significantly reduce the bitstream size.
[0211] The proposed inter-voxel redundancy reduction uses four operating modes signaled by the bitstream variable earlyVoxelMode[v]. In mode 0 ("no reference"), the list of reflection sequences for source voxel earlyVoxelS[v] and listener voxel earlyVoxelL[v] is transmitted as an array with path index p and order index o using a generic codebook in variables earlyVoxelNumPaths[v], earlyVoxelOrder[v][p], and earlyVoxelSurf[v][p][o]. In the other operating modes, the difference between the reference and current list of reflection sequences is transmitted.
[0212] In mode 1 ("x-axis reference"), the list of reflection sequences of the current source voxel and its neighboring listener voxels in the negative x-axis direction is used as a reference. Along with the list of additional reflection sequences, a list of indices specifying the reference list entries that need to be removed is sent.
[0213] Mode 2 ("y-axis based") differs from Mode 1 by using neighboring listener voxels in the negative y-axis direction.
[0214] Mode 3 ("z-axis reference") differs from Mode 1 by using neighboring listener voxels in the negative z-axis direction.
[0215] The index list earlyVoxelIndicesRemoved[v] specifying the reference list entries that need to be removed can be coded more efficiently if a zero-terminated list of differences earlyVoxelIndicesRemovedDiff[v] is sent instead. This reduces entropy, as smaller values become more likely and larger values become less likely, resulting in a more pronounced distribution. The transformation is performed via accumulation.
[0216] [Table 7] The following describes the syntax of the generalized codebook.
[0217] Some payloads, such as payloadEarlyReflections(), utilize individual codebooks defined within the bitstream using the following syntax:
[0218] [Table 8]
[0219] The codeword list "codeList" is transmitted using the following recursive tree traversal algorithm, where the keyword "Bitarray" is used as an alias for a bit sequence of a particular length. Additionally, the keyword "append()" indicates how to extend the length of the array by one or more elements added to the end.
[0220] [Table 9]
[0221] An instance of such a codebook, "exampleCodebook", is created as follows: exampleCodebook = genericCodebook(); In addition to the data fields of the returned data structure, the generic codebook has a method "get_symbol()" that reads a valid codeword from the bitstream, i.e., the nth element of codeList[], and returns the corresponding symbol, i.e., symbolList[n]. Use of this method is illustrated as follows: exampleVariable = exampleCodebook.get_symbol(); Below, a proposed syntax for the Early Reflection payload is presented.
[0222] [Table 10]
[0223] [Table 11]
[0224] [Table 12] In the following, a proposed data structure is presented: the Early Reflection Payload Data Structure.
[0225] earlyTriangleCullingDistanceOrder1 Triangle culling distance for 1st order reflections. earlyTriangleCullingDistanceOrder2 Triangle culling distance for secondary reflections. earlySourceCullingDistanceOrder1 Source culling distance for 1st order reflections. earlySourceCullingDistanceOrder2 Source culling distance for secondary reflections. earlyVoxelGridOriginX The x component of the Cartesian coordinate of the voxel grid origin [0,0,0]. earlyVoxelGridOriginY The y component of the Cartesian coordinate of the voxel grid origin [0,0,0]. earlyVoxelGridOriginZ The z component of the Cartesian coordinate of the voxel grid origin [0,0,0]. earlyVoxelGridPitchX Voxel grid spacing along the x-axis (voxel width). earlyVoxelGridPitchY Voxel grid spacing along the y-axis (voxel length). earlyVoxelGridPitchZ Voxel grid spacing along the z-axis (voxel height). earlyVoxelGridShapeX The number of voxels along the x-axis. earlyVoxelGridShapeY The number of voxels along the y axis. earlyVoxelGridShapeZ The number of voxels along the z-axis. earlyHasSurfaceData Flag indicating the presence of earlySurfaceData. earlySurfaceDataLength The length of the earlySurfaceData block in bytes. earlyHasVoxelData A flag indicating the presence of earlyVoxelData. earlyVoxelDataLength The length of the earlySurfaceData block in bytes. earlySurfaceDistOffset The offset of earlySurfaceDist in mm. numberOfSurfaces The number of surfaces. earlySurfaceLengthFaceIdx The length of the array of earlySurfaceFaceIdx. earlySurfaceFaceIdx A list of triangle IDs. earlySurfaceAzi An array with azimuth angles specifying the surface normal in spherical coordinates (Hessian standard system). earlySurfaceEle An array with elevation angles specifying the surface normal in spherical coordinates (Hessian standard system). earlySurfaceDist Array with distance values (Hessian standard system) numberOfVoxelPairs The number of source and listener voxel pairs with available voxel data. earlyVoxelL An array with the voxel indices of the listener. earlyVoxelS An array with the source voxel indices. earlyVoxelMode An array that specifies the encoding mode for the voxel data. earlyVoxelIndicesRemovedDiff A differential encoding removal list that specifies the indices of the reference reflection sequence list that should be removed. earlyVoxelNumPaths The number of reflection paths. earlyVoxelOrder A 2D array specifying the reflection order. earlyVoxelSurf A reflection sequence given as a 3D array of surface indices.
[0226] In the following, a renderer stage that takes early reflections into account is proposed and terminology and definitions are provided. Voxel Grid: The renderer uses the voxel data to speed up computationally complex visibility checks of reflected sound propagation paths. The scene is rasterized to a regular grid with grid spacing that can be defined separately for each dimension. Each voxel is identified by a unique voxel ID, and a sparse database is used to store pre-computed data for a given source / listener voxel pair. The relevant variables and data structures are: earlyVoxelGridOriginX earlyVoxelGridOriginY earlyVoxelGridOriginZ earlyVoxelGridPitchX · earlyVoxelGridPitchY earlyVoxelGridPitchZ earlyVoxelGridShapeX earlyVoxelGridShapeY earlyVoxelGridShapeZ These variables are the voxel coordinates V = [v x ,v y ,v z ] T This is the basis of the equation. For any point P=[p x ,p y ,p z ] T For , the corresponding voxel coordinates are calculated by the following rounding operation to the nearest integer:
[0227] [Table 13] Voxel coordinates can be converted to voxel indices
[0228] [Table 14] This representation is used, for example, in the sparse voxel database earlyVoxelDatabase[l][s][p] of listener voxel ID l and source voxel IDs.
[0229] Culling Distance: The encoder can use source and / or triangle distance culling to speed up the pre-computation of voxel data. The culling distance is encoded in the bitstream to allow the renderer to smoothly fade out reflections that reach the culling threshold used. The relevant variables and data structures are: ·earlyTriangleCullingDistanceOrder1 ·earlyTriangleCullingDistanceOrder2 ·earlySourceCullingDistanceOrder1 ·earlySourceCullingDistanceOrder2
[0230] Surface Data: Surface data is geometric data that defines the reflective surfaces from which sound is reflected. The relevant variables and data structures are as follows: ·earlySurfaceIdx[s]; ·earlySurfaceFaceIdx[s][f]; earlySurface_N0[s] earlySurface_d[s]
[0231] The surface index earlySurfaceIdx[s] identifies the surface and is referenced by the sparse voxel database earlyVoxelDatabase[l][s][p]. The triangle ID list earlySurfaceFaceIdx[s][f] defines the triangles of the StaticMesh that belong to this surface. One of these triangles must be hit for the specular reflection visibility test to succeed. The reflective face of each surface is given in the Hessian standard system using the surface normal N0 and surface distance d transformed as follows: int max_steps_azi = 1 << 12; int max_steps_ele = 1 << 11; int num_steps_azi = 144 * (max_steps_azi / 144); int num_steps_ele = 72 * (max_steps_ele / 72); int shift_ele = num_steps_ele / 2; float quant2azi = double(2.0 * M_PI) / double(num_steps_azi); float quant2ele = double(M_PI) / double(num_steps_ele); float quant2dist = 0.001f; for (int s = 0; s < numberOfSurfaces; s++) { earlySurfaceIdx[s] = s; float azi = earlySurfaceAzi[s] * quant2azi; float ele = (earlySurfaceEle[s] - shift_ele) * quant2ele; earlySurface_N0[s][0] = -1.0 * sin(azi) * cos(ele); earlySurface_N0[s][1] = sin(ele); earlySurface_N0[s][2] = -1.0 * cos(azi) * cos(ele); earlySurface_d[s] = (earlySurfaceDist[s] + dist_offset) * quant2dist;
[0232] Voxel Data The early reflection voxel data is a sparse voxel database containing a list of reflection sequences of potentially visible image sources for a given source voxel and listener voxel pair. An entry in the database can be undefined if the given source voxel and listener voxel pair is not specified in the bitstream, can be an empty list, or can contain a list of surface connection IDs. The relevant variables and data structures are: numberOfVoxelPairs earlyVoxelL[v] earlyVoxelS[v] earlyVoxelMode[v] ·earlyVoxelIndicesRemovedDiff[v][k] earlyVoxelNumPaths[v] earlyVoxelOrder[v][p] earlyVoxelSurf[v][p][o]
[0233] The sparse voxel database earlyVoxelDatabase[l][s][p] is derived from these variables by the following algorithm: int delta_x = voxelCoordinateToVoxelIndex( {1, 0, 0} ); int delta_y = voxelCoordinateToVoxelIndex( {0, 1, 0} ); int delta_z = voxelCoordinateToVoxelIndex( {0, 0, 1} ); int delta_list[4] = { 0, -delta_x, -delta_y, -delta_z}; for (int v = 0; v < numberOfVoxelPairs; v++) { PathList path_list; int l = earlyVoxelL[v]; int s = earlyVoxelS[v]; int mode = earlyVoxelMode[v]; if (mode != 0) { int l_ref = l + delta_list[mode]; path_list = earlyVoxelDatabase[l_ref][s]; / / generate list with removed items in reverse order int numberOfIndicesRemoved = length(earlyVoxelIndicesRemovedDiff[v]) - 1; int listIndicesRemoved[numberOfIndicesRemoved]; int val = -1; for (int k = 0; k < numberOfIndicesRemoved; k++) { val += earlyVoxelIndicesRemovedDiff[v][k]; listIndicesRemoved[numberOfIndicesRemoved - 1 - k] = val; } / / remove reflection sequences for (int k = 0; k < numberOfIndicesRemoved; k++) { path_list.erase(listIndicesRemoved[k]); } } / / add reflection sequences for (int p = 0; p < earlyVoxelNumPaths[v]; p++) { path_list.append(earlyVoxelSurf[v][p]); } / / add sorted path list to sparse voxel database path_list = shortlex_sort(path_list); int num_paths = length(path_list); for (int p = 0; p < num_paths; p++) { earlyVoxelDatabase[l][s][p] = path_list[p]; } }
[0234] In this algorithm, the function voxelCoordinateToVoxelIndex() represents the conversion from voxel coordinates to voxel indexes, the keyword PathList represents a list of integer arrays that can be modified by the methods append() to add an element to the end of the list and erase() to remove a list element at a given position, and the function shortlex_sort() represents a sorting function that sorts a given list of reflection sequences into shortlex order.
[0235] Complexity Assessment While the runtime complexity of the renderer is unaffected by the proposed changes, the decoder is simplified as the JSON data parsing step is no longer required. Basis of benefits To verify the correct functioning of the proposed method and prove its technical advantages, we encoded all "Test 1" and "Test 2" scenes and compared the size of the early reflection metadata with the encoding results of the P13 encoder.
[0236] Data Compression Table 2 lists the size of the payload Early Reflections for the P13 encoder ("Old Size / Bytes") and the variant of the P13 encoder using the proposed encoding method ("New Size / Bytes"). The last column lists the achieved compression ratio, i.e., the ratio between the old and new payload sizes.
[0237] In all cases, the proposed method results in smaller payload sizes. For all scenes that reflect scene objects, i.e., scenes with mesh data, compression ratios of over 10 were achieved. For some scenes ("SingerInTheLab" and "VirtualBasketball"), compression ratios approaching or even exceeding 100 were achieved.
[0238] [Table 15] In the following, the total bitstream savings are considered.
[0239] The table below lists the total bitstream size savings in percentage. On average, the total bitstream size was reduced by 21.33%. When considering only scenes with mesh data, the total bitstream size was reduced by 28.91% on average.
[0240] [Table 16]
[0241] Data Verification and Quantization Error The table below lists the results of data validation tests on an extended test set that includes all "Test 4" scenes and additional scenes that did not make it into the official test repository, comparing decoded metadata, such as earlySurfaceData and earlyVoxelData, with the output of the P13 decoder. For P13 payloads, we combined connected surface data with surface data to allow for comparison with the new encoding method. The validation result "same structure" means that both payloads have the same reflective surface and the data differs by the expected quantization error.
[0242] For all scenes, the decoded earlyVoxelData was identical, and the decoded earlySurfaceData was either identical or structurally identical.
[0243] [Table 17]
[0244] The table below lists the minimum, mean, median, and maximum quantization error in mm of the transmission plane normal N0 after transformation to Cartesian coordinates. A maximum quantization error of 1.095 mm corresponds to an angular deviation of 0.063°. Using a resolution of 0.088° per quantization step, and therefore a maximum quantization error of 0.044° per axis, the observed results are in good agreement with the theoretical values.
[0245] The maximum angular deviation of 0.063° for the surface normal vector N0 is so small that the transmission can be considered quasi-lossless.
[0246] [Table 18]
[0247] The table below lists the minimum, mean, median, and maximum quantization errors in mm for the transmission surface distance. With a resolution of 1 mm per quantization step, the observed maximum deviation of 0.519 mm agrees well with the expected maximum of 0.5 mm. The overshoot can be explained by the limited accuracy of the single-precision floating-point variables used, which do not provide sufficient sub-millimeter resolution for large-scale scenes such as "Park," "Parking Lot," and "Recreation."
[0248] The maximum deviation of the surface distance d, 0.519 mm, is so small that the transmittance can be considered quasi-lossless.
[0249] [Table 19]
[0250] In one embodiment, a binary encoding method is provided for earlySurfaceData() and earlyVoxelData() as part of the early reflections metadata within payloadEarlyReflections(). For a test set containing 30 AR and VR scenes, the decoded data was compared to data decoded by a P13 decoder, and only the expected quantization error was observed. The quantization error for the surface data was so small that the transmittance can be considered quasi-lossless. The transmitted voxel data was identical.
[0251] In all cases, the proposed method results in smaller payload sizes. For all scenes reflecting scene objects, i.e., scenes with mesh data, compression ratios of over 10 were achieved. For some scenes ("Singer In The Lab" and "Virtual Basketball"), compression ratios approaching or even exceeding 100 were achieved. For all "Test 1" and "Test 2" scenes, the proposed encoding method provides an average 21.33% reduction in overall bitstream size over P13. Considering only scenes reflecting mesh data, the proposed encoding method provides an average 28.91% reduction in overall bitstream size over P13.
[0252] The proposed encoding method does not affect the runtime complexity of the renderer. Furthermore, the proposed replacement also reduces the library dependencies of the reference software, since generating and parsing JSON documents is no longer required.
[0253] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0254] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partly in hardware, or at least partly in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.
[0255] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0256] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code can be stored on, for example, a machine-readable carrier.
[0257] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0258] In other words, therefore, an embodiment of the inventive methods is a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0259] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0260] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can for example be arranged to be transmitted via a data communication connection, for example via the Internet.
[0261] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0262] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0263] Further embodiments according to the invention comprise an apparatus or system configured to transfer (e.g. electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may for example be a computer, a mobile device, a memory device, etc. The apparatus or system may for example comprise a file server for transferring the computer program to the receiver.
[0264] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0265] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0266] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0267] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented as descriptions and explanations of the embodiments herein.< / bool> < / int> < / bool> < / int>
Claims
1. An apparatus (100) for generating one or more audio output signals from one or more encoded audio signals, said apparatus (100) comprising: at least one entropy decoding module (110; 116, 118) for decoding the encoded additional audio information to obtain decoded additional audio information, if the encoded additional audio information is entropy coded; a signal processor (120) for generating the one or more audio output signals in response to the one or more encoded audio signals and in response to the decoded additional audio information.
2. The device (100) at least one non-entropy decoding module (111) for decoding the encoded additional audio information to obtain the decoded additional audio information if the encoded additional audio information is not entropy coded; a selector (115) for selecting one of the at least one entropy decoding module (110; 116, 118) and the at least one non-entropy decoding module (111) for decoding the encoded additional audio information depending on whether the encoded additional audio information is entropy coded or not; The apparatus (100) of claim 1, further comprising:
3. the encoded additional audio information comprises augmented reality or virtual reality data; 3. The apparatus (100) of claim 1 or 2.
4. the encoded additional audio information is dependent on a real listening environment, a virtual listening environment, or an augmented listening environment; The apparatus (100) according to any one of claims 1 to 3.
5. the encoded additional audio information comprises propagation information dependent on one or more propagations of one or more sound waves along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment. The apparatus (100) of claim 4.
6. the propagation information is reflection information dependent on one or more reflections at one or more reflecting objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment; 6. The apparatus (100) of claim 5.
7. the propagation information is diffraction information dependent on one or more diffractions at one or more diffractive objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment; 6. The apparatus (100) of claim 5.
8. the encoded additional audio information includes data for rendering early reflections; the signal processor (120) is configured to generate the one or more audio output signals in response to the data for rendering early reflections. An apparatus (100) according to any one of claims 1 to 7.
9. the signal processor (120) is configured to generate a binaural signal comprising two binaural channels as the one or more audio output signals; An apparatus (100) according to any one of claims 1 to 8.
10. the at least one entropy decoding module (110; 116, 118) comprising a Huffman decoding module (116) for decoding the encoded additional audio information if the encoded additional audio information is Huffman coded; The apparatus (100) according to any one of claims 1 to 9.
11. the at least one entropy decoding module (110; 116, 118) comprises an arithmetic decoding module (118) for decoding the coded additional audio information if the coded additional audio information is arithmetically coded, An apparatus (100) according to any one of claims 1 to 10.
12. the selector (115) is configured to select one of the at least one non-entropy decoding module (111), the Huffman decoding module (116), and the arithmetic decoding module (118) for decoding the encoded additional audio information. The device (100) according to claim 2, claim 10 and claim 11.
13. the at least one non-entropy decoding module (111) comprises a fixed length decoding module for decoding the encoded additional audio information if the encoded additional audio information is fixed length coded; An apparatus (100) according to any one of claims 3 to 12, further dependent on claim 2.
14. the device (100) is configured to receive selection information; the selector (115) is configured to select one of the at least one entropy decoding module (110; 116, 118) and the at least one non-entropy decoding module (111) according to the selection information. An apparatus (100) according to any one of claims 3 to 13, further dependent on claim 2.
15. the apparatus (100) is configured to receive a codebook or coding tree on which the encoded additional audio information depends, the at least entropy decoding module (110; 116, 118) is configured to decode the encoded additional audio information using the codebook or using the coding tree.
15. The apparatus (100) according to any one of claims 1 to 14.
16. the device (100) is configured to receive an encoding of a structure of the coding tree on which the encoded additional audio information depends, the at least entropy decoding module (110; 116, 118) is configured to reconstruct a plurality of codewords of the coding tree according to the structure of the coding tree, the at least entropy decoding module (110; 116, 118) is configured to decode the encoded additional audio information using the codewords of the coding tree.
16. The apparatus (100) of claim 15.
17. the apparatus (100) further comprises a memory storing a codebook or coding tree; the at least entropy decoding module (110; 116, 118) is configured to decode the encoded additional audio information using the codebook or using the coding tree.
15. The apparatus (100) according to any one of claims 1 to 14.
18. the apparatus (100) is configured to receive the encoded additional audio information, the additional audio information including a plurality of transmitted symbols and an offset value; the at least one non-entropy decoding module (111) is configured to decode the encoded additional audio information using the plurality of transmitted symbols and using the offset value.
18. The apparatus (100) of any one of claims 1 to 17.
19. the data for rendering early reflections comprises information regarding the location of one or more walls in an environment, the walls being one or more real or virtual walls; the signal processor (120) is configured to generate the one or more audio output signals in response to the information regarding the position of one or more walls. Apparatus (100) according to any one of claims 9 to 18, further dependent on claim 8.
20. the information for each wall of the one or more walls comprises information about an azimuth angle and / or an elevation angle of the wall, the azimuth angle of the wall being entropy coded and / or the elevation angle of the wall being entropy coded; one or more of the entropy decoding modules of the at least one entropy decoding module (110; 116, 118) is configured to decode the entropy coded azimuth angle of the wall and / or the entropy coded elevation angle of the wall.
20. The apparatus (100) of claim 19.
21. wherein the one or more of the at least one entropy decoding modules (110; 116, 118) are configured to decode the entropy coded azimuth angle of the wall and / or the entropy coded elevation angle of the wall using the codebook or the coding tree.
21. Apparatus (100) according to claim 20, further dependent on any one of claims 15 to 17.
22. the encoded additional audio information comprises voxel position information, the position information comprising information regarding one or more positions of one or more voxels of a plurality of voxels within a three-dimensional coordinate system; the signal processor (120) is configured to generate the one or more audio output signals in response to the voxel position information.
22. The apparatus (100) of any one of claims 1 to 21.
23. The at least one entropy decoding module (110; 116, 118) is configured to decode the additional audio information that has been entropy coded, the additional audio information being: a list of triangle indices, the array length of the list of triangle indices, an array having azimuth angles specifying surface normals in spherical coordinates; an array having elevation angles specifying the surface normal in spherical coordinates; an array having distance values; an array having listener positions; an array having one or more sound source positions; a removal list or removal set specifying a set of reflection sequences to be removed or an index of a reflection sequence in the reference reflection sequence list to be removed; the number of reflection sequences or paths; an array specifying the reflection order; A reflex sequence; 23. The apparatus (100) of any one of claims 1 to 22, comprising at least one of:
24. An apparatus (200) for encoding one or more audio signals and additional audio information, said apparatus (200) comprising: an audio signal encoder (210) for encoding the one or more audio signals to obtain one or more encoded audio signals; and at least one entropy coding module (220) for encoding the additional audio information using entropy coding to obtain encoded additional audio information.
25. The device (200) at least one non-entropy coding module (221) for coding said additional audio information to obtain said encoded additional audio information; a selector (215) for selecting one of the at least one entropy coding module (220) and the at least one non-entropy coding module (221) for coding the additional audio information according to a symbol distribution within the additional audio information to be coded; 25. The apparatus (200) of claim 24, further comprising:
26. the encoded additional audio information comprises augmented reality or virtual reality data.
26. Apparatus (200) according to claim 24 or 25.
27. the encoded additional audio information is dependent on a real listening environment, a virtual listening environment, or an augmented listening environment; 27. Apparatus (200) according to any one of claims 24 to 26.
28. the encoded additional audio information comprises propagation information dependent on one or more propagations of one or more sound waves along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment.
28. The apparatus (200) of claim 27.
29. the propagation information is reflection information dependent on one or more reflections at one or more reflecting objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment; 29. The apparatus (200) of claim 28.
30. the propagation information is diffraction information dependent on one or more diffractions at one or more diffractive objects of one or more sound waves propagating along one or more propagation paths in the real listening environment, the virtual listening environment, or the augmented listening environment; 29. The apparatus (200) of claim 28.
31. the encoded additional audio information includes data for rendering early reflections; 31. Apparatus (200) according to any one of claims 24 to 30.
32. the at least one entropy coding module (220) comprising a Huffman coding module (226) for coding the additional audio information using Huffman coding; 32. Apparatus (200) according to any one of claims 24 to 31.
33. the at least one entropy coding module (220) comprising an arithmetic coding module (228) for coding the additional audio information using arithmetic coding; 33. Apparatus (200) according to any one of claims 24 to 32.
34. the selector (215) is configured to select one of the at least one non-entropy coding module (221), the Huffman coding module (226), and the arithmetic coding module (228) for encoding the additional audio information.
34. Apparatus (200) according to claim 25, claim 32 and claim 33.
35. 26. The method according to claim 25, wherein the at least one non-entropy coding module (221) comprises a fixed length coding module for coding the additional audio information.
35. Apparatus (200) according to any one of claims 24 to 34.
36. the device (200) is configured to generate selection information indicating one of the at least one entropy coding module (220) and the at least one non-entropy coding module (221) employed to code the additional audio information.
36. Apparatus (200) according to any one of claims 24 to 35, further dependent on claim 25.
37. the device (200) is configured to transmit a codebook or coding tree employed to encode the additional audio information.
37. Apparatus (200) according to any one of claims 24 to 36.
38. the device (200) is configured to transmit an encoding of a structure of the coding tree on which the encoded additional audio information depends.
38. The apparatus (200) of claim 37.
39. the apparatus (200) further comprises a memory storing a codebook or coding tree; the at least entropy decoding module (220) is configured to encode the additional audio information using the codebook or using the coding tree.
37. Apparatus (200) according to any one of claims 24 to 36.
40. the at least one entropy coding module (220) is configured to code the additional audio information such that the coded additional audio information comprises a plurality of transmitted symbols and an offset value.
40. Apparatus (200) according to any one of claims 24 to 39.
41. the data for rendering early reflections includes information about the location of one or more walls in the environment, the walls being one or more real or virtual walls; 41. Apparatus (200) according to any one of claims 24 to 40, further dependent on claim 31.
42. the information for each wall of the one or more walls comprises information about an azimuth angle and / or an elevation angle of the wall, the azimuth angle of the wall being entropy coded and / or the elevation angle of the wall being entropy coded; one or more entropy coding modules of the at least one entropy coding module (220) configured to code the additional audio information such that the coded additional audio information comprises an entropy coded azimuth angle of the wall and / or an entropy coded elevation angle of the wall.
42. The apparatus (200) of claim 41.
43. the one or more entropy coding modules are configured to code the entropy coded azimuth angle of the wall and / or the entropy coded elevation angle of the wall using the codebook or the coding tree.
43. Apparatus (200) according to claim 42, further dependent on any one of claims 37 to 39.
44. the encoded additional audio information includes voxel position information, the position information including information regarding one or more positions of one or more voxels of a plurality of voxels within a three-dimensional coordinate system; 44. Apparatus (200) according to any one of claims 24 to 43.
45. The at least one entropy decoding module (220) is configured to encode the additional audio information using the entropy coding, and the encoded additional audio information comprises: a list of triangle indices, the array length of the list of triangle indices, an array having azimuth angles specifying surface normals in spherical coordinates; an array having elevation angles specifying the surface normal in spherical coordinates; an array having distance values; an array having listener positions; an array having one or more sound source positions; a removal list or removal set specifying a set of reflection sequences to be removed or an index of a reflection sequence in the reference reflection sequence list to be removed; the number of reflection sequences or paths; an array specifying the reflection order; A reflex sequence; 45. The apparatus (200) of any one of claims 24 to 44, comprising at least one of:
46. - an apparatus (200) according to any one of claims 24 to 45 for encoding one or more audio signals and additional audio information to obtain one or more encoded audio signals and encoded additional audio information; An apparatus (100) according to any one of claims 1 to 23 for generating from the one or more encoded audio signals one or more audio output signals in dependence on the encoded additional audio information; A system with.
47. 1. A method for generating one or more audio output signals from one or more encoded audio signals, said method comprising: - decoding the encoded additional audio information to obtain decoded additional audio information if the encoded additional audio information is entropy coded; generating the one or more audio output signals in response to the one or more encoded audio signals and in response to the decoded additional audio information.
48. 1. A method for encoding one or more audio signals and additional audio information, said method comprising: encoding the one or more audio signals into one or more encoded audio signals; encoding the additional audio information using entropy coding to obtain encoded additional audio information.
49. 49. A computer program for performing the method of claim 47 or 48 when the computer program is run on a computer or signal processor.