Method and apparatus representing a space of interest of an audio scene, computer readable storage medium
By defining a space of interest in the audio scene and using grammatical elements for decoding and processing circuits, the problem of the inability to represent audio scenes in existing technologies is solved, and accurate processing and presentation of audio scenes are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2021-09-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies cannot effectively apply the concept of region of interest to audio scenes, especially in the process of audio encoding, decoding and presentation, where they cannot accurately represent and process the space of interest in an audio scene.
By defining the space of interest for the audio scene, processing circuits decode the audio scene data, and using syntax elements to represent subsets of multiple items in the audio scene, parts of the audio content are identified and presented, including the listener space, audio channel configuration, and audio object configuration.
It achieves accurate representation and processing of audio scenes, improves the effects of audio encoding, decoding and presentation, and meets the needs of different audio scenarios.
Smart Images

Figure CN115589787B_ABST
Abstract
Description
[0001] By citation and inclusion in this article
[0002] This application claims priority to U.S. Patent Application No. 17 / 489,212, filed September 29, 2021, entitled "Method and Apparatus for Representing a Space of Interest in an Audio Scene," and to U.S. Provisional Application No. 63 / 184,571, filed May 5, 2021, entitled "Representing a Space of Interest in an Audio Scene." The entire disclosure of the earlier application is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to embodiments of audio scene representation. Background Technology
[0004] The background description provided herein is intended to present the general context of this disclosure. The extent of the work of the currently attributed inventors described in the background section and in various aspects of this specification does not indicate that it was prior art at the time of this disclosure filing, nor is it expressly or implied that it was acknowledged as prior art to this disclosure.
[0005] A region of interest (ROI) is a sample region within a dataset identified for a specific purpose. The concept of ROI is commonly used in many application areas, such as medical imaging, geographic information systems, computer vision, and optical feature recognition.
[0006] While the concept of Region of Interest (ROI) can be used for one-dimensional audio signals, it may not be directly applicable in audio scenarios. This disclosure provides a method for representing the space of interest (ROI) of an audio scene. Summary of the Invention
[0007] This disclosure provides means for representing a space of interest (SOI) of an audio scene. One means includes: processing circuitry that decodes audio scene data of the audio scene, the audio scene data including (i) audio content representing a plurality of items of the audio scene, and (ii) a first syntax element indicating the type of a subset of the plurality of items, the subset of items representing the SOI of the audio scene. The processing circuitry determines a portion of the audio content for the subset of the plurality of items based on the type of the subset of items indicated in the first syntax element. The processing circuitry presents the determined portion of the audio content.
[0008] In one embodiment, the first syntax element indicates that the type of a subset of the plurality of items is one of a type associated with listener space, a type associated with audio channel configuration, or a type associated with audio object configuration.
[0009] In one embodiment, the audio scene data includes a second syntax element indicating the number of subsets of the plurality of items.
[0010] In one embodiment, the second syntax element indicates that the number of subsets of the plurality of items is greater than 1, and the audio scene data includes a third syntax element that indicates the identifier index of each item in the subset of the plurality of items.
[0011] In one embodiment, the first syntax element indicates that the type of a subset of the plurality of items is a type associated with the listener space, and the audio scene data includes a fourth syntax element indicating whether to signal the subtype of the listener space.
[0012] In one embodiment, the fourth syntax element indicates signaling a subtype of the listener space, and the audio scene data includes a fifth syntax element indicating the subtype of the listener space.
[0013] In one embodiment, the fourth syntax element indicates that no signal is sent to the subtype of the listener space, and the subtype of the listener space is determined based on the video scene.
[0014] In one embodiment, the subtype of the listener space is either a type associated with the best point of the audio scene or a type associated with the auditory space.
[0015] Various aspects of this disclosure provide means for representing a space of interest in an audio scene. One such means includes: a decoding module for decoding audio scene data of the audio scene, the audio scene data including (i) audio content representing a plurality of items of the audio scene, and (ii) a first syntax element indicating the type of a subset of the plurality of items, the subset of items representing the space of interest in the audio scene; a determining module for determining a portion of the audio content for the subset of the plurality of items based on the type of the subset of items indicated in the first syntax element; and for presenting the determined portion of the audio content.
[0016] This disclosure provides methods for representing the space of interest of an audio scene. In one method, audio scene data of the audio scene is decoded, the audio scene data including (i) audio content representing a plurality of items of the audio scene, and (ii) a first syntax element indicating the type of a subset of the plurality of items, the subset of items representing the space of interest of the audio scene; a portion of the audio content for the subset of the plurality of items is determined based on the type of the subset of items indicated in the first syntax element; and the determined portion of the audio content is presented.
[0017] Various aspects of this disclosure provide a non-volatile computer-readable storage medium for storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one or a combination of methods representing a space of interest for an audio scene. Attached Figure Description
[0018] Other features, properties, and various advantages of the disclosed subject matter will become further apparent from the following detailed description and accompanying drawings, wherein:
[0019] Figure 1 Exemplary optimal points for audio scenarios according to embodiments of this disclosure are shown;
[0020] Figure 2 An example of an auditory space with a limited elevation range according to an embodiment of the present disclosure is shown;
[0021] Figure 3 An example of an auditory space having a spherical shape according to an embodiment of the present disclosure is shown;
[0022] Figure 4 An example of an auditory space having a rolling ball shape according to an embodiment of the present disclosure is shown;
[0023] Figure 5 An exemplary flowchart according to an embodiment of this disclosure is shown; and
[0024] Figure 6 This is a schematic diagram of a computer system according to an embodiment of the present disclosure. Detailed Implementation
[0025] I. Representation of the Space of Interest in an Audio Scene
[0026] It should be noted that the methods included in this disclosure can be used alone or in combination. These methods can be used in part or in whole.
[0027] According to various aspects of this disclosure, the space of interest can be defined as the boundary of the space considered in an audio scene. The space of interest can be used for audio encoding and decoding, audio processing, audio presentation, etc.
[0028] An audio scene is a semantically consistent segment of sound, characterized by one or more dominant sound sources. An audio scene can be modeled as a set of sound sources. In some embodiments, an audio scene may be dominated by a subset of the set of sound sources.
[0029] In some embodiments, the space of interest can be represented by the space that the listener can move to. For example, the entire space can be divided into one or more areas that the listener can move to and other areas that the listener cannot move to. The space of interest can therefore be represented by the set of areas that the listener can move to.
[0030] In one embodiment, the space of interest can be represented by one or more sweet spots in the audio scene, where one person (e.g., a listener) is fully able to hear the audio mix generated by the audio mixer in the manner desired to be heard.
[0031] Figure 1 An exemplary optimal point for an audio scene according to an embodiment of this disclosure is shown. Figure 1 In this context, the optimal point in the audio scene is the intersection of the areas covered by the audio sources marked 1 through 7. Therefore, in Figure 1 In this context, the optimal listening point is indicated by a circle around the chair. In some cases, such as international recommendations, the optimal listening point is referred to as the reference listening point.
[0032] In some embodiments, the space of interest can be represented by the auditory space.
[0033] In one embodiment, the space of interest can be represented by an auditory space having a finite range of elevation. For example, the space of interest can be represented by two numbers, where the auditory space lies within an elevation range between these two numbers.
[0034] Figure 2 An example of an auditory space with an elevation between 0.0 meters and 4.0 meters is shown.
[0035] In one embodiment, the space of interest can be represented by an auditory space having a rectangular prism. This representation could be the coordinates of two opposite vertices of the rectangular prism. Alternatively, it could be the coordinates of one vertex of the rectangular prism, along with the values of its height, width, and length. In some cases, the rectangular prism may not always be vertical or horizontal, thus directional information about the prism can be described.
[0036] In one embodiment, the space of interest can be represented by an auditory space having a polyhedral shape. This representation can be the coordinates of the vertices of the polyhedral shape. The representation can also be a set of surfaces of the polyhedral shape.
[0037] In one embodiment, the space of interest can be represented by a spherical auditory space centered on the listener's position, such as... Figure 3 As shown. This representation can be the coordinates of the center of the sphere and the value of the sphere's radius.
[0038] In one embodiment, the space of interest can be represented by an auditory space having a rolling sphere shape. The center of the rolling sphere is along the listener's walking path, such as... Figure 4 As shown. This representation can be a function describing the walking path and the radius of the rolling sphere.
[0039] In one embodiment, the space of interest can be represented by a combination of audio channels in a multi-channel audio stream. For example, the representation could be a set of left and right front channels in a 7.1 audio stream.
[0040] In one embodiment, the space of interest can be represented by a combination of audio objects. For example, a hospital audio scene may include audio objects such as doors, tables, chairs, a television (TV), a radio, doctors, and patients. In this example, the space of interest can be represented by a set of doors, doctors, and patients.
[0041] According to various aspects of this disclosure, the space of interest can be represented by a set of two or three types of items: the space that the listener can move to (referred to as the listener space), audio channels, and audio objects. In other words, the space of interest of an audio scene can be represented by a set of listener space, audio channels, and / or audio objects.
[0042] In some embodiments, a first syntax element in the audio scene data (e.g., the space_of_interest_type flag) may be signaled to indicate whether the space of interest is the listener space, the audio channel configuration, or the audio object configuration.
[0043] In some embodiments, a second syntax element in the audio scene data of the audio scene can be signaled to indicate the number of items of each type. For example, the second syntax element can be one of the following three values: listener_space_count, audio_channel_config_count, and audio_object_config_count, which represent the number of listener spaces, the number of audio channel configurations, and the number of audio object configurations, respectively.
[0044] In one embodiment, the value of listener_space_count can be set to 0 when the listener space does not exist in the space of interest of the audio scene.
[0045] In one embodiment, the value of audio_channel_config_count can be set to 0 when the audio channel configuration does not exist in the space of interest of the audio scene.
[0046] In one embodiment, the value of audio_object_config_count can be set to 0 when the audio object configuration does not exist in the space of interest of the audio scene.
[0047] In some embodiments, when the second syntax element indicates that the total number of items of the same type is greater than 1, a third syntax element in the audio scene data of the audio scene can be signaled to indicate the identifier index of each item of the same type.
[0048] In one embodiment, when listener_space_count is greater than 1, the third syntax element can be listener_space_id, which can be signaled to indicate the identifier index of each listener space.
[0049] In one embodiment, when listener_space_count equals 1, there is exactly one listener space in the space of interest of the audio scene.
[0050] In one embodiment, when audio_channel_config_count is greater than 1, the third syntax element can be audio_channel_config_id, which can be signaled to indicate the identifier index of each audio channel configuration.
[0051] In one embodiment, when audio_channel_config_count equals 1, there is exactly one audio channel configuration in the space of interest of the audio scene.
[0052] In one embodiment, when audio_object_config_count is greater than 1, the third syntax element can be audio_object_config_id, which can be signaled to indicate the identifier index of each audio object configuration.
[0053] In one embodiment, when audio_object_config_count equals 1, there is exactly one audio object configuration in the space of interest of the audio scene.
[0054] According to various aspects of this disclosure, audio and video signals can be correlated. Therefore, the listener space of the audio scene can be set according to the corresponding video scene.
[0055] In one embodiment, the listener space of the audio scene can be set to be the same as the ROI of the video scene.
[0056] In one embodiment, the listener space of the audio scene can be part of the ROI of the video scene.
[0057] In one embodiment, the listener space of the audio scene can be outside the ROI of the video scene.
[0058] In one embodiment, a fourth syntax element (e.g., `listener_space_flag`) in the audio scene data of the audio scene can be signaled to indicate the relationship between the listener space of the audio scene and other components (e.g., the video scene). If the fourth syntax element `listener_space_flag` is set to true, it means that the listener space is the audio listener space and can be fully represented in subsequent syntax elements (e.g., the fifth syntax element `listener_space_subtype`). If the fourth syntax element `listener_space_flag` is set to false, it means that the listener space of the audio scene can be inferred from elsewhere without signaling. For example, the listener space of the audio scene can be the same as the ROI of the video scene in an audio-video scene, and the listener space of the audio scene can be copied from the ROI of the video scene.
[0059] For a listener space item, the fifth syntax element listener_space_subtype can be signaled to indicate that the item is one of the following: an optimal point, a listening space with a finite elevation range, a listening space with a rectangular prism shape, a listening space with a polyhedral shape, a listening space with a spherical shape, a listening space with a rolling ball shape, etc.
[0060] Table 1 shows an example syntax table representing the space of interest for an audio scene. In Table 1, the syntax element `space_of_interest_type` indicates the type of item in the space of interest for the audio scene. This item can be a listener space, an audio channel configuration, or an audio object configuration. The syntax elements `listener_space_count`, `audio_channel_config_count`, and `audio_object_config_count` represent the total number of listener spaces, the total number of audio channel configurations, and the total number of audio object configurations, respectively. The syntax elements `listener_space_id`, `audio_channel_config_id`, and `audio_object_config_id` represent the identifier indexes of the listener space, the audio channel configuration, and the audio object configuration, respectively. The syntax element `listener_space_flag` indicates whether the listener space can be represented by a subtype of the listener space. The syntax element `listener_space_subtype` represents the subtype of the listener space. The subtype of the listener space can be one of the following: one or more optimal points, an auditory space with a finite elevation range, an auditory space with a rectangular prism shape, an auditory space with a polyhedral shape, an auditory space with a spherical shape, an auditory space with a rolling ball shape, etc.
[0061] Table 1
[0062]
[0063] For audio encoders, decoders, renderers, or other processors, a fixed-length flag `space_of_interest_selection` can be signaled for each listener space, audio channel, and audio object to indicate whether the corresponding item is enabled for a given audio encoder, decoder, renderer, or other processor. For example, a "1" value in the flag can indicate that the corresponding item (listener space, audio channel, or audio object) is enabled, and a "0" value in the flag can indicate that the corresponding item is disabled.
[0064] In this embodiment, the audio channel configuration can be a collection of multiple audio channels, which can be further indicated by the identifier index of these channels. Alternatively, the audio channel configuration can be a specific audio channel.
[0065] In this embodiment, the audio object configuration can be a collection of multiple audio objects, which can be further indicated by the identifier index of these objects. Alternatively, the audio object configuration can be a specific audio object.
[0066] Table 2 shows another example syntax table representing the space of interest for an audio scene.
[0067] Table 2
[0068]
[0069]
[0070]
[0071] II. Flowchart
[0072] Figure 5 A flowchart of a general example process (500) according to an embodiment of the present disclosure is shown. In various embodiments, the process (500) is executed by processing circuitry, such as... Figure 6 The processing circuit shown. In some embodiments, the process (500) is implemented as software instructions, so that when the processing circuit executes the software instructions, the processing circuit executes the process (500).
[0073] The process (500) typically begins at step (S501), where the process (500) decodes audio scene data of an audio scene. The audio scene data includes (i) audio content representing a plurality of items of the audio scene, and (ii) a first syntax element indicating the type of a subset of the plurality of items. The subset of items represents the space of interest of the audio scene. The process (500) then proceeds to step (S520).
[0074] In step (S520), process (500) determines a portion of the audio content for the subset of the plurality of items based on the type of the subset of the plurality of items indicated in the first syntax element. Then, process (500) proceeds to step (S530).
[0075] In step (S530), process (500) presents the determined portion of the audio content. Then, process (500) ends.
[0076] In one embodiment, the first syntax element indicates that the type of a subset of the plurality of items is one of a type associated with listener space, a type associated with audio channel configuration, or a type associated with audio object configuration.
[0077] In one embodiment, the audio scene data includes a second syntax element indicating the number of subsets of the plurality of items.
[0078] In one embodiment, the second syntax element indicates that the number of subsets of the plurality of items is greater than 1, and the audio scene data includes a third syntax element that indicates the identifier index of each item in the subset of the plurality of items.
[0079] In one embodiment, the first syntax element indicates that the type of a subset of the plurality of items is a type associated with the listener space, and the audio scene data includes a fourth syntax element indicating whether to signal the subtype of the listener space.
[0080] In one embodiment, the fourth syntax element indicates signaling a subtype of the listener space, and the audio scene data includes a fifth syntax element indicating the subtype of the listener space.
[0081] In one embodiment, the fourth syntax element indicates that no signal is sent to the subtype of the listener space, and the subtype of the listener space is determined based on the video scene.
[0082] In one embodiment, the subtype of the listener space is either a type associated with the best point of the audio scene or a type associated with the auditory space.
[0083] III. Computer System
[0084] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 6 A computer system (600) is shown, which is adapted to implement certain embodiments of the disclosed subject matter.
[0085] The computer software can be encoded using any suitable machine code or computer language, and code including instructions can be created through mechanisms such as assembly, compilation, and linking. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.
[0086] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0087] Figure 6 The components shown for the computer system (600) are exemplary in nature and are not intended to limit the scope or functionality of the computer software used to implement embodiments of this disclosure. Nor should the configuration of the components be construed as having any dependency or requirement on any component or combination thereof shown in the exemplary embodiments of the computer system (600).
[0088] The computer system (600) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through tactile input (e.g., keyboard input, swiping, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from still cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0089] Human-machine interface input devices may include one or more of the following (only one is shown): keyboard (601), mouse (602), touchpad (603), touch screen (610), data glove (not shown), joystick (605), microphone (606), scanner (607), camera (608).
[0090] The computer system (600) may also include certain human-machine interface (HMI) output devices. Such HMI output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such HMI output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (610), data gloves (not shown), or joystick (605), but may also include tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (609), headphones (not shown)), visual output devices (e.g., screens (610) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each with or without touchscreen input functionality, each with or without tactile feedback functionality—some of which may output two-dimensional or more three-dimensional visual outputs by means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown). These visual output devices (e.g., screens (610)) may be connected to the system bus (648) via a graphics adapter (650).
[0091] The computer system (600) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs (CD / DVD ROM / RW) (620) or similar media (621) with CD / DVD, thumb drives (622), removable hard disk drives or solid-state drives (623), conventional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.
[0092] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0093] The computer system (600) may also include a network interface (654) leading to one or more communication networks (655). For example, the one or more communication networks (655) may be wireless, wired, or optical. The one or more communication networks (655) may also be local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), vehicular and industrial networks, real-time networks, latency-tolerant networks, and so on. Examples of the one or more communication networks (655) also include Ethernet, wireless LANs, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANbus), and so on. Some networks typically require an external network interface adapter for connection to certain general-purpose data ports or peripheral buses (649) (e.g., a USB port on the computer system (600)); other systems are typically integrated into the core of the computer system (600) via a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). By using any of these networks, the computer system (600) can communicate with other entities. This communication can be unidirectional, for receiving only (e.g., wireless television), unidirectional, for sending only (e.g., CAN bus to certain CAN bus devices), or bidirectional, such as via a local area or wide area digital network to other computer systems. Each of the aforementioned networks and network interfaces can use certain protocols and protocol stacks.
[0094] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core (640) of the computer system (600).
[0095] The core (640) may include one or more central processing units (CPU) (641), graphics processing units (GPUs) (642), dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) (643), task-specific hardware accelerators (644), etc. These devices, as well as read-only memory (ROM) (645), random access memory (646), internal mass storage (e.g., internal non-user-accessible hard disk drives, solid-state drives, etc.) (647), etc., can be connected via a system bus (648). In some computer systems, the system bus (648) may be accessed as one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (648) or connected via a peripheral bus (649). Peripheral bus architectures include External Peripheral Component Interconnect (PCI), Universal Serial Bus (USB), etc.
[0096] The CPU (641), GPU (642), FPGA (643), and accelerator (644) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (645) or RAM (646). Transient data can also be stored in RAM (646), while permanent data can be stored, for example, in internal mass storage (647). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (641), GPUs (642), mass storage (647), ROM (645), RAM (646), etc.
[0097] The computer-readable medium may contain computer code for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purposes of this disclosure, or they may be media and code well-known and usable by those skilled in the art of computer software.
[0098] By way of example and not limitation, a computer system having an architecture (600), particularly a core (640), can provide functionality as a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the aforementioned user-accessible mass storage, as well as specific memory of the core (640) that is non-volatile, such as internal mass storage (647) or ROM (645). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the core (640). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the core (640), particularly the processor therein (including a CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (646) and modifying such data structures according to software-defined processes. Alternatively or as an alternative, the computer system may provide logic hardwired or otherwise incorporated into circuitry (e.g., an accelerator (644)) that may replace or operate with the software to perform the specific process or a specific portion of the specific process described herein. References to software may include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuitry storing the execution of software (such as an integrated circuit (IC)), circuitry containing execution logic, or both. This disclosure includes any suitable combination of hardware and software.
[0099] While this disclosure has described several exemplary embodiments, various modifications, arrangements, and equivalent substitutions of the embodiments are within the scope of this disclosure. Therefore, it should be understood that those skilled in the art can design various systems and methods that, while not explicitly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.
Claims
1. A method for representing an interest space of an audio scene, characterized in that, The method includes: The audio scene data of the audio scene is decoded, the audio scene data including (i) audio content representing multiple items of the audio scene, (ii) a first syntax element indicating that the type of a subset of the multiple items is a type associated with the listener space, and (iii) a second syntax element indicating whether to signal the subtype of the listener space. Based on the type of the subset of items associated with the listener space indicated in the first grammatical element, a portion of the audio content for the subset of the plurality of items is determined; and The defined portion of the audio content is presented.
2. The method according to claim 1, characterized in that, The audio scene data includes a third grammatical element indicating the number of subsets of the plurality of items.
3. The method according to claim 2, characterized in that, The third syntax element indicates that the number of subsets of the plurality of items is greater than 1, and the audio scene data includes a fourth syntax element that indicates the identifier index of each item in the subset of the plurality of items.
4. The method according to claim 1, characterized in that, The second syntax element indicates signaling the subtype of the listener space, and the audio scene data includes a fifth syntax element indicating the subtype of the listener space.
5. The method according to claim 1, characterized in that, The second syntax element indicates that no signal is sent to the subtype of the listener space, and the subtype of the listener space is determined based on the video scene.
6. The method according to claim 1, characterized in that, The subtype of the listener space is either a type associated with the best point of the audio scene or a type associated with the auditory space.
7. A device for representing a space of interest in an audio scene, characterized in that, The device includes: The processing circuit is configured to perform the method described in any one of claims 1-6.
8. A device for representing a space of interest in an audio scene, characterized in that, The device includes: A decoding module is used to decode the audio scene data of the audio scene, the audio scene data including (i) audio content representing multiple items of the audio scene, (ii) a first syntax element indicating that the type of a subset of the multiple items is a type associated with the listener space, and (iii) a second syntax element indicating whether to signal the subtype of the listener space. A determining module is configured to determine a portion of the audio content for the subset of the plurality of items based on the type of the subset of items associated with the listener space indicated in the first syntactic element; and The defined portion of the audio content is presented.
9. The apparatus according to claim 8, characterized in that, The audio scene data includes a third grammatical element indicating the number of subsets of the plurality of items.
10. The apparatus according to claim 9, characterized in that, The third syntax element indicates that the number of subsets of the plurality of items is greater than 1, and the audio scene data includes a fourth syntax element that indicates the identifier index of each item in the subset of the plurality of items.
11. The apparatus according to claim 8, characterized in that, The second syntax element indicates signaling the subtype of the listener space, and the audio scene data includes a fifth syntax element indicating the subtype of the listener space.
12. The apparatus according to claim 8, characterized in that, The second syntax element indicates that no signal is sent to the subtype of the listener space, and the subtype of the listener space is determined based on the video scene.
13. The apparatus according to claim 8, characterized in that, The subtype of the listener space is either a type associated with the best point of the audio scene or a type associated with the auditory space.
14. A non-volatile computer-readable storage medium, characterized in that, Used to store instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Audio Object Processing Based on Spatial Listener Information
US20180098173A1