Media processing method, apparatus and storage medium

By performing layered description and processing of the space of interest in the audio scene, the problem of matching audio with physical motion in virtual reality and augmented reality is solved, achieving a realistic interactive experience and efficient rendering effects.

CN116438813BActive Publication Date: 2026-05-26TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-06-02
Publication Date
2026-05-26

Smart Images

  • Figure CN116438813B_ABST
    Figure CN116438813B_ABST
Patent Text Reader

Abstract

Aspects of the disclosure provide methods and apparatuses for audio processing. In some examples, an apparatus for media processing includes processing circuitry. The processing circuitry receives an audio input associated with a hierarchical description of a space of interest in an audio scene. The space of interest includes a plurality of subspaces. The hierarchical description includes a first layer and a second layer. The first layer has a common node with a first value that is a common attribute value of two or more of the plurality of subspaces. The second layer has separate nodes respectively associated with each of the plurality of subspaces. The processing circuitry determines the plurality of subspaces of the space of interest based on the hierarchical description and renders an audio output based on the audio input in response to a location of a subject of the audio scene being in the space of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This disclosure claims priority to U.S. Patent Application No. 17 / 751425, filed May 23, 2022, entitled “LAYERED DESCRIPTION OF SPACE OF INTEREST,” which claims the benefit of U.S. Provisional Application No. 63 / 217442, filed July 1, 2021, entitled “Layered Description of Space of Interest,” the entire disclosure of the earlier application of which is incorporated herein by reference. Technical Field

[0003] This disclosure describes embodiments that are generally related to audio processing. Background Technology

[0004] The background description provided herein is for the purpose of presenting the general content of this disclosure. The extent of the work of the currently named inventors described in the background section and various aspects of this specification does not indicate that it was prior art at the time of filing of this application, nor is it expressly or implied that it was acknowledged as prior art to this disclosure.

[0005] In virtual reality or augmented reality applications, to give users the feeling of being present in the application's virtual world, audio within the application's virtual scene is perceived as real-world, with the sounds originating from virtual characters within the associated virtual scene. In some examples, the user's physical movements in the real world are perceived as having matching movements within the application's virtual scene. Furthermore, and importantly, users can interact with the virtual scene using audio that is perceived as realistic and matches their experience in the real world. Summary of the Invention

[0006] This disclosure provides methods and apparatus for audio processing. In some examples, the apparatus for media processing includes processing circuitry. The processing circuitry receives audio input associated with a hierarchical description of a space of interest in an audio scene. The space of interest includes multiple subspaces. The hierarchical description includes a first layer and a second layer. The first layer has common nodes, which have first values, which are common attribute values ​​of two or more of the multiple subspaces. The second layer has individual nodes associated with each of the multiple subspaces. The processing circuitry determines the multiple subspaces of the space of interest based on the hierarchical description and renders audio output based on the audio input in response to the location of a subject of the audio scene within the space of interest.

[0007] In some examples, multiple subspaces are defined by rectangles with at least position, orientation, and size attributes.

[0008] According to some aspects of this disclosure, the public node identifier is for the name of the attribute, the first value is the attribute value of the attribute, and the processing circuit can retrieve the first value from the public node in the first layer as the attribute value of the attribute for the subspace in the plurality of subspaces.

[0009] According to some aspects of this disclosure, the public node identifies the name of the attribute and the index of the subfield of the attribute, the first value is the subfield attribute value for the subfield of the attribute, and the processing circuit retrieves the first value from the public node in the first layer as the subfield attribute value for the subfield of the attribute in the subspace of the multiple subspaces.

[0010] In some examples, a common node with a first value is common to multiple subspaces, and the processing circuit retrieves the first value from the common node in the first layer as an attribute value for each of the multiple subspaces.

[0011] In some examples, a common node with a first value is common to a subset of multiple subspaces. The processing circuit, in response to a lack of a value for an attribute in a first individual node associated with a first subspace, retrieves the first value from the common nodes in the first layer as the attribute value for the attribute in the first subspace. Furthermore, in response to the presence of a second value associated with an attribute in a second individual node associated with a second subspace, the processing circuit retrieves the second value associated with the attribute in the second individual node.

[0012] In some examples, a common node with a first value is common to a subset of multiple subspaces. In response to a first individual node associated with a first subspace lacking an attribute value, the processing circuit retrieves the first value from the common nodes in the first layer as the attribute value of the first subspace. Furthermore, the processing circuit retrieves the difference associated with the attribute of the second subspace from a second individual node associated with the second subspace, and calculates a second value for the attribute of the second subspace based on the first value and the difference.

[0013] In some examples, the processing circuitry receives a bitstream carrying a hierarchical description of the audio input and the space of interest as metadata for the audio input; and decodes the bitstream to obtain the hierarchical description of the audio input and the space of interest.

[0014] In some examples, the processing circuitry ignores the audio input and does not render it in response to the fact that the location of the subject of the audio scene is outside the space of interest.

[0015] This disclosure also provides a non-transitory computer-readable storage medium for storing instructions that, when executed by a computer, cause the computer to perform a method for audio processing. Attached Figure Description

[0016] Other features, properties, and various advantages of the disclosed subject will become more apparent from the following detailed description and accompanying drawings, in which:

[0017] Figure 1 A schematic diagram is shown of an environment using 6 degrees of freedom (6 DoF) in some examples.

[0018] Figure 2 A block diagram of a media system according to an embodiment of the present disclosure is shown.

[0019] Figure 3 An audio scene, referred to as the canyon scene in some examples, is shown.

[0020] Figure 4 A description of the space of interest in the canyon scene is shown.

[0021] Figure 5 A syntax for a hierarchical description of a space of interest according to embodiments of this disclosure is shown.

[0022] Figure 6 A hierarchical description of the space of interest is shown in some examples.

[0023] Figure 7 A syntax for a hierarchical description of a space of interest according to embodiments of this disclosure is shown.

[0024] Figure 8 A hierarchical description of the space of interest is shown in some examples.

[0025] Figure 9 A hierarchical description of the space of interest is shown in some examples.

[0026] Figure 10 A hierarchical description of the space of interest is shown in some examples.

[0027] Figure 11 A flowchart outlining some embodiments of the process according to this disclosure is shown.

[0028] Figure 12 A flowchart outlining some embodiments of the process according to this disclosure is shown.

[0029] Figure 13 A flowchart outlining some embodiments of the process according to this disclosure is shown.

[0030] Figure 14 This is a schematic diagram of a computer system according to one embodiment. Detailed Implementation

[0031] This disclosure provides techniques for describing the space of interest (OOI) in an audio scene. Specifically, these techniques can provide a hierarchical description of the OIO in an audio scene. This hierarchical description of the OIO in an audio scene can provide compression information for audio encoding, transmission, and rendering.

[0032] Typically, an audio scene is a collection of semantically consistent sound segments represented by several primary sound sources. Therefore, an audio scene can be modeled as a set of sound sources. In some examples, an audio scene is dominated by several sets of sound sources. The space of interest in an audio scene can be defined by the boundaries of the space of interest considered within the audio scene. The space of interest in an audio scene can be utilized in audio encoding, processing, rendering, and other processes.

[0033] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0034] According to some aspects of this disclosure, some technologies attempt to create or mimic the physical world through digital simulations known as immersive media. Immersive media processing can be implemented according to immersive media standards, such as the Moving Picture Experts Group Immersive (MPEG-I) series of standards, including "immersive audio," "immersive video," and "system support." Immersive media standards can support VR or AR presentations, where users can navigate and interact with an environment using six degrees of freedom (6 DoF), including spatial navigation (x, y, z) and user head orientation (yaw, pitch, roll).

[0035] Figure 1 The diagram illustrates an environment using 6 degrees of freedom (6 DoF) in some examples. 6 DoF can be represented by spatial navigation (x, y, z) and user head orientation (yaw, pitch, roll).

[0036] According to one aspect of this disclosure, immersive media can be used to give users the feeling that they actually exist in a virtual world. In some examples, the audio of the scene is perceived as being in the real world, where the sound originates from associated visual characters. For example, the correct location and distance of the sound within the scene are perceived. The user's physical movement in the real world is perceived as having matching movement within the virtual world scene. Furthermore, the user can interact with the scene, evoking sounds that are perceived as real and match the user's experience in the real world.

[0037] Typically, a region of interest (ROI) comprises samples in a dataset that are identified for a specific purpose. The concept of ROI can be used in many application areas, such as medical imaging, geographic information systems, computer vision, and optical character recognition.

[0038] In an audio scene, a space of interest (SOI) can be described for a specific audio purpose. The SOI of an audio scene can be associated with an audio source that can cause audio effects within that SOI. In some examples, the SOI can be defined by an audio scene generator. The audio scene generator can define the SOI in three-dimensional (3D) space and the audio input as an audio source that can cause audio effects within the SOI. The audio input and the description of the SOI can be provided to an audio encoder. The audio encoder can encode the audio input into a bitstream, and the SOI can be included as metadata associated with the encoded audio. The bitstream can be provided to a client device. The client device can decode the audio content from the bitstream and render the audio according to the SOI. For example, when a game player moves to a SOI in a virtual world, audio content associated with that SOI is played.

[0039] Figure 2 A block diagram of a media system (200) according to an embodiment of the present disclosure is shown. The media system (200) can be used in a variety of applications, such as immersive media applications, augmented reality (AR) applications, virtual reality applications, video game applications, sports game animation applications, teleconferencing and telepresence applications, media streaming applications, etc.

[0040] The media system (200) includes a media server device (210) and multiple media client devices, such as Figure 2 The media client devices (260A) and 260B shown can be connected via a network (not shown). In one example, the media server device (210) may include one or more devices with audio and video encoding capabilities. In one example, the media server device (210) includes a single computing device, such as a desktop computer, laptop computer, server computer, tablet computer, etc. In another example, the media server device (210) includes a data center, server cluster, etc. The media server device (210) can receive video and audio content and compress the video and audio content into one or more encoded bitstreams according to a suitable media encoding standard. The encoded bitstreams can be transmitted via the network to the media client devices (260A) and 260B.

[0041] Media client devices (e.g., media client devices (260A) and (260B)) each include one or more devices with video and audio encoding capabilities for media applications. In one example, each media client device includes a computing device such as a desktop computer, laptop computer, server computer, tablet computer, wearable computing device, head-mounted display (HMD) device, etc. The media client device can decode the encoded bitstream according to a suitable media encoding standard. The decoded video and audio content can be used for media playback.

[0042] The media server device (210) can be implemented using any suitable technology. Figure 2 In the example, the media server device (210) includes processing circuitry (230) and interface circuitry (211) coupled together.

[0043] The processing circuitry (230) may include any suitable processing circuitry, such as one or more central processing units (CPUs), one or more graphics processing units (GPUs), application-specific integrated circuits (ASICs), etc. Figure 2 In one example, the processing circuitry (230) can be configured to include various encoders, such as an audio encoder (240), a video encoder (not shown), etc. In one example, one or more CPUs and / or GPUs can execute software to act as the audio encoder 240. In another example, the audio encoder (240) can be implemented using an application-specific integrated circuit (ASIC).

[0044] Interface circuit (211) can connect media server device (210) to a network. Interface circuit (211) may include a receiving unit for receiving signals from the network and a transmitting unit for sending signals to the network. For example, interface circuit (211) can transmit signals carrying encoded bit streams to other devices, such as media client device (260A), media client device (260B), etc., via the network. Interface circuit (211) can receive signals from media client devices such as media client devices (260A) and (260B).

[0045] The network is suitably coupled to the media server device (210) and media client devices (e.g., media client devices (260A) and (260B)) via wired and / or wireless connections, such as Ethernet connections, fiber optic connections, WiFi connections, cellular network connections, etc. The network may include network server devices, storage devices, network devices, etc. The components of the network are suitably coupled together via wired and / or wireless connections.

[0046] Media client devices (e.g., media client devices (260A) and (260B)) are configured to decode the encoded bitstream. In one example, each media client device may perform video decoding to reconstruct a displayable sequence of video frames and may perform audio decoding to generate an audio signal for playback.

[0047] Media client devices, such as media client devices (260A) and (260B), can be implemented using any suitable technology. Figure 2 The examples show a media client device (260A), but not limited to a user device with a headset display that can be used by user A, and a media client device (260B), but not limited to a smartphone used by user B.

[0048] exist Figure 2 In the middle, the media client device (260A) includes, for example, Figure 2 The interface circuitry (261A) and processing circuitry (270A) are coupled together as shown. The media client device (260B) includes, for example, the interface circuitry (261A) and processing circuitry (270A) coupled together as shown. Figure 2 The interface circuit (261B) and processing circuit (270B) are shown coupled together.

[0049] The interface circuit (261A) connects the media client device (260A) to a network. The interface circuit (261A) may include a receiving section for receiving signals from the network and a transmitting section for sending signals to the network. For example, the interface circuit (261A) may receive signals carrying data, such as a signal carrying an encoded bitstream from the network.

[0050] The processing circuitry (270A) may include suitable processing circuitry, such as a CPU, GPU, application-specific integrated circuit, etc. The processing circuitry (270A) may be configured to include various components, such as an audio decoder (271A), a renderer (272A), etc.

[0051] In some examples, the audio decoder (271A) can decode the audio content in the encoded bitstream by selecting a decoding tool suitable for the encoding scheme. Furthermore, the renderer (272A) can generate a final digital product suitable for a media client device (260A) based on the audio content decoded from the encoded bitstream. It should be noted that the processing circuitry (270A) may include other suitable components (not shown), such as mixers, post-processing circuitry, etc., for further audio processing.

[0052] Similarly, the interface circuit (261B) can connect the media client device (260B) to the network. The interface circuit (261B) may include a receiving unit for receiving signals from the network and a transmitting unit for sending signals to the network. For example, the interface circuit (261B) can receive signals carrying data, such as signals carrying encoded bit streams from the network.

[0053] The processing circuitry (270B) may include suitable processing circuitry, such as a CPU, GPU, application-specific integrated circuit, etc. The processing circuitry (270B) may be configured to include various components, such as an audio decoder (271B), a renderer (272B), etc.

[0054] In some examples, the audio decoder (271B) can decode the audio content in the encoded bitstream by selecting a decoding tool suitable for the encoding scheme. Furthermore, the renderer (272B) can generate a final digital product suitable for a media client device (260B) based on the audio content decoded from the encoded bitstream. It should be noted that the processing circuitry (270A) may include other suitable components (not shown), such as a mixer, post-processing circuitry, etc., for further audio processing.

[0055] According to some aspects of this disclosure, a hierarchical description of the space of interest in an audio scene is used in the media system (200). Media server device (210), media client devices (e.g., media client devices (260A) and (260B)) can process the hierarchical description of the space of interest in the audio scene. For example, processing circuitry (230), processing circuitry (270A), processing circuitry (270B), etc., can determine the space of interest in the audio scene based on the hierarchical description of the space of interest in the audio scene.

[0056] In some examples, for an audio scene, the media server device (210) receives audio input and a hierarchical description of the space of interest in the audio scene from an audio source (201) (e.g., an audio injection server, an audio scene generator device, etc.). In some examples, the audio source (201) includes computing circuitry, such as a desktop computer, a laptop computer, a server computer, a tablet computer, etc. The computing circuitry can generate audio input for the audio scene and generate a hierarchical description of the space of interest in the audio scene. The audio input for the audio scene and the hierarchical description of the space of interest in the audio scene can be provided to the media server device (210).

[0057] In some embodiments, the media server device (210) may determine the appropriate media content to be sent to the media client device, encode the media content into a bitstream, and send the bitstream to the media client device. In some examples, the media server device (210) may determine the audio content for the media client device (260A) based on information provided by the media client device (260A), such as scene information in an application (e.g., a game scene, a VR scene, etc.). The media server device (210) may determine the audio input for the scene (referred to as an audio scene in audio processing) as the audio content. The media server device (210) may encode the audio content along with other suitable information into a bitstream. In one example, the bitstream may carry the encoded audio content for the audio scene and a hierarchical description of the space of interest (SOI) of the audio scene. The hierarchical description of the SOI of the audio scene may be metadata for the audio content. The media server device (210) may send the bitstream to the media client device (260A). It should be noted that in some examples, the encoded audio content and the hierarchical description of the SOI of the audio scene may be sent separately.

[0058] When the media client device (260A) receives encoded audio content and a hierarchical description of the space of interest (SOI) of the audio scene, the audio decoder (271A) can decode the audio content. The renderer (272A) can then generate a final digital product suitable for the media client device (260A) based on the audio content and the SOI of the audio scene. For example, the renderer (272A) can render the audio content when a subject in the application moves into the SOI of the audio scene. In one example, the audio content is ignored and not rendered when the subject moves out of the SOI.

[0059] In some examples, the space of interest (POI) may have an irregular shape. For ease of description, the POI can be described as a combination of several regions (also called subspaces) with regular shapes. Each of the several subspaces can be described using attributes. According to some aspects of this disclosure, some of these subspaces may share some attribute values. In some examples, the hierarchical description of the POI of an audio scene can describe the shared attribute values ​​of two or more subspaces in layers separate from the individual attribute values ​​of the subspaces, thus allowing for a more compact description of the POI.

[0060] In the following description, rectangles are used to describe subspaces of regular shape that represent the space of interest. It should be noted that in some examples, other regular shapes, such as spheres, cylinders, cubes, etc., may be used.

[0061] Figure 3 An audio scene, sometimes referred to as a canyon scene in some examples, is shown. Within an audio scene, the space of interest can be specified using one or more rectangles. Figure 3 In the example, the space of interest (310) in the canyon scene can be described by four overlapping rectangles (301)-(304).

[0062] In some examples, four attributes can be used to define each rectangle: identifier (id) attribute, position attribute, orientation attribute, and size attribute.

[0063] The identifier attribute of a rectangle can have a value that indicates the rectangle. For example, rectangle (301) can be identified by the identifier "box:Box1"; rectangle (302) can be identified by the identifier "box:Box2"; rectangle (303) can be identified by the identifier "box:Box3"; and rectangle 304 can be identified by the identifier "box:Box4".

[0064] In some examples, the position attribute of the rectangle can include three values ​​corresponding to the coordinates of the rectangle's center position in 3D space, such as x, y, and z. In one example, the first value with index "1" corresponds to the x-coordinate of the center position, the second value with index "2" corresponds to the y-coordinate of the center position, and the third value with index "3" corresponds to the z-coordinate of the center position.

[0065] In some examples, the orientation property of the rectangle may include three values ​​corresponding to the rotation angles along the X, Y, and Z axes at the center of the rectangle. In one example, the first value with index "1" corresponds to the rotation angle along the Y axis at the center of the rectangle, the second value with index "2" corresponds to the rotation angle along the X axis at the center of the rectangle, and the third value with index "3" corresponds to the rotation angle along the Z axis at the center of the rectangle.

[0066] In some examples, the size property of the rectangle may include three values ​​corresponding to the side lengths along the X, Y, and Z axes. In one example, the first value with index "1" corresponds to the side length of the rectangle along the X-axis, the second value with index "2" corresponds to the side length of the rectangle along the Y-axis, and the third value with index "3" corresponds to the side length of the rectangle along the Z-axis.

[0067] Figure 4 A description (400) of the space of interest (310) in the canyon scene is shown. The description (400) includes individual nodes for rectangles (301)-(304) respectively. Specifically, the description (400) includes a description node (401) for rectangle (301), a description node (402) for rectangle (302), a description node (403) for rectangle (303), and a description node (404) for rectangle (304).

[0068] It should be noted that rectangles (301)-(304) have the same (common) attribute values, which are listed in the individual nodes of rectangles (301)-(304).

[0069] Some aspects of this disclosure provide a hierarchical description of the space of interest in an audio scene. The hierarchical description may include a first layer for common attribute values ​​and a second layer for non-common attribute values. In the hierarchical description, the common attribute values ​​of two or more subspaces (e.g., bounding boxes) may be explicitly listed in the first layer. For all relevant descriptions of a subspace, the common attribute values ​​may be signaled only once. Other non-common attribute values ​​may be listed separately in the second layer. The hierarchical description is more compact than description (400).

[0070] In some embodiments, a first layer of description for common attribute values ​​may be presented first, followed by a second layer of description for non-common attribute values ​​for each individual rectangle. The first layer of description may include one or more common nodes that describe the common attribute values, respectively. The second layer of description may include individual nodes for each of the multiple rectangles, respectively. Common attribute values ​​presented as common attribute values ​​in common nodes will not be listed again in the individual nodes for rectangles unless necessary, for example, to avoid duplicate information. For each rectangle in the spatial description, common attribute values ​​already in common nodes will not be listed. Instead, the rectangles will share common attribute values ​​listed in common nodes by default.

[0071] Figure 5 A syntax (500) for a hierarchical description of a space of interest according to an embodiment of the present disclosure is shown. The syntax includes a first layer (510) for descriptions of common attribute values ​​shared by two or more rectangles within a rectangle, and a second layer (520) for descriptions of non-common attribute values ​​of individual rectangles.

[0072] The first layer (510) includes multiple public nodes. Each public node may include a name for identifying an attribute (e.g., shown by (511)) and a value for the public attribute value (e.g., shown by (512)).

[0073] In one example, the four rectangles of interest have the same dimensions, such as "20.0 2.50 15.0". Figure 6 A hierarchical description (600) of a space of interest for four rectangular boxes of the same size is shown using syntax (500). The hierarchical description (600) includes a first layer (610) of descriptions of common nodes shared by two or more rectangular boxes and a second layer (620) of descriptions of non-common attribute values ​​for each individual rectangular box.

[0074] Specifically, the first layer (610) includes a common node. The common node has a name “size” (e.g., shown by (611)) for identifying the size attribute and a value “20.0 2.50 15.0” (e.g., shown by 612) that specifies the common attribute value as the size of the rectangular frame. The second layer (620) includes descriptions of non-common attribute values ​​for each individual rectangular frame, such as position and orientation attributes. For example, the second layer (620) includes four individual nodes for each of the four rectangular frames.

[0075] exist Figure 6In the example, individual nodes do not include the size attribute. The rectangles identified by "box:Box1", "box:Box2", "box:Box3", and "box:Box4" can reference the common node to retrieve the size attribute value "20.0 2.50 15.0" for the size attribute.

[0076] In some examples, an attribute can have multiple subfields. For instance, the size attribute of a rectangle might have a first subfield for the length of its side along the X-axis, a second subfield for the length of its side along the Y-axis, and a third subfield for the length of its side along the Z-axis. In some examples, two or more rectangles may not share the entire size attribute, but they may share one or more subfields of the size attribute.

[0077] In one example, Figure 3 Canyon scene examples and Figure 4 In the description, the four rectangles (301)-(304) of the space of interest (310) have the same height of 2.5 (side length along the Y-axis), which is the second subfield in the size attribute. In one example, the height information (side length along the Y-axis) of the rectangles (301)-(304) can be regarded as a common attribute (e.g., a common subfield attribute) and can be notified only once by signaling.

[0078] In another example, the four rectangles (301)-(304) share some subfields of the direction attribute. Specifically, the four rectangles (301)-(304) share the second subfield value (e.g., "0") and the third subfield value (e.g., "0") of the direction attribute. In one example, the second and third subfields of the direction attribute can be considered as common subfields of the direction attribute and are signaled only once for all rectangles (301)-(304).

[0079] In one embodiment, the hierarchical description of the space of interest may include a first level describing common attribute values ​​and / or common subfield attribute values, and a second level describing non-common attribute values ​​and / or non-common subfield attribute values ​​for individual rectangles. For example, common subfields of an attribute may be listed in a common node (also referred to as a parent in one example) in the first level. A common node may include the name of the attribute, an index for the common subfield, and the value for the common subfield attribute value. In the hierarchical description of the space of interest, for a rectangle, if the subfields of an attribute are already listed in a common node, the subfield values ​​of the attribute within the rectangle will not be listed in the individual node associated with the rectangle in the second level. Instead, in one example, the attributes of the rectangle will be shared by default with the common subfield attribute values ​​listed in the common node.

[0080] Figure 7A syntax (700) for a hierarchical description of a space of interest according to an embodiment of the present disclosure is shown. The syntax includes a first layer (710) for a description of common attribute values ​​and / or common subfield attribute values ​​shared by two or more rectangles within a rectangle, and a second layer (720) for a description of non-common attribute values ​​and / or non-common subfield attribute values ​​for individual rectangles.

[0081] The first layer (710) includes multiple common nodes. Each common node may include a name for identifying an attribute (e.g., shown by (711)), an index for identifying a subfield (e.g., shown by (713)), and a value for the common subfield attribute value (e.g., shown by (712)).

[0082] Figure 8 A hierarchical description (800) for a space of interest (310) having four rectangles (301)-(304) is shown according to the syntax (700). The hierarchical description (800) includes a first layer (810) for descriptions of common nodes shared by two or more rectangles, and a second layer (820) for descriptions of non-common attribute values ​​and / or non-common subfield attribute values ​​of the individual rectangles.

[0083] Specifically, the first layer (810) includes a first common node for the subfields of the orientation attribute and a second common node for the subfields of the size attribute. The first common node has a name “orientation” for identifying the orientation attribute (e.g., shown by (811)), an index “23” for identifying the second and third subfields (e.g., shown by (813)), and a value “0.00 0.00” for specifying the common subfield attribute value (e.g., shown by (812)). Thus, the first common node lists the common value of the second subfield of the orientation attribute as “0.00” and the common value of the third subfield of the orientation as “0.00”.

[0084] Similarly, the second common node has a name "size" to identify the size attribute, an index "2" to identify the second subfield, and a value "2.50" to specify the value of the common subfield attribute. Therefore, the second common node lists the common value of the second subfield of the size attribute as "2.50".

[0085] The second layer (820) includes descriptions of non-public attribute values ​​and / or non-public subfield attribute values ​​for individual rectangles, such as position, orientation, and size attributes. Figure 8In the example, the second layer (820) includes four separate nodes for each of the four rectangles. The four rectangles share information in a common node. For each separate node, the direction attribute includes a first subfield value but excludes the second and third subfield values. The second and third subfield values ​​of the direction attribute can be retrieved from the first common node of the first layer (810).

[0086] Similarly, in the second layer (820), for each individual node, the size attribute includes a first subfield value and a third subfield value, but excludes a second subfield value. The second subfield value of the size attribute can be retrieved from the second common node in the first layer (810).

[0087] It is important to note that, Figure 8 In the example, the information in the first and second common nodes is shared by all four rectangles. In some embodiments, the information in the common nodes of the first layer does not need to be shared by all individual nodes in the second layer.

[0088] In one embodiment, the hierarchical description of the space of interest may include a first layer describing common attribute values ​​and / or common subfield attribute values ​​shared by subsets of individual nodes for subspaces (e.g., rectangles), and a second layer describing non-common attribute values ​​and / or non-common subfield attribute values ​​for individual subspaces (e.g., rectangles). If a subspace does not share common attribute values ​​specified in common nodes in the first layer, the individual nodes associated with the subspace in the second layer may list individual attribute values ​​that differ from the common attribute values.

[0089] In the second layer, for a rectangle, if the attribute value is the same as the value listed in the common node, the attribute will not be listed; if the attribute value of the rectangle is different from the common attribute, the attribute value of the rectangle can be listed.

[0090] In one example, rectangular box (301) has a side length of “29.66” along the X-axis, which is different from the other three rectangular boxes (302)-(304) which have a side length of “19.23” along the X-axis.

[0091] Figure 9 A hierarchical description (900) for a space of interest (310) having four rectangular boxes (301)-(304) is shown according to syntax 700. The hierarchical description (900) includes a first layer (910) for descriptions of common nodes for common attribute values ​​shared by two or more rectangular boxes, and a second layer (920) for descriptions of individual nodes for non-common attribute values ​​and / or non-common subfield attribute values ​​of individual rectangular boxes.

[0092] Specifically, the first layer (910) includes a first common node for the subfields of the orientation attribute and a second common node for the subfields of the size attribute. The first common node has a name “orientation” to identify the orientation attribute, an index “2 3” to identify the second and third subfields, and a value “0.00 0.00” to specify the common subfield attribute value. Therefore, the first common node lists the common value of the second subfield of the orientation attribute as “0.00” and the common value of the third subfield of the orientation attribute as “0.00”.

[0093] Furthermore, the second common node has a name “size” to identify the size attribute, an index “1 2” to identify the first and second subfields, and a value “19.23 2.50” to specify the common subfield attribute value for a subset of the rectangle (e.g., rectangles (302)-(304)). Therefore, the second common node lists the common value of the first and second subfields of the size attribute as “19.23 2.50”.

[0094] The second layer (920) includes descriptions of individual nodes for non-public attribute values ​​and / or non-public subfield attribute values ​​(e.g., position, orientation, and size attributes) for individual rectangles. In the second layer (920), for each individual node, the orientation attribute includes a first subfield value but excludes a second and third subfield value. The second and third subfield values ​​of the orientation attribute can be retrieved from the first public node in the first layer.

[0095] In the second layer (920), for individual nodes associated with a subset of rectangles (e.g., box:Box2, box:Box3, and box:Box4), the size attribute includes a third subfield value but excludes the first and second subfield values. The first and second subfield values ​​of the size attribute can be referenced from and retrieved from the second common node in the first layer (910).

[0096] In the second layer (920), in a separate node for a rectangle (e.g., box:Box1), the size attribute includes a first subfield value, a second subfield value, and a third subfield value. For example, as shown in (921), the size attribute of the rectangle "box:Box1" is listed together with all three subfield values ​​"29.66 2.50 20.47". Therefore, the information in a separate node can override the information in a common node.

[0097] In one embodiment, the hierarchical description of the space of interest may include a first layer describing common attribute values ​​and / or common subfield attribute values ​​shared by subsets of individual nodes for subspaces (e.g., rectangles), and a second layer describing non-common attribute values ​​and / or non-common subfield attribute values ​​for individual subspaces (e.g., rectangles). If a subspace does not share common attribute values ​​specified in common nodes in the first layer, the individual nodes associated with that subspace in the second layer may list differences between the non-common attribute values ​​of that subspace and the common attribute values ​​in the first layer.

[0098] In the second layer, for a rectangle, if an attribute value is the same as an attribute value listed in the common node, then that attribute value is not listed; if the attribute value of the rectangle is different from the common attribute value, then the difference between the attribute value of the rectangle and the common attribute value can be listed.

[0099] In one example, rectangle (301) has a side length of “29.66” along the X-axis, which is different from the other three rectangles (302)-(304) which have a side length of “19.23” along the X-axis. The difference between the side length of “29.66” and the side length of “19.23” is “10.43”.

[0100] Figure 10 A hierarchical description (1000) for a space of interest (310) having four rectangular boxes (301)-(304) is shown according to the syntax (700). The hierarchical description (1000) includes a first layer (1010) for descriptions of common nodes shared by two or more rectangular boxes and a second layer (1020) for descriptions of non-common attribute values ​​and / or non-common subfield attribute values ​​of individual rectangular boxes.

[0101] Specifically, the first layer (1010) includes a first common node for the subfields of the orientation attribute and a second common node for the subfields of the size attribute. The first common node has a name "orientation" to identify the orientation attribute, an index "2 3" to identify the second and third subfields, and a value "0.00 0.00" to specify the common subfield attribute values. Therefore, the first common node lists the common value of the second subfield of the orientation attribute as "0.00" and the common value of the third subfield of the orientation attribute as "0.00".

[0102] Furthermore, the second common node has a name “size” to identify the size attribute, an index “1 2” to identify the first and second subfields, and a value “19.23 2.50” to specify the common subfield attribute value for a subset of the rectangle (e.g., rectangles (302)-(304)). Therefore, the second common node lists the common value of the first and second subfields of the size attribute as “19.23 2.50”.

[0103] The second layer (1020) includes descriptions of individual nodes for non-public attribute values ​​and / or non-public subfield attribute values ​​(such as position, orientation, and size attributes) for individual rectangles. In the second layer (1020), within each individual node, the orientation attribute includes a first subfield value but excludes second and third subfield values. The second and third subfield values ​​of the orientation attribute can be retrieved from the first public node in the first layer (1010).

[0104] In the second layer (1020), in individual nodes for subsets of rectangles (e.g., box:Box2, box:Box3, and box:Box4), the size attribute includes a third subfield value but excludes the first and second subfield values. The first and second subfield values ​​of the size attribute can be referenced from and retrieved from the second common node in the first layer (1010).

[0105] In the second layer (1020), in the individual node associated with the rectangle (e.g., box:Box1), the size attribute includes a first subfield difference, a second subfield difference, and a third subfield value. For example, as shown by (1021), the size attribute of the rectangle "box:Box1" is listed as "10.43 0.00 20.47". By checking with the second common node, the first subfield value of the size attribute of "box:Box1" can be recovered as the sum of "19.23" and "10.43", which equals 29.66, and the second subfield value of the size attribute of "box:Box1" can be recovered as the sum of "2.5" and "0", which equals 2.5.

[0106] Figure 11 A flowchart outlining a process (1100) according to an embodiment of the present disclosure is shown. The process (1100) may be performed by an audio source device, such as an audio source device (201). In some embodiments, the process (1100) is implemented as software instructions, so that the processing circuitry executes the process (1100) when the software instructions are executed. The process begins at (S1101) and proceeds to (S1110).

[0107] In (S1110), it is determined that two or more subspaces of a plurality of subspaces of the space of interest have common attribute values ​​for the attribute.

[0108] In some examples, multiple subspaces are rectangular boxes. Each rectangular box can be defined by at least position, orientation, and size attributes.

[0109] In (S1120), common nodes are formed in the first layer of the hierarchical description of the space of interest. Common nodes include common attribute values ​​for attributes.

[0110] In (S1130), attributes are removed from individual nodes associated with two or more subspaces. The individual nodes are located in the second layer of the hierarchical description of the space of interest.

[0111] In one example, the common node identifies the name of the attribute and the common attribute value. The common attribute value can be removed from a single node associated with two or more subspaces.

[0112] In another example, a common node identifies the attribute name, the index of the attribute's subfields, and the common attribute value as the common subfield attribute value. Common subfield attribute values ​​can be removed from individual nodes associated with two or more subspaces.

[0113] In some examples, a common node is common to multiple subspaces. Common attribute values ​​can be removed from each individual node associated with multiple subspaces.

[0114] In some examples, a common node is common to subsets of multiple subspaces. Common attribute values ​​can be removed from each individual node associated with a subset of multiple subspaces. For a subspace not in a subset, an individual node associated with that subset can include attribute values ​​that are different from the common attribute values ​​for the attribute.

[0115] In some examples, a common node is common to subsets of multiple subspaces. Common attribute values ​​can be removed from each individual node associated with a subset of multiple subspaces. For subspaces not within subsets, individual nodes associated with subsets can include the difference between the (subspace-specific) attribute value and the common attribute value.

[0116] Then, the process proceeds to (S1199) and terminates.

[0117] The process (1100) can be adjusted as appropriate. Steps in the process (1100) can be modified and / or omitted. Additional steps can be added. Any implementation in a suitable order can be used.

[0118] Figure 12A flowchart outlining a process (1200) according to an embodiment of the present disclosure is shown. The process (1200) may be executed by a media server device, such as a media server device (210). In some embodiments, the process (1200) is implemented as software instructions, so that the processing circuitry executes the process (1200) when the software instructions are executed. The process begins at (S1201) and proceeds to (S1210).

[0119] At (S1210), an audio input for an audio scene and a hierarchical description of the space of interest associated with the audio input are received. The space of interest includes multiple subspaces. The hierarchical description includes a first layer and a second layer. The first layer includes common nodes with a first value, which is a common attribute value of two or more subspaces. The second layer includes individual nodes that are associated with the multiple subspaces respectively.

[0120] In some examples, multiple subspaces are rectangular boxes. Each rectangular box can be defined by at least position, orientation, and size attributes.

[0121] In S1220, multiple subspaces of the space of interest are determined based on the hierarchical description.

[0122] In one example, the common node identifies the name of the attribute, and the first value is the attribute value. The first value can be retrieved from the common node in the first level as the attribute value for an attribute in multiple subspaces.

[0123] In another example, the common node identifies the name of the attribute and the index of the attribute's subfields, and the first value is the subfield attribute value for the subfield of the attribute. The first value can be retrieved from the common node in the first level as the subfield attribute value for the subfield of the attribute in multiple subspaces.

[0124] In some examples, a common node with a first value is common to multiple subspaces. The first value can be retrieved from the common node in the first level as the attribute value of an attribute in each of the multiple subspaces.

[0125] In some examples, a common node with a first value is common to subsets of multiple subspaces. In response to a first individual node associated with a first subspace lacking a value for an attribute, the first value is retrieved from the common nodes in the first level as the attribute value for the attribute in the first subspace. Furthermore, in response to a second individual node associated with a second subspace containing a second value associated with an attribute, the second value associated with the attribute in the second individual node is retrieved from the second individual node.

[0126] In some examples, a common node with a first value is common to subsets of multiple subspaces. In response to a first individual node associated with a first subspace lacking a value for an attribute, the first value is retrieved from the common nodes in the first level as the attribute value for the first subspace. Furthermore, the difference associated with the attribute of the second subspace is retrieved from the second individual node associated with the second subspace. Then, a second value for the attribute of the second subspace is calculated based on the first value and the difference, such as the sum of the first value and the difference.

[0127] In (S1230), in response to information provided from the client device, a bitstream carrying the audio input and a hierarchical description of the space of interest is sent to the client device. In one example, the information provided from the client device indicates a scene change in the audio scene. In another example, the information provided from the client device indicates that the movement of a subject in the application is associated with the space of interest, such as moving to the space of interest, etc.

[0128] Then, the process proceeds to (S1299) and terminates.

[0129] Process (1200) can be adjusted appropriately. Steps in process (1200) can be modified and / or omitted. Additional steps can be added. It can be implemented using any suitable order.

[0130] Figure 13 A flowchart outlining a process (1300) according to an embodiment of the present disclosure is shown. The process (1300) may be executed by a media client device, such as media client device (260A), media client device (260B), etc. In some embodiments, the process (1300) is implemented as software instructions, so that the processing circuitry executes the process (1300) when the software instructions are executed. The process begins at (S1301) and proceeds to (S1310).

[0131] At (S1310), audio input is received associated with a hierarchical description for a space of interest in an audio scene. The space of interest comprises multiple subspaces. The hierarchical description comprises a first layer and a second layer. The first layer has a common node, which has a first value, which is a common attribute value of two or more subspaces among the multiple subspaces. The second layer has individual nodes associated with each of the multiple subspaces respectively.

[0132] In some examples, a bitstream carrying a hierarchical description of the audio input and the space of interest (e.g., metadata for the audio input) is received. This bitstream is then decoded to obtain the hierarchical description of the audio input and the space of interest.

[0133] In some examples, multiple subspaces are rectangular boxes. Each rectangular box can be defined by at least position, orientation, and size attributes.

[0134] In (S1320), multiple subspaces of the space of interest are determined based on the hierarchical description.

[0135] In one example, the common node identifies the name of the attribute, and the first value is the attribute value. The first value can be retrieved from the common node in the first level as the attribute value for an attribute in multiple subspaces.

[0136] In another example, the common node identifies the name of the attribute and the index of the attribute's subfields, and the first value is the subfield attribute value for the subfield of the attribute. The first value can be retrieved from the common node in the first level as the subfield attribute value for the subfield of the attribute in multiple subspaces.

[0137] In some examples, a common node with a first value is common to multiple subspaces. The first value can be retrieved from the common node in the first level as the attribute value for an attribute in each of the multiple subspaces.

[0138] In some examples, a common node with a first value is common to subsets of multiple subspaces. In response to a first individual node associated with a first subspace lacking a value for an attribute, the first value is retrieved from the common nodes in the first level as the attribute value for the first subspace. Furthermore, in response to the presence of a second value associated with an attribute in a second individual node, the second value associated with the attribute of the second subspace is retrieved from the second individual node associated with the second subspace.

[0139] In some examples, a common node with a first value is common to subsets of multiple subspaces. In response to a first individual node associated with a first subspace lacking a value for an attribute, the attribute value for which the first value is an attribute of the first subspace is retrieved from the common nodes in the first level. Furthermore, the difference associated with the attribute of the second subspace is retrieved from the second individual node associated with the second subspace. Then, a second value for the attribute of the second subspace is calculated based on the first value and the difference.

[0140] At (S1330), in response to the location of the subject of the audio scene being within the space of interest, audio output is rendered based on the audio input. For example, the audio scene corresponds to a game scene in a game application, and the subject of the audio scene is the game player in the game application. In response to the game player moving into the space of interest of the game scene, audio output is rendered. In some examples, in response to the location of the subject of the audio scene being outside the space of interest, the audio input is ignored and no rendering is performed.

[0141] Then, the process proceeds to (S1399) and terminates.

[0142] Process (1300) can be adjusted as appropriate. Steps in process (1300) can be modified and / or omitted. Additional steps can be added. Any suitable order can be used for implementation.

[0143] The above-described techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 14 A computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0144] Computer software can be coded using any suitable machine code or computer language, which can be assembled, compiled, linked or similarly to create code including instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0145] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.

[0146] Figure 14 The components shown for the computer system (1400) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of the computer software used to implement embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any component or combination of components shown in the exemplary embodiments of the computer system (1400).

[0147] The computer system (1400) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keyboard input, swiping, data glove movement), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from still-view cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0148] Input human-machine interface devices may include one or more of the following: keyboard (1401), mouse (1402), touchpad (1403), touch screen (1410), data glove (not shown), joystick (1405), microphone (1406), (scanner 1407), and camera (1408) (only one of each is shown).

[0149] The computer system (1400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback from a touchscreen (1410), a data glove (not shown), or a joystick (1405), but may also include tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., screens (1410), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability, some of which may be able to output two-dimensional visual output or greater than three-dimensional output via, for example, a stereoscopic output device; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0150] The computer system (1400) may also include human-accessible storage devices and their associated media, such as CD / DVD ROM / RW (1420) or similar media (1421) with CD / DVD, thumb drives (1422), removable hard disk drives or solid-state drives (1423), conventional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated ROMs / ASICs / PLDs based on devices such as security dongles (not shown), etc.

[0151] Those skilled in the art will also understand that the term "computer-readable medium" used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0152] The computer system (1400) may also include an interface (1454) for connecting to one or more communication networks (1455). The network may be, for example, wireless, wired, or optical. Examples of local, wide area, metropolitan area, vehicular and industrial, real-time, and latency-tolerant networks include local area networks such as Ethernet and wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television; and vehicular and industrial networks including CANBus. Some networks typically require external network interface adapters attached to certain general-purpose data ports or peripheral buses (1449) (e.g., USB ports of the computer system (1400)); others are typically integrated into the core of the computer system (1400) by attaching to system buses as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcast television), send-only (e.g., CANbus to some CANbus device), or bidirectional (e.g., other computer systems using local or wide area digital networks). Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0153] The aforementioned human-machine interface device, human-accessible storage device and network interface can be attached to the core (1440) of the computer system (1400).

[0154] The core (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units (1443) in the form of field-programmable gate arrays (FPGAs), hardware accelerators (1444) for certain tasks, graphics adapters (1450), and the like. These devices, along with read-only memory (ROM) (1445), random access memory (1446), and internal mass storage (1447) such as internal non-user-accessible hard disk drives, SSDs, etc., can be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessed as one or more physical plugs to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1448) or connected to the core's system bus (1448) via a peripheral device bus (1449). In one example, a screen (1410) may be connected to a graphics adapter (1450). The architecture for peripheral buses includes PCI, USB, etc.

[0155] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (1445) or RAM (1446). Transient data can also be stored in RAM (1446), while permanent data can be stored, for example, in internal mass storage (1447). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage (1447), ROM (1445), RAM (1446), etc.

[0156] Computer-readable media may contain computer code for performing operations of various computer implementations. The media and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be those well known and available to those skilled in the art of computer software.

[0157] By way of example and not limitation, a computer system having an architecture (1400), particularly a core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with the user-accessible mass storage described above, as well as some memory of the non-transitory core (1440), such as kernel mass storage (1447) or ROM (1445). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the core (1440). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the core (1440), specifically causing the processor therein (including a CPU, GPU, FPGA, etc.), to execute the specific processing described herein or a specific portion of a specific processing, including defining data structures stored in RAM (1446) and modifying such data structures according to the software-defined processing. In addition, or alternatively, a computer system may provide functionality as a hardwired or otherwise embodied logic result in circuitry (e.g., an accelerator (1444)), which may replace or operate with software to perform the specific processing or a specific portion of the specific processing described herein. References to software may include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0158] Although several exemplary embodiments have been described in this disclosure, there are also variations, substitutions, and various alternatives that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.

Claims

1. A method of media processing in a device, characterized by, The method includes: Receive audio input associated with a hierarchical description of a space of interest in an audio scene, the space of interest comprising multiple subspaces, the hierarchical description comprising a first layer and a second layer, the first layer having a common node, the common node having a first value, the first value being a common attribute value of two or more of the multiple subspaces, and the second layer having individual nodes respectively associated with each of the multiple subspaces. The plurality of subspaces of the space of interest are determined based on the hierarchical description; and The audio output is rendered based on the audio input, in response to the location of the subject of the audio scene within the space of interest.

2. The method of claim 1, wherein, The plurality of subspaces are rectangular boxes defined by at least position, orientation and size attributes.

3. The method of claim 1, wherein, The public node identifier refers to the name of the attribute, and the first value is the attribute value of the attribute, and determining the plurality of subspaces includes: The first value is retrieved from the common node in the first layer as the attribute value for the attribute of the subspace in the plurality of subspaces.

4. The method according to claim 1, characterized in that, The public node identifies the name of the attribute and the index of the subfield of the attribute, the first value is the subfield attribute value for the subfield of the attribute, and determining the plurality of subspaces includes: The first value is retrieved from the public node in the first layer as the subfield attribute value of the subfield of the attribute for the subspace in the plurality of subspaces.

5. The method according to claim 1, characterized in that, The public node having the first value is public to the plurality of subspaces, and determining the plurality of subspaces further includes: Retrieve the first value from the common node in the first layer as an attribute value for an attribute in each of the plurality of subspaces.

6. The method according to claim 1, characterized in that, The common nodes having the first value are common to subsets of the plurality of subspaces, and determining the plurality of subspaces further includes: In response to a lack of a value for an attribute in a first individual node associated with the first subspace, the first value is retrieved from the common nodes in the first layer as the attribute value for the attribute in the first subspace; and In response to the existence of a second value associated with an attribute in a second separate node associated with the second subspace, the second value associated with the attribute for the second subspace is retrieved from the second separate node.

7. The method according to claim 1, characterized in that, The common nodes having the first value are common to subsets of the plurality of subspaces, and determining the plurality of subspaces further includes: In response to the first individual node associated with the first subspace lacking an attribute value, the first value is retrieved from the public node in the first layer as the attribute value of the attribute in the first subspace; Retrieve the difference associated with the attribute of the second subspace from the second separate node associated with the second subspace; and Calculate a second value for the attribute for the second subspace based on the first value and the difference.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Receive a bitstream carrying the audio input and a hierarchical description of the space of interest as metadata for the audio input; and Decode the bitstream to obtain the audio input and a hierarchical description of the space of interest.

9. The method according to any one of claims 1 to 7, characterized in that, The method further includes: In response to the fact that the location of the subject of the audio scene is outside the space of interest, the audio input is ignored and no rendering is performed.

10. A media processing device, characterized in that, The device includes: Memory, used to store instructions; and A processor for invoking instructions stored in the memory to implement the method according to any one of claims 1-9.

11. A non-transitory computer-readable storage medium for storing instructions, characterized in that, When executed by at least one processor, the instructions cause the at least one processor to perform the method according to any one of claims 1-9.