Method and apparatus for transmitting and receiving immersive audio media in wireless communication system supporting distributed rendering

By employing edge servers for distributed rendering and incorporating audio spatial information in RTP packet headers, the challenges of transmitting and receiving immersive audio media are addressed, resulting in reduced power consumption and enhanced immersive audio experiences in wireless communication systems.

WO2025095544A1PCT designated stage expired Publication Date: 2025-05-08SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016692
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-10-29
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current wireless communication systems face challenges in efficiently transmitting and receiving immersive audio media, particularly for 3D media, due to high battery consumption and complex rendering processes, and existing methods struggle to effectively introduce 3D audio transmission for immersive experiences.

Method used

The proposed solution involves using edge servers for distributed rendering, where audio spatial information is extracted and included in the RTP packet headers, enabling efficient transmission and reception of immersive audio media. This method utilizes existing stereo-based workflows and metadata to facilitate 3D ambisonic audio rendering.

Benefits of technology

This approach reduces power consumption on terminals, enhances the efficiency of immersive audio media transmission, and enables effective 3D ambisonic audio rendering, providing a more immersive experience in wireless communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016692_08052025_PF_FP_ABST
    Figure KR2024016692_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and an apparatus for transmitting and receiving immersive audio media in a wireless communication system supporting distributed rendering. A method performed by a transmission device for transmitting audio data in a communication system according to an embodiment of the present disclosure comprises the processes of: extracting, from at least one sound source, audio spatial information including location information of the sound source; and transmitting, to an application server, an RTP packet in which a payload includes audio data encoded from the at least one sound source and a header includes the audio spatial information, wherein on the basis of the audio spatial information, audio data to be provided to a reception device is split-rendered into ambisonic audio by the application server.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for transmitting and receiving immersive audio media in a wireless communication system supporting distributed rendering

[0001] The present disclosure relates to a method and apparatus for providing immersive audio media in a wireless communication system.

[0002] The recent proliferation of 5G communication systems has led to a growing demand for services such as high-capacity media like 3D and real-time media transmission. Furthermore, new types of devices, such as augmented reality (AR) glasses, are emerging to effectively display 3D media, rather than traditional smartphones. These devices require lightweight designs to ensure comfortable wear. However, 3D media requires significantly more complex rendering processing than conventional 2D methods, which leads to increased battery consumption.

[0003] Edge computing is a proposed technology that allows operators and / or third-party services to be hosted close to access points, such as base stations, thereby reducing end-to-end network latency and load, enabling efficient service provision. This edge computing technology reduces data processing time by processing data generated from UEs in real time at the site where the data was generated, rather than transmitting it to a central cloud network (hereinafter referred to as the "central cloud"). For example, edge computing can be applied to technologies such as autonomous vehicles, which require rapid processing in various situations that may occur during driving. Edge computing is a network architecture concept that enables cloud computing functions and service environments, and the network for edge computing can be deployed close to the UE. Edge computing offers advantages such as reduced latency, increased bandwidth, reduced backhaul traffic, and the prospect of new services compared to cloud environments. The 5G or 6G or higher core network (CN) proposed by the 3rd Generation Partnership Project (3GPP) can expose network information and functions to edge computing applications (hereinafter, edge applications). The edge computing can be understood as a technology for Mobile Edge Computing that establishes a data connection to an Edge Data Network (EDN) and accesses an Edge Application Server (EAS) running on an Edge Hosting Environment or Edge Computing Platform operated by an Edge Enabler Server (EES) of the Edge Data Network (EDN) to use data services.

[0004] An alternative to addressing these issues is utilizing edge servers (e.g., the edge application servers) in edge networks. This approach utilizes edge servers, which are capable of complex computations compared to AR glasses, to share the rendering process between the device and the edge server, reducing power consumption on the device while enabling effective 3D media rendering. This split-process rendering method is called distributed (split) rendering, and the series of processes required for this is called a workflow.

[0005] While various technologies and standards are being discussed for distributed rendering for video packet transmission, among the video and audio data that constitute media, there is little research on audio packets. Specifically, for video packets, methods and devices for pre-rendering 3D video into 2D by transmitting pose information (e.g., position and orientation) of the receiver to an edge server are common. However, for audio packets, the prevailing stereo-based workflow makes it difficult to introduce a 3D audio transmission method for immersive viewing.

[0006] The present disclosure provides a method and apparatus for efficiently transmitting / receiving immersive audio media in a wireless communication system.

[0007] Additionally, the present disclosure provides a method and device for efficiently transmitting / receiving immersive audio media in a wireless communication system supporting split rendering.

[0008] Additionally, the present disclosure provides a method and device for transmitting / receiving immersive audio media using audio spatial information of a sound source in a wireless communication system.

[0009] A method performed in a transmitting device for transmitting audio data in a communication system according to an embodiment of the present disclosure includes the steps of extracting audio spatial information including location information of the sound source from at least one sound source, and transmitting an RTP (real-time transport protocol) packet including encoded audio data from the at least one sound source in a payload and the audio spatial information in a header to an application server, wherein audio data to be provided to a receiving device is split-rendered as ambisonic audio by the application server based on the audio spatial information.

[0010] In addition, in a communication system according to an embodiment of the disclosure, a transmitting device for transmitting audio data includes a transceiver, and a processor configured to extract audio spatial information including location information of the sound source from at least one sound source, and transmit an RTP packet including encoded audio data from the at least one sound source in a payload and including the audio spatial information in a header to an application server through the transceiver, wherein audio data to be provided to a receiving device is distributedly rendered as ambisonic audio in the application server based on the audio spatial information.

[0011] According to various embodiments of the present disclosure, a device and method can be provided that enable more efficient 3D Ambisonic audio rendering in a receiving device by utilizing metadata such as separately transmitted audio spatial information. Since the corresponding metadata is included in the RTP (extension) header as described above and transmitted in real time together with the actual audio packet, the receiving device can utilize this audio spatial information to effectively render Ambisonic audio, or perform Ambisonic audio rendering through split rendering in an application server.

[0012] Figure 1 is a diagram illustrating a workflow for Ambisonic audio acquisition and playback.

[0013] FIG. 2 is a diagram illustrating an example of a system structure that provides audio spatial information for transmission and reception of immersive audio media in a wireless communication system according to an embodiment of the present disclosure.

[0014] FIG. 3 is a diagram showing an example configuration of a pre-processor (processing) in a transmitting device that transmits immersive audio media according to an embodiment of the present disclosure.

[0015] FIG. 4 is a diagram showing an example configuration of a post-processor (processing) in a receiving device that receives immersive audio media according to an embodiment of the present disclosure.

[0016] FIG. 5 is a diagram illustrating a procedure for split rendering of stereo-based audio media in a wireless communication system according to an embodiment of the present disclosure.

[0017] FIG. 6 is a diagram for explaining audio spatial information according to an embodiment of the present disclosure.

[0018] FIG. 7 is a diagram illustrating an example of an RTP (extension) header structure including audio spatial information according to an embodiment of the present disclosure.

[0019] FIG. 8 is a diagram showing an example of reference points of left / right directional angles and up / down directional angles included in audio spatial information according to an embodiment of the present disclosure.

[0020] FIG. 9 is a diagram illustrating another example of an RTP (extension) header structure including audio spatial information according to an embodiment of the present disclosure.

[0021] FIG. 10 is a diagram showing an example configuration of an electronic device in a wireless communication system according to an embodiment of the present disclosure.

[0022] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the attached drawings. It should be noted that, where possible, identical components are represented by identical reference numerals throughout the attached drawings. Furthermore, detailed descriptions of well-known functions and configurations that may obscure the gist of the present disclosure will be omitted.

[0023] Electronic devices according to various embodiments of the present disclosure may be various types of electronic devices. Electronic devices may include, for example, portable communication devices (e.g., smartphones), portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to embodiments of the present disclosure are not limited to the aforementioned devices. It should be understood that the various embodiments of the present disclosure and the terminology used therein are not intended to limit the technical features described in this document to a specific embodiment, but include various modifications, equivalents, or alternatives of the embodiment. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the item, unless the relevant context clearly indicates otherwise.

[0024] In this disclosure, phrases such as "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can each include any one of the items listed together in that phrase, or all possible combinations thereof. Terms such as "first", "second", or "first" or "second" may be used merely to distinguish the corresponding component from other corresponding components and do not limit the corresponding components in any other respect (e.g., importance or order). When a (e.g., a first) component is referred to as "coupled" or "connected" to another (e.g., a second) component, with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0025] The term "module" as used herein may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integrally configured component or a minimum unit or part of the component that performs one or more functions. According to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0026] Various embodiments of the present disclosure may be implemented as software (e.g., a program) including one or more instructions stored in a storage medium (e.g., built-in memory or external memory) readable by the electronic device. For example, a processor of the electronic device may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the device to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The storage medium readable by the electronic device may be provided in the form of a non-transitory storage medium. Here, "non-transitory" only means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0027] According to one embodiment, the method according to the various embodiments disclosed in the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., CD-ROM, DVD-ROM), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0028] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include a single or multiple entities. According to various embodiments, one or more components or operations of the above-described components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each component of the plurality of components identically or similarly to those performed by the corresponding component of the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0029] The 5G network technology and edge computing technology described in the drawings and description of the present disclosure may refer to standard specifications (e.g., TS 23.558) defined by the International Telecommunication Union (ITU) or 3GPP. In addition, according to embodiments of the present disclosure, an electronic device may refer to various devices used by a user. For example, the electronic device may refer to a user equipment (UE), a mobile station, a subscriber station, a remote terminal, a wireless terminal, or a user device. For convenience, the embodiments of the present disclosure will be described below by exemplifying the electronic device as a terminal (UE).

[0030] Figure 1 is a diagram illustrating a workflow for acquisition and playback of Ambisonic audio.

[0031] Ambisonic is one of the methods for recording and playing back audio in 360-degree surround (i.e., 3D (dimension)). Referring to Fig. 1, in order to implement three-dimensional sound according to not only the left and right directions of the sound source but also the height and distance, the transmitting device (transmitting end) can acquire the sound source using an Ambisonic microphone (101) suitable for 3D audio from the sound source acquisition process. Thereafter, the transmitting device encodes the sound source using an audio codec (102) suitable for Ambisonic and transmits it as an audio packet. The receiving device (receiving end) can render the received audio packet using an Ambisonic decoder (103) to output / implement three-dimensional sound through an Ambisonic speaker (104). At this time, the Ambisonic speaker (104) can be a generally known 5.1 channel speaker or an Ambisonic headset depending on the audio playback environment.

[0032] However, since recording of the sound source in a professional studio is generally required to obtain the sound source through an Ambisonic microphone (101), it is not easy to utilize Ambisonic. Therefore, as a complementary method, the transmitting device encodes the sound source created by the existing well-known stereo microphone (111) and stereo encoder (112) and transmits it as an audio packet. The receiving device that receives the audio packet uses a stereo decoder (113) to decode the audio packet, and the decoded audio data is rendered into 3D Ambisonic audio through a separate post processor (114) and outputted through an Ambisonic speaker (114). This method is being utilized as an easy-to-use workflow for anyone to easily create and play 3D Ambisonic audio. For example, when conducting a video conference between multiple users wearing AR glasses, even if the AR glasses do not have a separate Ambisonic microphone, the AR glasses on the transmitting side can acquire surrounding sound sources, encode them, and transmit them as audio packets, and the AR glasses on the receiving side participating in the video conference can render the audio packets as 3D Ambisonic audio, enabling spatial three-dimensional sound configuration as if listening from the transmitting side's position.

[0033] The stereo-based 3D audio workflow (111 to 115) described in Fig. 1 has the advantage of being able to simply implement 3D audio by rendering 3D Ambisonic audio through separate post processing on the receiving device even without audio equipment used in a professional studio. However, it requires a separate processor requiring a high amount of computation, such as a post-processor (114) that performs rendering. In addition, since the original sound source itself implements virtual 3D audio based on left / right stereo sound sources, it inevitably has the imperfection of not being able to solve the rendering problem with an algorithm alone. Therefore, the present disclosure proposes a method of providing metadata including separate audio spatial information on a transmitting device so as to facilitate 3D Ambisonic audio playback on a receiving device while utilizing an existing stereo-based workflow.

[0034] FIG. 2 is a diagram illustrating an example of a system structure for providing audio spatial information for transmitting and receiving immersive audio media in a wireless communication system according to an embodiment of the present disclosure. In the present disclosure, audio media may be referred to by various names such as audio data, audio packets, and audio streams. The system of FIG. 2 includes a transmitting device (200) for transmitting an ambisonic-based audio stream and a receiving device (210) for receiving an ambisonic-based audio stream.

[0035] Referring to FIG. 2, a transmitting device (200) can acquire stereo data through a stereo microphone array (201) including a plurality of stereo microphones. A pre-processor (202) extracts (audio) spatial information including positional information of a sound source (e.g., azimuth angle, elevation angle, etc.) from the stereo data and outputs it as an encapsulation (204) (206). In addition, the pre-processor (202) outputs an audio stream corresponding to the stereo data to a stereo encoder (203) (205). The stereo encoder (203) compresses the audio stream and outputs the compressed audio stream as an encapsulation (204). The stereo encoder (203) can use a known encoder. Encapsulation (204) generates an RTP (Real-time Transport Protocol) packet that includes the compressed audio stream in the payload. At this time, encapsulation (204) can generate an RTP packet by inserting the spatial information into the header of the RTP packet. The transmitting device (200) can transmit the RTP packet to the network through a transmitter (or transceiver) not shown. Encapsulation

[0036] A receiving device (210) receives an RTP packet through a receiver (or transceiver) not shown, and decapsulation (211) separates header information including the spatial information and an RTP payload including a compressed audio stream from the received RTP packet. At this time, decapsulation (211) outputs the RTP payload to a stereo decoder (212) (220), and outputs the spatial information included in the header information to a post-processor (213) (221). The stereo decoder (212) restores the compressed audio stream included in the RTP payload into a stereo audio stream and outputs the same to the post-processor (213). The post-processor (213) can use the (audio) spatial information to (re)configure the stereo audio stream into 3D ambisonic audio. The ambisonic audio reconstructed in this way by the receiving device (210) is output as a three-dimensional sound through the left and right speakers (216, 217) through the HRTF (Head-Related Transfer Function) (214, 215) to match the left and right ears when the user uses a headset.

[0037] FIG. 3 is a diagram illustrating an example configuration of a pre-processor (processing) in a transmitting device that transmits immersive audio media according to an embodiment of the present disclosure. FIG. 3 illustrates an example configuration of the pre-processor (202) in FIG. 2.

[0038] Referring to FIG. 3, an audio stream recorded in real time from a plurality of microphones of a stereo microphone array (201) in a transmitting device (200) is output as a stereo audio stream to be transmitted to a receiving device (210) through a mixer (301). The output stereo audio stream (i.e., mix signal) undergoes general audio signal processing, such as noise canceling, through an audio processing unit (302) and is transmitted to a stereo encoder (203). Meanwhile, a spatial information extraction unit (303) extracts audio spatial information including at least one of an azimuth angle and an elevation angle of a sound source from audio signals collected from a plurality of microphones of a stereo microphone array (201), and transmits the extracted audio spatial information to an encapsulation (204). The audio spatial information is inserted into an RTP header of an RTP packet through an encapsulation (204).

[0039] FIG. 4 is a diagram illustrating an example configuration of a post-processor (processing) in a receiving device for receiving immersive audio media according to an embodiment of the present disclosure. FIG. 4 illustrates an example configuration of the post-processor (213) in FIG. 2.

[0040] Referring to FIG. 4, a stereo audio stream restored through a stereo decoder (212) in a receiving device (210) is converted into a surround sound signal through an Ambisonic rendering unit (401). At this time, the Ambisonic rendering unit (401) receives the audio spatial information separately transmitted from the transmitting device (200) through decapsulation (211) and performs Ambisonic rendering to convert the stereo audio stream into a surround sound signal. In the Ambisonic rendering process, the receiving device (210) can additionally utilize audio connection information (Speaker configuration). For example, the receiving device (210) can use internal device information to determine which speaker system is currently being used, so if a 5.1 channel speaker is connected to the receiving device (210), the stereophonic sound signal is configured as a 5.1 audio stream corresponding thereto, or if an ambisonic headset (not shown) is connected, signals for the left and right speakers corresponding thereto are configured through the Demux (402) and output to the speakers (216, 217).

[0041] The above-described post-processor (213) may be included in the receiving device (210) as in the embodiment of FIG. 2, or may be included in an edge server (e.g., an application server) of an edge network for distributing the computational processing of the receiving device (210) as in the embodiment to be described later. Distributing the function corresponding to the post-processor (213) to the application server of the network in this way is referred to as split rendering for audio services. When split rendering is used, after the edge server (e.g., the application server) renders stereo audio into ambisonic audio, the edge server can transmit the processed ambisonic audio to the receiving device. As an optional embodiment, the edge server can perform stereo re-encoding on the ambisonic audio and transmit it to the receiving device.

[0042] FIG. 5 is a diagram illustrating a procedure for split rendering of stereo-based audio media in a wireless communication system according to an embodiment of the present disclosure. In the example of FIG. 5, UE1 may correspond to the transmitting device (200), UE2 may correspond to the receiving device (210), and RTC AS (Real-Time Communication Application Server) may correspond to the Edge server (e.g., application server). Although the Edge server is exemplified as the RTC AS in the present disclosure, the RTC AS is not limited to the Edge server and may be various types of application servers. The RTC AS may communicate with the transmitting device (200) and the receiving device (210) via a network such as a core network of a 5G system (not shown), a local network, or the Internet. The embodiment of FIG. 5 may be performed in combination with at least one of the embodiments of FIGS. 2 to 4 described above.

[0043] Referring to FIG. 5, a method of playing back 3D sound in UE2 by split-rendering audio acquired from UE1 through an Edge server corresponding to an RTC AS will be described. In step 501, UE1 can find network address information to connect to counterpart UE2 in, for example, video conferencing, online games, augmented reality services, etc. (remote endpoint discovery) through RTC AS (e.g., WebRTC). Through the remote endpoint discovery, UE1 can inquire about RTC AS and acquire the network address of UE2 from RTC AS. In step 502, media configuration information to be used by UE1 can be transmitted to UE2 through SDP (Session Description Protocol). The media configuration information may include, for example, at least one of codec information used for media (e.g., AVC (advanced video coding), HEVC (High Efficiency Video Coding), etc.) and resolution information. At this time, UE1 can propose / inquire UE2 to send the audio spatial information according to the present disclosure. In step 503, UE2 selects a mode appropriate for the reception environment based on the media configuration information offered by UE1 and responds to UE1. The mode appropriate for the reception environment may be distinguished, for example, based on whether audio spatial information is used, and the SDP answer to UE1 may include information indicating whether audio spatial information is used. In this process, if an ambisonic speaker is connected to UE2 (e.g., 5.(e.g., 1-channel speaker or ambisonic headset) to the transmitting device UE1 to confirm / respond to the SDP answer and confirm that UE1 includes audio spatial information in the RTP packet and transmits it, and if not, use of the audio spatial information can be rejected. Hereinafter, the procedure of the present disclosure will be described assuming that UE2 confirms the use of audio spatial information to UE1. Step 504 is a process of configuring an Edge server for post-processing for ambisonic rendering. In step 504, UE2 requests the Edge server for the resources necessary for configuring a surround sound and goes through a process of confirming the available resources from the RTC AS. After the above preparation steps are completed, in step 505, the RTC AS notifies UE1 that all setups for split rendering are complete. In step 505, the RTC AS can selectively send configuration information related to Split rendering with the configured UE2 to UE1. UE1 can configure the necessary audio space information based on the configuration information returned from RTC AS regarding Split rendering.

[0044] When a Web-based RTC connection (WebRTC connection) is established between UE1, UE2, and RTC AS according to the above process, an audio stream can be transmitted from UE1 to UE2 through RTC AS through the process of steps 506 to 510. Specifically, in step 506, UE1 extracts necessary audio spatial information through audio pre-processing described in the examples of FIGS. 2 and 3, and transmits the audio spatial information to RTC AS by including it in an RTP header in step 507. In step 508, RTC AS performs post-processing described in the examples of FIGS. 2 and 4 to render stereo audio into an ambisonic audio signal, and compresses the ambisonic audio signal into an audio stream required to transmit it to UE2. In the example of FIG. 5, it is assumed that UE2, which is a receiving device, uses an ambisonic headset. In this case, since UE2 needs a stereo signal to be transmitted to each left and right speaker of the ambisonic headset, in step 509, the RTC AS re-encodes the ambisonic audio signal back into a stereo audio signal, and in step 510, transmits the re-encoded stereo audio signal to UE2. In step 511, UE2, which receives the re-encoded stereo audio signal, simply transmits the received audio signal to the left and right speakers without separate processing for ambisonic sound, thereby playing / outputting ambisonic sound.

[0045] FIG. 6 is a diagram illustrating audio spatial information according to an embodiment of the present disclosure. Various sound sources may exist around a transmitting device, and audio transmitted from these various sound sources is mixed and recorded by at least one microphone in the transmitting device. For example, in an environment such as a smartphone with one or two existing microphones, it is impossible to separate spatial information from such mixed sound sources. However, in an environment such as AR glasses where multiple microphones are available, it is possible to separate spatial information of each sound source even if it is a mixed signal. Here, the spatial information of the sound source refers to the location information of the sound source as described above, and at least one of the azimuth angle (612) and the elevation angle (611) can be estimated as the location information of the sound source by using the time difference in which the sound sources (601, 602) reach each microphone.

[0046] FIG. 7 is a diagram illustrating an example of an RTP (extension) header structure including audio spatial information according to an embodiment of the present disclosure.

[0047] Referring to FIG. 7, the length of the extended header is indicated using the Length field of the existing 32-bit RTP basic header, and the example of FIG. 7 shows a case where spatial information for two sound sources is included in the RTP (extended) header. The RTP (extended) header is an example of 8-bit local IDs (ID#1, ID#2) of each sound source, local lengths of spatial information for each sound source (Length#1, #2), and left-right directional angles (Azimuth#1, Azimuth#2) and up-down directional angles (Elevation#1, Elevation#2) of each sound source in 16 bits.

[0048] FIG. 8 is a diagram showing an example of reference points of left / right directional angles and up / down directional angles included in audio spatial information according to an embodiment of the present disclosure.

[0049] Referring to Fig. 8, in the case of left / right directional angles, the direction facing the user of the transmitting device (e.g., AR glasses) is set to 0 degrees, and +180 degrees are displayed in a clockwise direction. Therefore, the range of left / right directional angle values ​​can have a range of values ​​from -180 degrees to +180 degrees. Similarly, in the case of up / down directional angles, the range of values ​​can be from -90 degrees to +90 degrees, with the direction parallel to the ground surface being set to 0 degrees.

[0050] FIG. 9 is a diagram illustrating another example of an RTP (extension) header structure including audio spatial information according to an embodiment of the present disclosure.

[0051] Referring to Fig. 9, in addition to the directional angle for each local sound source, a polar response value (Polar response) for the corresponding directional angle may be included in the RTP (extended) header. The example of Fig. 9 illustrates a case in which Polar response values ​​(Polar response#1, Polar response#2) for two sound sources are additionally included in the RTP (extended) header of Fig. 7. The Polar response value is a structural intrinsic parameter of the microphone and represents gain information of the microphone for a specific direction. In general, since a microphone has a specific pattern rather than being omnidirectional with the same gain value for all directions, transmitting this polar response information to a receiving device or an edge server enables more precise stereophonic sound reproduction / output based on the polar response information.

[0052] [Table 1] below illustrates an example of SDP negotiation / negotiation parameters according to the present disclosure. The SDP negotiation / negotiation parameters may be included in the SDP offer / answer transmitted and received between UE1 and UE2 in steps 502 and 503 of the procedure of FIG. 5.

[0053]

[0054] Referring to the above [Table 1], an example of the RTP (Real-time Transport Protocol) / AVP (Audio VIdeo Profile) protocol is shown through an m-line related to Audio, and may include content for transmitting audio spatial information defined in 3gpp:xr-audio-doa in the 3GPP standard as an example by the present disclosure. When UE1 transmits an SDP offer including the corresponding attribute, and UE2 replies an SDP answer including the corresponding line to UE1 if 3D sound reproduction is possible, the use of audio spatial information can be negotiated between the transmitting device (UE1) and the receiving device (UE2).

[0055] According to various embodiments of the present disclosure, a method and device can be provided that enable more efficient 3D Ambisonic audio rendering in a receiving device by utilizing metadata such as separately transmitted audio spatial information. Since the corresponding metadata is included in the RTP (extension) header as described above and transmitted in real time together with the actual audio packet, the receiving device can utilize this audio spatial information to effectively perform Ambisonic audio rendering, or perform Ambisonic audio rendering through split rendering in an application server. The technical effects obtainable in the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.

[0056] FIG. 10 is a diagram illustrating an example configuration of an electronic device in a wireless communication system according to an embodiment of the present disclosure. The electronic device may be one of the transmitting device, receiving device, or application server.

[0057] The electronic device of FIG. 10 may include a processor (1001), a transceiver (1003), and a memory (1005). The processor (1001), the transceiver (1003), and the memory (1005) of the electronic device may operate to transmit and receive immersive audio media according to the configurations and methods described in the embodiments of FIGS. 2 to 9. However, the components of the electronic device are not limited to the examples described above. For example, the electronic device may include more or fewer components than the components described above. In addition, the processor (1001), the transceiver (1003), and the memory (1005) may be implemented in the form of a single chip.

[0058] The transceiver (1003) is a general term for a receiver and transmitter of an electronic device, and can transmit and receive signals with a counterpart electronic device. At this time, the transmitted and received signals may include at least one of control information and data. In addition, the transceiver (1003) may receive a signal, output it to the processor (1001), and transmit the signal output from the processor (1001). In addition, the transceiver (1003) of FIG. 10 may include an RF transmitter that up-converts and amplifies the frequency of a transmitted signal, and an RF receiver that low-noise amplifies and frequency-downconverts the received signal. In addition, the transceiver (1003) may include a communication interface for performing wired communication on a network when the electronic device is an application server. In addition, the transceiver (1003) may receive a signal, output it to the processor (1001), and transmit the signal output from the processor (1001) to the counterpart electronic device through a network. The memory (1005) can store programs and data required to perform the procedures / operations described in the embodiments of FIGS. 2 to 9. The memory (1005) can store control information or data included in signals acquired from the electronic device. The memory (1005) can be configured as a storage medium or a combination of storage media, such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD. In addition, the processor (1001) can control a series of processes so that the electronic device can operate according to at least one of the embodiments of FIGS. 2 to 9. In addition, the processor (1001) can perform the function of at least one of the components exemplified in the embodiments of FIGS. 2 to 4. The processor (1001) can include at least one processor.

[0059] The methods according to the embodiments described in the claims or specification of the present disclosure may be implemented in the form of hardware, software, or a combination of hardware and software. If implemented in software, a computer-readable storage medium storing one or more programs (software modules) may be provided. The one or more programs stored in the computer-readable storage medium are configured for execution by one or more processors within an electronic device. The one or more programs include instructions that cause the electronic device to execute the methods according to the embodiments described in the claims or specification of the present disclosure.

[0060] These programs (software modules, software) may be stored in random access memory, non-volatile memory including flash memory, read only memory (ROM), electrically erasable programmable read only memory (EEPROM), magnetic disc storage devices, compact disc-ROMs (CD-ROMs), digital versatile discs (DVDs) or other forms of optical storage devices, magnetic cassettes, or may be stored in memories formed by a combination of some or all of these. In addition, each configuration memory may include multiple copies. The above program may be stored on an attachable storage device that is accessible via a communication network such as the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a storage area network (SAN), or a combination thereof. This storage device may be connected to a device performing an embodiment of the present disclosure via an external port. Additionally, a separate storage device on the communication network may be connected to a device performing an embodiment of the present disclosure.

[0061] In the specific embodiments of the present disclosure described above, components included in the disclosure are expressed in the singular or plural form, depending on the specific embodiment presented. However, the singular or plural expressions are selected to suit the presented situation for convenience of explanation, and the present disclosure is not limited to singular or plural components. Components expressed in the plural form may be composed of singular elements, or components expressed in the singular form may be composed of plural elements.

[0062] The embodiments disclosed in this specification and drawings above are merely specific examples presented to facilitate easy explanation and understanding of the contents of the present disclosure, and are not intended to limit the scope of the present disclosure. Therefore, the scope of the present disclosure should be interpreted to include all modifications or variations derived based on the present disclosure, in addition to the embodiments disclosed herein.

Claims

1. A method performed in a transmitting device for transmitting audio data in a communication system, A process of extracting audio spatial information including location information of the sound source from at least one sound source; and A process for transmitting an RTP (real-time transport protocol) packet including encoded audio data from at least one sound source in a payload and including the audio spatial information in a header to an application server, A method for split rendering audio data to be provided to a receiving device as ambisonic audio from the application server based on the above audio spatial information.

2. In paragraph 1, A method wherein the location information includes at least one of an azimuth angle and an elevation angle for the at least one sound source.

3. In paragraph 1, A method wherein the above audio spatial information further includes a polar response for the at least one sound source, the polar response representing a gain of a microphone of the transmitting device.

4. In paragraph 1, A method further comprising a process in which the transmitting device performs negotiation with a receiving device from which the audio data is to be received regarding whether to use the audio spatial information.

5. In paragraph 1, The above extraction process is, A process of collecting audio signals from a plurality of microphones of the transmitting device for at least one sound source; and A method comprising a process of extracting at least one of an azimuth angle and an elevation angle of at least one sound source from the audio signals.

6. In a transmitting device that transmits audio data in a communication system, Transmitter and receiver; and Extracting audio spatial information including location information of the sound source from at least one sound source, A processor configured to transmit, through the transceiver, to an application server, an RTP (real-time transport protocol) packet including encoded audio data from at least one sound source in a payload and including the audio spatial information in a header, A transmitting device in which audio data to be provided to a receiving device is split-rendered as ambisonic audio by the application server based on the above audio spatial information.

7. In paragraph 6, A transmitting device wherein the above location information includes at least one of an azimuth angle and an elevation angle for the at least one sound source.

8. In paragraph 6, A transmitting device, wherein said audio spatial information further includes a polar response for said at least one sound source, said polar response representing a gain of a microphone of said transmitting device.

9. In paragraph 6, The above processor is a transmitting device further configured to perform negotiation with a receiving device through the transceiver regarding whether to use the audio space information with respect to which the audio data is to be received.

10. In paragraph 6, The above processor, For at least one sound source, audio signals are collected from a plurality of microphones of the transmitting device, A transmitting device configured to extract at least one of an azimuth angle and an elevation angle of at least one sound source from the audio signals.

11. A method performed on an application server for transmitting and receiving audio data in a communication system, A process for receiving, from a transmitting device, an RTP (real-time transport protocol) packet including encoded audio data from at least one sound source in its payload and audio spatial information including location information of the sound source in its header; and A method comprising a process of split rendering audio data into ambisonic audio using the above audio spatial information.

12. In paragraph 11, The method performed on the above application server is: A process of generating an ambisonic audio signal by compressing the above ambisonic audio into an audio stream; and A method comprising the step of transmitting the above ambisonic audio signal to a receiving device.

13. In paragraph 11, The method performed on the above application server is: A process of generating an ambisonic audio signal by compressing the above ambisonic audio into an audio stream; A process of re-encoding the above ambisonic audio signal to generate a stereo audio signal transmitted to the left speaker and / or the right speaker; and A method comprising the step of transmitting the stereo audio signal to a receiving device.

14. In paragraph 11, A method wherein the location information includes at least one of an azimuth angle and an elevation angle for the at least one sound source.

15. In paragraph 11, A method wherein the above audio spatial information further includes a polar response for the at least one sound source, the polar response representing a gain of a microphone of the transmitting device.

Citation Information

Patent Citations

  • Spatial Audio for Two-Way Audio Environments

    JP2021528001A

  • Methods for learning algebra in elementary and middle school curriculum by individual ability using mobile applications

    KR1020220123821A

  • Biodegradable mask sheet, biodegradable mask with same, and manufacutirng method thereof

    KR1020230009558A

  • Spade studs soccer shoes

    KR1020230040766A

  • Methods and systems for immersive 3DOF / 6DOF audio rendering

    WO2023187208A1